I opened Nature on January 28 expecting another incremental eval paper and instead got a benchmark that calls itself Humanity's Last Exam. The paper sits at Nature 649, pages 1139 to 1146. Center for AI Safety and Scale AI led it with nearly a thousand expert contributors across more than five hundred institutions in fifty countries. The public pitch is blunt. Older broad academic tests like MMLU have been pushed into the nineties by frontier models, so the field needed a harder closed-ended exam that still grades itself. HLE is 2,500 multimodal questions across more than a hundred subjects, built to be Google-proof, closed-ended, and auto-gradable as multiple choice or short answer. There is a public set and a private holdout. The site is lastexam.ai. Frontier models posted low accuracy at the start. That is the story the authors want you to remember. The story I keep circling is whether naming something the last exam is science or marketing with footnotes.
What the exam actually is
HLE is not a vibes leaderboard and it is not an open-ended writing contest. The questions are closed-ended on purpose. Multiple choice and short answers can be graded automatically, which matters when labs want to rerun the suite every release train without hiring a room of graders. Multimodal means text is not the whole game. Figures, diagrams, and other non-text inputs show up across the subject spread. That is a real step past exams that only reward fluent paragraph generation.
The Google-proof claim is the part evaluators care about most. If a model can retrieve the answer from the open web, you are not measuring reasoning. You are measuring search with better manners. The consortium designed items that resist casual lookup. I cannot audit every item from my desk, and I will not pretend I can. What I can say is the design goal matches the failure mode that killed confidence in older suites once scores crossed ninety percent.
Public plus private holdout is the other structural choice worth naming. A fully public set invites overfitting, prompt templates, and quiet contamination. A private holdout keeps a yardstick that marketing decks cannot memorize. This is not a new idea in machine learning. It is still the right idea when the audience for a score includes investors, procurement teams, and journalists who only read the top line.
Nearly a thousand expert contributors sounds like a census more than a committee. Five hundred-plus institutions and fifty countries are the scale numbers attached to that claim. Broad authorship helps when you are trying to cover more than a hundred subjects without letting one lab's syllabus dominate. It also creates the usual coordination tax. Someone had to decide what counts as expert-hard versus merely obscure.
I keep comparing HLE to the moment a class switches from homework everyone finished early to a final that actually sorts the room. MMLU and friends became that early homework. Hitting ninety percent-plus made them poor instruments for telling frontier systems apart. HLE is an attempt to restore separation with closed-ended academic difficulty rather than with agent tool use or long-horizon tasks.
The Nature publication date, January 28, 2026, matters because it moves the suite out of blog-post limbo into a venue that forces methods sections and peer review. Labs can still game incentives. Peer review does not stop that. It does raise the cost of shipping a selfie benchmark dressed as science.
Why frontier scores started low
The early result that traveled farthest is simple. Frontier models were not close to saturating HLE when the paper landed. Low accuracy is the feature, not the bug, if your goal is headroom. A benchmark that everyone clears in six months is a press cycle, not a measuring stick.
Closed-ended academic hardness is a different skill stack from chat polish. A model can sound confident on a product demo and still miss short answers that demand precise subject knowledge, careful reading of a figure, or a calculation that does not forgive hand-waving. HLE leans into that gap. That is why the initial scores felt like cold water after a year of soft leaderboard inflation.
I am skeptical of treating any single number as the intelligence ranking of record. Accuracy on 2,500 items is still one slice of capability. It does not tell you whether an agent can keep a week-long project coherent. It does not tell you refusal quality under pressure. It does tell you whether the model can clear expert-authored, auto-graded academic obstacles that were built after the last wave of suites got too easy.
The private holdout is especially important once scores start climbing. Public leaderboard chasing is a sport. Holdout drift is how you notice the sport stopped measuring the thing you cared about. If labs publish only public-set wins and stay quiet on holdout, treat that silence as a signal.
lastexam.ai is where the project points outsiders. In a field drowning in PDF-only artifacts, a public face helps. It also invites the inevitable screenshot wars. I would rather argue about a documented suite with a holdout than about anonymous Discord evals with no item control.
For builders, the practical read is calibration. If your internal eval looks solved and HLE looks brutal, your internal eval is probably too friendly. If both look solved, either you are ahead of the public frontier or your harness is contaminated. Those are very different management conversations.
The name is dramatic. The need is real.
Humanity's Last Exam is a title that begs for dunks. Last implies finality. Humanity is doing a lot of rhetorical work for a multiple-choice and short-answer suite. Benchmarks get saturated. Naming one the finale does not stop the next paper from arriving eighteen months later with a harder set and a louder acronym.
I still think the underlying motivation is sound. Broad closed-ended academic benchmarks were dying as discriminators. The field needed a replacement that stayed gradable at scale and hard enough that frontier systems look distinct. HLE is the most visible attempt to be that replacement in early 2026. Dramatic branding does not erase the measurement problem it is trying to solve.
My skeptical checklist is short. Watch for saturation curves. Watch for train-test contamination rumors. Watch whether private holdout results are reported with the same enthusiasm as public wins. Watch whether labs substitute HLE for messy real-world tasks that matter more to users. A great academic exam can become a cargo-cult KPI if product teams worship it.
Useful calibration is the defense. When MMLU stopped separating models, people kept quoting it anyway because the brand was familiar. HLE resets the ceiling. Even if the name ages badly, a harder closed-ended suite with multimodal items and a holdout is better instrumentation than another victory lap on a solved test.
I will keep using HLE the way I use any serious eval. As one instrument among several, not as prophecy. If a lab claims general expertise and posts weak HLE numbers, I want an explanation. If a lab posts strong HLE numbers and ships a product that fails basic tool use, I want a different explanation. Scores are evidence. They are not the whole case.
For now, Nature publication, a huge contributor network, 2,500 multimodal questions, and initially low frontier accuracy make HLE worth tracking. Just do not let the title do your thinking. The last exam narrative sells. The holdout and the item design are the parts that earn trust.
- LLMs




