Meta simply launched Muse Spark. The announcement says it beats GPT-5.4 on well being duties, ranks top-five globally on the Artificial Analysis Intelligence Index, and scores 89.5% on one thing referred to as GPQA Diamond.
Eleven months in the past, Meta mentioned nearly an identical issues about Llama 4, earlier than individuals really used it and the numbers collapsed.
So what are these benchmarks? How do the scores get calculated? And why does a mannequin that tops each leaderboard typically really feel mediocre the second you utilize it?
This information explains what the largest AI benchmarks really measure, together with MMLU, GPQA Diamond, HumanEval, SWE-bench, HealthBench, Humanity’s Last Exam, and Chatbot Arena. It additionally explains how benchmark scores are calculated, why some exams matter greater than others, and the way AI labs can inflate benchmark outcomes with out enhancing real-world efficiency.
What Is an AI Benchmark?
A benchmark is only a standardized check. A set set of questions or duties, given to each AI mannequin in the identical means, scored the identical means. The thought is that if everybody takes the identical check, you may examine the outcomes pretty. But there is a apply the AI group has began calling benchmaxxxing: squeezing each potential level out of a benchmark by analysis decisions, cherrypicked settings, and coaching methods that enhance the rating with out essentially enhancing the mannequin.
We’ll get into the specifics of how this works as we undergo every benchmark.
MMLU and MMLU-Pro: The Knowledge Test
What it’s: Over 15,000 multiple-choice questions throughout 57 topics. Law, drugs, chemistry, historical past, economics, laptop science. Four reply decisions per query.
What an precise query appears like:
A 60-year-old man presents with progressive weak spot, hyporeflexia, and fasciculations in each legs. MRI exhibits anterior horn cell degeneration. Which of the next is the most probably prognosis? (A) Multiple sclerosis (B) Amyotrophic lateral sclerosis (C) Guillain-Barré syndrome (D) Myasthenia gravis
The mannequin outputs a letter. The check runner checks if it matches the reply key.
How the rating is calculated: Before every query, the mannequin is proven 5 instance questions with appropriate solutions, that is referred to as 5-shot prompting. Then comes the actual query. Score = appropriate solutions ÷ whole questions, expressed as a proportion.
Why it is practically ineffective in 2026: Top fashions now rating above 88% on MMLU. GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro are all bunched collectively above 87%. The check can not separate them, it is like utilizing a toilet scale to measure the burden distinction between two individuals of comparable construct. Technically potential, virtually meaningless.
Researchers responded by constructing MMLU-Pro: similar topics, more durable questions, ten reply decisions as a substitute of 4, with choices designed to look believable even to educated people. On MMLU-Pro, the gaps between fashions begin exhibiting up once more.
→ When you see MMLU in a press launch in 2026, it is principally padding. It’s additionally the benchmark most probably to be inflated by coaching knowledge contamination: fashions have had three years of web knowledge that overlaps closely with MMLU-style questions.
Build your personal no-code agent totally free
Try right here

GPQA Diamond: The Scientific Reasoning Test
This is essentially the most credible educational benchmark in use at the moment. The means it was constructed is what makes it reliable.
How the questions had been made: Researchers employed PhD scientists in biology, physics, and chemistry. Each scientist wrote a query in their very own discipline. Then a second PhD scientist in the identical discipline tried to reply it. If that second knowledgeable obtained it unsuitable, the query handed the filter. Then three extra individuals, good non-domain specialists given limitless web entry and half-hour, tried to reply it. If in addition they failed, the query made it into the Diamond subset.
The end result: 198 questions that require you to truly purpose by exhausting science. You can’t Google them. The solutions aren’t in Wikipedia.
What an precise query appears like:
Two quantum states with energies E1 and E2 have a lifetime of 10⁻⁹ sec and 10⁻⁸ sec, respectively. We wish to clearly distinguish these two power ranges. Which of the next might be their power distinction to allow them to be clearly resolved? (A) 10⁻⁸ eV (B) 10⁻⁹ eV (C) 10⁻⁴ eV (D) 10⁻¹¹ eV
To reply this, it is advisable to know the energy-time uncertainty precept from quantum mechanics, calculate the pure linewidths of the power ranges, and verify which power distinction is massive sufficient to resolve them. The reply is (A), however you may’t discover that by looking. You should derive it.
How the rating is calculated: Same letter-pick system as MMLU. The mannequin is informed to purpose step-by-step and should finish its response with “ANSWER: LETTER” – capital letters solely. If the mannequin would not comply with that actual format, it will get zero for that query no matter whether or not the reasoning was appropriate. This strict formatting rule is intentional: it forces fashions to decide to a selected reply moderately than hedging.
The benchmark in numbers:
- Random guessing: 25% (4 decisions)
- Smart non-experts with web entry: 34%
- PhD-level area specialists: 65%
- GPT-4 when it launched (2023): 39%
- Muse Spark at the moment: 89.5%
- Gemini 3.1 Pro: 94.3%
- Claude Opus 4.6: 92.8%
That soar from 39% to 89% in three years is actual. These fashions have genuinely gotten higher at scientific reasoning. But Muse Spark continues to be about 5 factors behind Gemini on this check, throughout 198 questions. That’s roughly 10 questions. Meta calls this “competitive” which is technically correct.

HumanEval: The Basic Coding Test
What it’s: 164 Python programming issues. Each downside is a perform signature with a docstring explaining what the perform ought to do.
What an precise query appears like:

The mannequin writes the perform physique. An automated check runner then executes the code in opposition to 10-15 hidden check instances, inputs with recognized appropriate outputs. Either each check case passes, or the issue fails.
How the rating is calculated: The principal metric is move@1: did the mannequin’s first try move all of the hidden exams? Score = variety of issues the place the code labored ÷ 164 whole issues.
Example of move vs. fail:
An accurate resolution for the above returns “fl” for [“flower”,”flow”,”flight”] and “” for [“dog”,”racecar”,”car”] and handles edge instances like an empty checklist. A mannequin that hardcodes the seen examples however fails on an edge case like a single-element checklist will get zero for that downside.
Why it is outdated: Top fashions now clear up 90%+ of those 164 issues. They’ve had years to coach on HumanEval-style duties. Researchers brazenly query what number of fashions might have seen these actual issues in coaching. Leading with HumanEval in 2026 is sort of a automotive firm main their security pitch with a check from 2015.
Build your personal no-code agent totally free
Try right here
SWE-bench: The Real Software Engineering Test
What it’s: Real GitHub points from actual open-source repositories. The mannequin is given the difficulty description and the complete codebase and should produce a code patch (a diff) that fixes the bug.
What an precise job appears like:
A developer information a GitHub subject within the sympy math library: “The simplify() function returns the wrong result when called on expressions containing nested Piecewise objects under certain conditions.”
The mannequin will get the difficulty textual content, navigates a codebase with hundreds of information, identifies the supply of the bug, and writes a patch. That patch is mechanically utilized to the codebase, and the present check suite runs to verify that the repair works and did not break the rest.
How the rating is calculated: Pass/fail on the subject degree. Score = proportion of points the place the mannequin’s patch handed all exams.
Why this benchmark issues greater than HumanEval: Because there is not any memorization shortcut. The repositories are actual, the bugs are actual, and the analysis setting is strictly managed. You both mounted the bug otherwise you did not.
Where Muse Spark stands right here: Meta’s personal weblog put up acknowledges “current performance gaps, specifically in coding workflows.” SWE-bench is nearly actually the place that exhibits up. Claude Opus 4.6 at present leads most coding evaluations.

Humanity’s Last Exam: The Frontier Reasoning Test
What it’s: Around 2,500 questions written by researchers particularly designed to exceed what present AI can reply: PhD-level and past, throughout math, science, historical past, and regulation.
Why Muse Spark highlights it: In its “Contemplating” mode, which launches a number of sub-agents working in parallel on totally different components of an issue, Muse Spark scored 50.2%. GPT-5.4 in its highest-effort mode scored 43.9%. Gemini’s Deep Think mode scored 48.4%.
This is Muse Spark’s most authentic lead throughout any benchmark. The hole is actual (6+ factors over GPT-5.4) and the benchmark is genuinely exhausting. One caveat: Contemplating mode makes use of considerably extra compute than a typical response. You’re paying, in time and in API value for that efficiency.
HealthBench: The Clinical Reasoning Test
What it’s: Clinical and medical reasoning duties evaluated by physicians. Questions cowl affected person symptom interpretation, drug interactions, therapy selections, and well being info accuracy.
How the rating is calculated: Unlike automated benchmarks, HealthBench solutions are graded in opposition to physician-defined requirements. The rating represents the share of solutions that met medical accuracy necessities.
The numbers: Muse Spark 42.8%. GPT-5.4 40.1%. Gemini 3.1 Pro 20.6%.
42.8%. GPT-5.4 scored 40.1%. Gemini 3.1 Pro scored 20.6%. This is Muse Spark’s most defensible lead in any benchmark. A 22-point hole over Gemini on a physician-graded check is important.
Build your personal no-code agent totally free
Try right here

Chatbot Arena: The Human Preference Test
This one is totally different from each different benchmark, and understanding the way it works explains the Llama 4 scandal.
What it exams: Whether a human consumer prefers one mannequin’s response over one other.
If Model A beats Model B in 60% of comparisons, Model A will get extra factors. Over time, after sufficient comparisons, the rankings stabilize right into a leaderboard.
Why this benchmark is gameable: Human customers are inclined to favor responses which can be lengthy, confident-sounding, and well-formatted, even when a shorter, extra correct reply would serve them higher. A mannequin that provides enthusiasm, makes use of daring textual content, and provides elaborately structured responses will rating higher on LMArena than a mannequin that provides a direct, appropriate reply in two sentences.
And that is what occurred with Llama 4.
The Llama 4 Incident
When Meta launched Llama 4 in April 2025, its announcement mentioned the mannequin ranked #2 on LMArena, simply behind Gemini 2.5 Pro, with an ELO rating of 1417. That quantity was technically correct, however the mannequin that earned that rating was not the one being launched to the general public.
The mannequin Meta submitted to LMArena was referred to as “Llama-4-Maverick-03-26-Experimental.” Researchers who later in contrast it in opposition to the publicly downloadable model discovered constant behavioral variations:
The experimental model (LMArena): verbose responses, heavy use of emojis, elaborate formatting, dramatic construction, lengthy embellishments even for easy questions.
The public model (what you’d really use): concise, plain, direct, no emojis.
LMArena’s voting system reliably most well-liked the primary type. Real customers in actual use instances most well-liked the second. When the precise public mannequin was individually added to the leaderboard, it ranked thirty second.
There’s one other quantity value realizing: when LMArena turned on Style Control, eradicating the formatting and size benefit, Llama 4 Maverick dropped from 2nd place to fifth. The mannequin’s content material high quality, stripped of its presentational packaging, was a lot much less spectacular.
LMArena acknowledged publicly: “Meta’s interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that ‘Llama-4-Maverick-03-26-Experimental’ was a customized model to optimize for human preference.” They up to date their submission guidelines after.
And on ARC-AGI: a benchmark designed to check real novel reasoning, not sample matching, Llama 4 Maverick scored 4.38% on ARC-AGI-1, and 0.00% on ARC-AGI-2. This was by no means within the press launch.
Build your personal no-code agent totally free
Try right here
How AI Labs Game Benchmark Scores: Goodhart’s Law and Benchmaxxxing

There’s a precept from economics referred to as Goodhart’s Law: when a measure turns into a goal, it stops being measure.
In plain English: the second everybody agrees that GPQA Diamond is the quantity that issues, labs begin optimizing particularly for GPQA Diamond. Scores go up however the real-world functionality might not transfer in any respect.
This has a reputation within the AI group now: benchmaxxxing. It’s the apply of compacting each potential level out of a benchmark by strategies that enhance the rating with out essentially enhancing the mannequin. Some of those strategies are authentic engineering and a few are nearer to the gaming Meta did with LMArena. The line is genuinely blurry, which is a part of what makes this difficult to name out.
That’s how benchmaxxxing really appears like in apply:
Cherry-picking which benchmarks to publish. Every mannequin will get evaluated on dozens of benchmarks internally. The ones that seem within the press launch are those the mannequin did properly on. The relaxation disappear. This is common, each lab does it. Llama 4’s ARC-AGI rating of 0.00% was not within the announcement.
Choosing favorable analysis settings. Many benchmarks could be run in numerous methods: totally different prompting kinds, totally different numbers of instance questions proven beforehand, totally different temperatures. Labs run all of the variants internally and publish the perfect end result. This is technically allowed however hardly ever disclosed.
Training on benchmark-adjacent knowledge. If you understand a benchmark exams quantum mechanics reasoning, you can also make certain your coaching set is heavy on quantum mechanics. The questions themselves aren’t within the coaching knowledge, however the information required to reply them is saturated. This is almost unattainable to tell apart from real functionality enchancment from the skin.
Benchmark contamination, the intense model. Sometimes precise benchmark questions, or near-identical variants, find yourself in coaching knowledge. This can occur unintentionally when coaching on web scrapes. It can even occur much less unintentionally. Susan Zhang, a former Meta AI researcher who later moved to Google DeepThoughts, shared analysis earlier in 2025 documenting how benchmark datasets could be contaminated by coaching corpus overlap. When a mannequin sees the query and reply throughout coaching, it is basically memorized the check. And the rating displays reminiscence, not reasoning.
Majority voting and repeated sampling. Some labs run every benchmark query a number of instances and take the most typical reply. A mannequin that scores 80% on one try may rating 88% throughout 32 makes an attempt. Meta particularly disclosed they do not do that for Muse Spark’s reported numbers, they use zero temperature, single makes an attempt.
The deepest downside with Goodhart’s Law in AI is that it creates a ratchet impact. Each new mannequin must beat the earlier one’s benchmark scores, or it is declared a failure. So each launch will get extra optimized for the benchmarks that exist, which makes these benchmarks much less informative over time, which drives the creation of more durable benchmarks, which then additionally get optimized for. MMLU was the gold normal in 2022 however it’s saturated now. GPQA Diamond changed it.
What Benchmarks Still Can’t Tell You
Speed. GPQA Diamond says nothing about whether or not the mannequin responds in 1 second or 10.
Cost. A mannequin scoring 92% at $15 per million tokens versus one scoring 89% at $1 per million tokens are totally different decisions relying on how a lot quantity you are working.
Consistency. A mannequin averaging 90% on a benchmark however producing catastrophically unsuitable solutions 2% of the time is a unique threat profile from one which scores 85% uniformly. Benchmarks report averages. Averages conceal tails.
Your particular job. None of those benchmarks had been designed in your paperwork, your prompts, or your customers. A mannequin that dominates GPQA Diamond may deal with an insurance coverage kind extraction job worse than a smaller, cheaper mannequin skilled on domain-specific knowledge.
Evaluate AI Models for Your Own Use Case
You can really consider the perfect mannequin for you, your self.
Take your ten or twenty most consultant duties: the precise prompts, paperwork, or questions you’d ship to the mannequin in apply. Run each mannequin you are contemplating on these actual inputs. Score the outputs your self (or have somebody with area experience do it.)
That single customized check will let you know greater than any benchmark desk in a press launch. Because benchmarks let you know the place a mannequin claims to face. Your check set tells you the place it really has to point out up.
Build your personal no-code agent totally free
Try right here
