What MMLU-Pro questions actually look like

dgx-spark

Every quality figure on this site — every accuracy percentage in every model comparison — comes from MMLU-Pro. It is worth knowing what that actually means, because “scored 74.9% on MMLU-Pro” is a sentence most people read without ever having seen one of the questions.

This post is the reference the other articles link to. It shows real questions from all fourteen categories, taken from the same runs the published scores come from, along with which of the models actually got them right.

The shape of it

MMLU-Pro is a 12,032-question multiple-choice benchmark across fourteen categories, built as a harder replacement for the original MMLU after models started saturating it. Every sweep here samples 200 questions per category, which is 2,800 questions per model.

Here is a complete question, exactly as the model receives it:

Polio can be eradicated by which of the following?

A. Herbal remedies B. Use of antibiotics C. Regular intake of vitamins D. Administration of tetanus vaccine E. Attention to sewage control and hygiene F. Natural immunity acquired through exposure G. Use of antiviral drugs

The gold answer is E, and every run we have per-question logs for answered E — four models, five sweeps, unanimous. That is what an easy question looks like: one plausible answer, six that are wrong in obvious ways.

Three things that change how you read a score

1. It is not really ten options

MMLU-Pro is described everywhere as a ten-option benchmark — that is its main advertised difference from MMLU’s four. Across the 2,800 questions in our sweeps, that holds for 81.5% of them:

Options Questions Share
10 2,282 81.5%
9 182 6.5%
4 162 5.8%
8 83 3.0%
7 43 1.5%
6 22 0.8%
5 19 0.7%
3 7 0.2%

Nearly one question in five has fewer than ten. Some have three. That matters for the floor: a model guessing uniformly at random scores 11.26%, not the 10% the framing implies. It is a small correction, but it is the number a “this model is barely above chance” claim should be measured against.

2. The questions come from different places, and it shows

MMLU-Pro is an aggregate. Its questions carry a src field naming where each came from — the original MMLU, STEMEZ, SciBench, TheoremQA — and the mix is wildly uneven by category:

Category Source mix Tokens per answer
engineering 94% STEMEZ 2,413
business 64% STEMEZ 1,587
biology 60% STEMEZ 1,497
chemistry 57% STEMEZ 1,859
economics 54% STEMEZ 1,365
physics 42% STEMEZ 1,739
psychology 33% STEMEZ 1,293
computer science 14% STEMEZ 1,562
law 100% original MMLU 2,181
history 100% original MMLU 1,577
philosophy 100% original MMLU 1,513
health 100% original MMLU 1,449
math 100% original MMLU 1,440
other 100% original MMLU 1,258

Tokens-per-answer is Nemotron 3.5 across the full fourteen-subject sweep. The STEMEZ questions are worked engineering problems — the model has to actually compute something — and engineering, at 94% STEMEZ, produces the longest answers of any category by a wide margin.

This is not trivia. Our benchmark caps generated answers at 3,072 tokens, and on engineering Nemotron ran past that cap on 47% of questions, getting cut off mid-calculation and scored wrong regardless of what it knew. Its engineering score of 45.0% is 84.9% among the answers it was allowed to finish. A category’s source mix predicts how badly a token budget will distort its score.

Law is the interesting exception: no STEMEZ at all, yet the second-longest answers. Legal questions are long to read as well as to reason about.

3. Wrong does not mean ignorant

Some wrong answers are wrong for reasons worth understanding. Here is a chemistry question two of the three models got wrong. It is twenty-three characters long:

The simplest alkene has

A. cis-trans isomers B. at least four sigma bonds C. no sigma bonds D. a trigonal planar configuration E. a tetrahedral configuration F. at least three pi bonds G. at least two pi bonds H. aromatic properties I. at least six sigma bonds

The gold answer is B, and it is unambiguous: the simplest alkene is ethene, C₂H₄, with four C–H sigma bonds plus one C–C sigma bond. Five is at least four. The trap is option I — five is not at least six.

Muse Glimmer and Nemotron both answered D; Qwen3.8-27B answered B and was scored correct; gpt-oss ran past the token cap and produced no answer at all. D is not right, but it is not arbitrary either: each carbon in ethene is sp²-hybridised and trigonal planar, and the molecule is flat. Strictly, “trigonal planar” describes a single central atom with three electron domains — BF₃ is the textbook case — and ethene has two such centres rather than being one. A model that has learned “ethene is trigonal planar” as a fact about its geometry picks D over a question that was actually asking it to count sigma bonds.

That is a distractor doing its job, not a broken key. The distinction matters, because the next example is the other kind.

Here is the other kind — a question from the other category, with three options, that every model missed in exactly the same way:

As the single-person public relations staff of a public transit agency, you are tasked with increasing the number of people who ride your buses each month. Your target audiences are lower-income individuals, college students, people with basic educations and people for whom English is a Second Language. When preparing to craft a primary message, what should you do first?

A. Consider which mediums would most effectively reach all of the target audiences. B. Consider how riding the bus could similarly affect all of the target audiences. C. Consider the public perception of bus transportation among the target audiences.

Gold is B. Every run answered C — Muse Glimmer at two reasoning strengths, Qwen3.8 at both FP8 and BF16, gpt-oss and Nemotron. Six sweeps, four models, not one dissent. A defensible case exists for C: this is a question about which step of a PR framework comes first, and the answer depends on which textbook you learned. Unanimous agreement on a non-gold answer is worth more attention than a scattered miss; it usually means the question, not the models.

The two are worth separating. The alkene question has a correct answer and a tempting wrong one; a better model gets it right, and Qwen3.8 did. The PR question has a defensible answer that is not the gold one, and no amount of capability fixes that. Only the second kind is a floor. A benchmark of this size has some of both, and from the score alone you cannot tell which you are looking at. When we report that two models are separated by 0.68 points, this is part of why we call that not significant.

One from every category

Real questions, one per category, with how many of the models that attempted it got it right:

Category Question Opts Source Correct
biology What is meant by the term muscle fatigue? 10 STEMEZ Biology 4/4
business The primary objective of team-based selling is to 10 MMLU marketing 4/4
chemistry The Pauli exclusion principle states that 8 MMLU high school chemistry 4/4
computer science Are Python variable names case-sensitive? 10 MMLU high school computer science 4/4
economics What are the major factors of production? 10 STEMEZ Economics 0/4
engineering The errors mainly caused by human mistakes are 3 MMLU electrical engineering 4/4
health Men are more likely than women to die from 10 MMLU human aging 2/4
history When did the first pharaohs emerge in Egypt? 4 MMLU prehistory 2/4
law What is personal (ratione personae) immunity? 10 MMLU international law 3/4
math A subset H of a group (G,*) is a group if 10 MMLU abstract algebra 3/4
other Which of these planets has no known moon? 9 MMLU miscellaneous 4/4
philosophy Hare claims that all moral arguments are: 4 MMLU philosophy 1/4
physics What is not true of Jupiter’s magnetic field? 10 MMLU astronomy 3/4
psychology Individuals with Moderate Mental Retardation 10 MMLU professional psychology 4/4

Four models, one run each. Per-question logging was switched on partway through this project, so Muse Glimmer and gpt-oss have no logs from their default sweeps and are represented by a reasoning-variant run — which changes how long they think, not which questions they face. This is not a difficulty ranking: a single question is not evidence about a model either way.

Two things stand out even here. “Are Python variable names case-sensitive?” is a real MMLU-Pro computer science question, sitting in the same category as graph algorithms and complexity theory. And engineering’s example has three options, not ten.

Two categories are unlike the rest

Across every model measured here, the same two categories separate them: engineering and law. Both produce the longest answers, and both are where the spread between models is widest.

They are hard for opposite reasons. Engineering is 94% worked numerical problems, so it punishes arithmetic slips and rewards patience — and it is the category most damaged by a token cap. Law is entirely original-MMLU professional exam questions, where the answers are long, similar to each other, and turn on a detail. Here is the kind of thing:

Defendant was arrested on February 1 and released one month later on March 1 after being charged with a felony. On December 1 of the same year as his arrest, he filed a motion to discharge since no trial or other action had occurred to that point. The court held a hearing 3 days after the motion was filed. Defendant should be

A. brought to trial within 10 days of the hearing on the motion to discharge. B. discharged because more than 175 days passed between his release from jail and the filing of the motion to discharge. … I. discharged because more than 200 days passed between arrest and the filing of the motion to discharge. J. brought to trial within 60 days of the filing of the motion to discharge.

Gold is A, and nothing scored it. Two runs picked a wrong option — Muse Glimmer chose B, gpt-oss chose H — and three produced no extractable answer at all, having run past the 3,072-token cap mid-reasoning: Nemotron, Qwen3.8, and Qwen3.8 again at BF16. On this question the majority failure is not a wrong answer, it is no answer. Six of the ten options are “discharged because more than N days passed”, varying only in the number and the event they count from. That is a question about a specific speedy-trial rule, and getting it right means knowing the recapture window, not reasoning about it.

How the answers are scored

Models are asked to finish with “the answer is (X)”, and the score is a regex against that. There is no partial credit and no judge model — the scoring is mechanical, which is what makes it reproducible. The details, including what that scoring gets wrong, are on the methodology page.

The part worth repeating here: a model that reasons correctly and then fails to state its answer in the expected form scores zero. Across the sweeps with per-question logs, the rate at which that happens ranges from 0.57% (Muse Glimmer at low reasoning strength) to 11.89% (Nemotron), and it is the single largest source of error in these numbers that has nothing to do with the models’ knowledge.

Attribution

MMLU-Pro is released by TIGER-Lab under the MIT licence, and the questions above are reproduced under it. The benchmark aggregates questions from the original MMLU, STEMEZ, SciBench and TheoremQA; the src field on each question names its origin, and the source column above reports it.

Dataset: TIGER-Lab/MMLU-Pro. Paper: arXiv:2406.01574.