Qwen3.8-27B at a 65,536-token cap: the one model that needed it

dgx-spark Qwen3.8-27B

This is the sixth and last model in a campaign that re-measured every quality number on this site with the answer budget raised twenty-one fold. The first five posts all reached some version of the same conclusion: the 3,072-token default was wrong, and 65,536 was more room than anyone needed.

Qwen3.8-27B is the exception. It is the one model that was still gaining at the top of the range.

The result

Subject (accuracy %) 3,072-token cap65,536-token cap Spread
engineering 64.0 80.0 16.0
other 68.0 76.5 8.5
law 63.5 70.5 7.0
chemistry 86.0 92.5 6.5
physics 87.0 91.5 4.5
history 78.5 75.0 3.5
business 87.5 90.5 3.0
math 91.5 94.5 3.0
computer science 81.5 84.0 2.5
health 73.5 76.0 2.5
biology 91.5 93.0 1.5
philosophy 72.5 73.5 1.0
psychology 80.5 81.0 0.5
economics 91.0 91.0 0.0
Mean 79.8 83.5 3.8
MMLU-Pro accuracy by subject, 200 questions per cell, Qwen3.8-27B at two generation caps. Spread is the gap between the two columns; on the Mean row it is the gap between the column means.

79.75 to 83.54, +3.79. Twelve subjects up, one flat, one down.

Engineering leads at +16.0, and the ordering below it tracks one thing:

Subjects Mean accuracy Mean gain
Answers grew more than 1.3× 11 78.77 → 83.86 +5.09
Answers grew less than 1.3× 3 83.33 → 82.33 −1.00

The three that did not grow are economics, history and psychology. Between them they moved −1.00 on average, which is noise around zero and is what a cap that never bound anything looks like. It is the same split Muse Glimmer showed, and the sixth model in a row to show it.

Subject (tokens/answer) 3,072-token cap65,536-token cap Spread
engineering 1896 7253 5357
law 1750 3960 2209
physics 934 2096 1162
chemistry 1151 2261 1110
math 824 1854 1030
computer science 886 1772 885
health 913 1682 769
philosophy 1096 1777 681
business 864 1505 641
other 827 1397 570
biology 678 897 218
economics 629 752 123
history 819 940 120
psychology 690 793 103
Mean generated tokens per answer at each cap, sorted by change. The subjects at the top are the ones that gained.

Where this model breaks the pattern

Every other cap-sensitive model in this campaign had stopped improving well before 65,536. The Qwen3.6 pair were finished by 8,192: their engineering scores went 52.0 → 83.0 → 84.5 → 82.5 and 60.0 → 81.5 → 82.5 → 79.0, flat after the first step. The campaign’s own driver notes that 65,536 was chosen because it is where Nemotron stopped being truncated, and for the others it was insurance rather than points.

Not here:

Subject (accuracy %) 3,0728,19216,38465,536
engineering 64.0 67.5 75.0 80.0
law 63.5 70.5 71.0 70.5
MMLU-Pro engineering and law, 200 questions per cell, Qwen3.8-27B against itself at four generation caps. The 8,192 and 16,384 arms measured these subjects only.

Engineering runs 64.0 → 67.5 → 75.0 → 80.0. Every step is worth something, including the last one, and the curve has not turned over. The token counts say the same: 1,896 → 3,441 → 5,052 → 7,253 per answer, still climbing at the top.

Subject (tokens/answer) 3,0728,19216,38465,536
engineering 1896 3441 5052 7253
law 1750 2909 3438 3960
Mean generated tokens per engineering and law answer at each cap. Engineering has not flattened; law has.

Law is the counterweight in the same table: 63.5 → 70.5 → 71.0 → 70.5, with answers growing 1,750 → 2,909 → 3,438 → 3,960. All of law’s gain arrives by 8,192 and the rest of the budget buys nothing. So the right cap is not even constant within one model — it is a property of how long that model reasons on that subject, and engineering is the only place in this grid where 65,536 earned its keep.

This is the measurement that justifies the campaign’s cap after the fact. Five models suggested 65,536 was over-provisioned. The sixth would have been under-measured at anything smaller, and there was no way to know which kind a model was without running it.

Law, and a debt from earlier in the campaign

Law fell for both Qwen3.6 models when the cap came off — −5.0 and −2.0 — which was alarming enough to be re-run and shown to be sampling churn: about one answer in ten changes between identical runs, and the flips split evenly.

That result was a negative one. It said the drops were not real; it could not say whether the cap does anything to law at all. Here it does: +7.0, with answers 2.26× longer, and a curve that rises and then flattens exactly as a cap-limited subject should. Law is cap-sensitive on this model, and the earlier drops really were noise rather than something specific about the subject.

Does it change the verdict against Qwen3.6-27B?

The original post on this model concluded that the successor is 2.79 points behind Qwen3.6-27B, losing on knowledge and judgement while gaining slightly on technical subjects. Both models have now been re-measured at the honest cap:

Subject (accuracy %) Qwen3.8-27BQwen3.6-27B Spread
philosophy 73.5 87.0 13.5
computer science 84.0 89.5 5.5
psychology 81.0 86.5 5.5
history 75.0 80.0 5.0
physics 91.5 95.5 4.0
law 70.5 73.5 3.0
chemistry 92.5 90.0 2.5
biology 93.0 91.0 2.0
health 76.0 78.0 2.0
business 90.5 92.0 1.5
math 94.5 93.0 1.5
other 76.5 78.0 1.5
engineering 80.0 79.0 1.0
economics 91.0 90.5 0.5
Mean 83.5 86.0 2.4
MMLU-Pro accuracy by subject, 200 questions per cell, both models at a 65,536-token generation cap and both FP8 with MTP speculative decoding.

83.54 against 85.96 — the gap closes from 2.79 to 2.42 and does not go away. The shape is unchanged too:

Qwen3.8 − Qwen3.6, at 65,536
Technical subjects −0.42
Knowledge and judgement −3.94

Philosophy is still −13.5, history −5.0, law −3.0. The earlier post’s finding — that this model trades knowledge for technical ability, and gave up ground precisely where models are most separable — survives a measurement in which both models gained three to four points.

What it cost

Arm What ran Questions Tokens generated Machine time
3,072-token cap 14 subjects · 3,072-token cap 2,800 2,791,722 8.1 h
65,536-token cap 14 subjects · 65,536-token cap 2,800 5,787,596 22.3 h
Total 5,600 8,579,318 30.5 h
Machine time for both arms, summed from data/subject-sweep/. Scoring time only: it excludes model loading and the rest periods between subjects.

8.1 hours to 22.3 hours for the same 2,800 questions, at 997 → 2,067 generated tokens per answer. Engineering alone took 6.07 hours of that — one subject consuming more than a quarter of the sweep, and the second most expensive category in the entire campaign.

Throughput fell from 95.3 to 72.0 tok/s on the same machine with the same drafter, which is the expected shape: longer answers mean a larger KV cache per slot at eight-way concurrency.

What this does not tell you

200 questions per subject, and this campaign has measured 8–10% of answers changing between identical runs. Treat single-subject moves under about 5 points as scatter. Engineering’s +16.0, other’s +8.5, law’s +7.0 and chemistry’s +6.5 clear that. History’s −3.5 does not, and its answers grew 15%, so there is no mechanism for the cap to have caused it.

The 3,072 baseline’s serving command was never recorded. That sweep predates the fix that makes the harness record what it actually ran, so the file says serving_command: null. The two arms are comparable on the evidence that survives — the baseline logged speculative-decoding acceptance of 0.77–0.85 and this arm 0.73–0.84, so both had a working drafter of the same kind — but it is weaker than being able to diff two commands, and it is the reason the 65,536 arm’s command was reconstructed rather than copied.

Engineering has not been shown to saturate. This post says the curve was still climbing at 65,536, not that it stops there. A larger cap might buy more. It would also cost more than six hours for one subject, and this campaign chose to stop.

Nothing here is about speed, and half the argument for this model was throughput — it answered the original 2,800 questions in half the wall clock of its predecessor. That comparison is in the original post and this re-measurement does not touch it.

Setup