Qwen3.8-27B at a 65,536-token cap: the one model that needed it
This is the sixth and last model in a campaign that re-measured every quality number on this site with the answer budget raised twenty-one fold. The first five posts all reached some version of the same conclusion: the 3,072-token default was wrong, and 65,536 was more room than anyone needed.
Qwen3.8-27B is the exception. It is the one model that was still gaining at the top of the range.
The result
| Subject (accuracy %) | 3,072-token cap | 65,536-token cap | Spread |
|---|---|---|---|
| engineering | 64.0 | 80.0 | 16.0 |
| other | 68.0 | 76.5 | 8.5 |
| law | 63.5 | 70.5 | 7.0 |
| chemistry | 86.0 | 92.5 | 6.5 |
| physics | 87.0 | 91.5 | 4.5 |
| history | 78.5 | 75.0 | 3.5 |
| business | 87.5 | 90.5 | 3.0 |
| math | 91.5 | 94.5 | 3.0 |
| computer science | 81.5 | 84.0 | 2.5 |
| health | 73.5 | 76.0 | 2.5 |
| biology | 91.5 | 93.0 | 1.5 |
| philosophy | 72.5 | 73.5 | 1.0 |
| psychology | 80.5 | 81.0 | 0.5 |
| economics | 91.0 | 91.0 | 0.0 |
| Mean | 79.8 | 83.5 | 3.8 |
79.75 to 83.54, +3.79. Twelve subjects up, one flat, one down.
Engineering leads at +16.0, and the ordering below it tracks one thing:
| Subjects | Mean accuracy | Mean gain | |
|---|---|---|---|
| Answers grew more than 1.3× | 11 | 78.77 → 83.86 | +5.09 |
| Answers grew less than 1.3× | 3 | 83.33 → 82.33 | −1.00 |
The three that did not grow are economics, history and psychology. Between them they moved −1.00 on average, which is noise around zero and is what a cap that never bound anything looks like. It is the same split Muse Glimmer showed, and the sixth model in a row to show it.
| Subject (tokens/answer) | 3,072-token cap | 65,536-token cap | Spread |
|---|---|---|---|
| engineering | 1896 | 7253 | 5357 |
| law | 1750 | 3960 | 2209 |
| physics | 934 | 2096 | 1162 |
| chemistry | 1151 | 2261 | 1110 |
| math | 824 | 1854 | 1030 |
| computer science | 886 | 1772 | 885 |
| health | 913 | 1682 | 769 |
| philosophy | 1096 | 1777 | 681 |
| business | 864 | 1505 | 641 |
| other | 827 | 1397 | 570 |
| biology | 678 | 897 | 218 |
| economics | 629 | 752 | 123 |
| history | 819 | 940 | 120 |
| psychology | 690 | 793 | 103 |
Where this model breaks the pattern
Every other cap-sensitive model in this campaign had stopped improving well before 65,536. The Qwen3.6 pair were finished by 8,192: their engineering scores went 52.0 → 83.0 → 84.5 → 82.5 and 60.0 → 81.5 → 82.5 → 79.0, flat after the first step. The campaign’s own driver notes that 65,536 was chosen because it is where Nemotron stopped being truncated, and for the others it was insurance rather than points.
Not here:
| Subject (accuracy %) | 3,072 | 8,192 | 16,384 | 65,536 |
|---|---|---|---|---|
| engineering | 64.0 | 67.5 | 75.0 | 80.0 |
| law | 63.5 | 70.5 | 71.0 | 70.5 |
Engineering runs 64.0 → 67.5 → 75.0 → 80.0. Every step is worth something, including the last one, and the curve has not turned over. The token counts say the same: 1,896 → 3,441 → 5,052 → 7,253 per answer, still climbing at the top.
| Subject (tokens/answer) | 3,072 | 8,192 | 16,384 | 65,536 |
|---|---|---|---|---|
| engineering | 1896 | 3441 | 5052 | 7253 |
| law | 1750 | 2909 | 3438 | 3960 |
Law is the counterweight in the same table: 63.5 → 70.5 → 71.0 → 70.5, with answers growing 1,750 → 2,909 → 3,438 → 3,960. All of law’s gain arrives by 8,192 and the rest of the budget buys nothing. So the right cap is not even constant within one model — it is a property of how long that model reasons on that subject, and engineering is the only place in this grid where 65,536 earned its keep.
This is the measurement that justifies the campaign’s cap after the fact. Five models suggested 65,536 was over-provisioned. The sixth would have been under-measured at anything smaller, and there was no way to know which kind a model was without running it.
Law, and a debt from earlier in the campaign
Law fell for both Qwen3.6 models when the cap came off — −5.0 and −2.0 — which was alarming enough to be re-run and shown to be sampling churn: about one answer in ten changes between identical runs, and the flips split evenly.
That result was a negative one. It said the drops were not real; it could not say whether the cap does anything to law at all. Here it does: +7.0, with answers 2.26× longer, and a curve that rises and then flattens exactly as a cap-limited subject should. Law is cap-sensitive on this model, and the earlier drops really were noise rather than something specific about the subject.
Does it change the verdict against Qwen3.6-27B?
The original post on this model concluded that the successor is 2.79 points behind Qwen3.6-27B, losing on knowledge and judgement while gaining slightly on technical subjects. Both models have now been re-measured at the honest cap:
| Subject (accuracy %) | Qwen3.8-27B | Qwen3.6-27B | Spread |
|---|---|---|---|
| philosophy | 73.5 | 87.0 | 13.5 |
| computer science | 84.0 | 89.5 | 5.5 |
| psychology | 81.0 | 86.5 | 5.5 |
| history | 75.0 | 80.0 | 5.0 |
| physics | 91.5 | 95.5 | 4.0 |
| law | 70.5 | 73.5 | 3.0 |
| chemistry | 92.5 | 90.0 | 2.5 |
| biology | 93.0 | 91.0 | 2.0 |
| health | 76.0 | 78.0 | 2.0 |
| business | 90.5 | 92.0 | 1.5 |
| math | 94.5 | 93.0 | 1.5 |
| other | 76.5 | 78.0 | 1.5 |
| engineering | 80.0 | 79.0 | 1.0 |
| economics | 91.0 | 90.5 | 0.5 |
| Mean | 83.5 | 86.0 | 2.4 |
83.54 against 85.96 — the gap closes from 2.79 to 2.42 and does not go away. The shape is unchanged too:
| Qwen3.8 − Qwen3.6, at 65,536 | |
|---|---|
| Technical subjects | −0.42 |
| Knowledge and judgement | −3.94 |
Philosophy is still −13.5, history −5.0, law −3.0. The earlier post’s finding — that this model trades knowledge for technical ability, and gave up ground precisely where models are most separable — survives a measurement in which both models gained three to four points.
What it cost
| Arm | What ran | Questions | Tokens generated | Machine time |
|---|---|---|---|---|
| 3,072-token cap | 14 subjects · 3,072-token cap | 2,800 | 2,791,722 | 8.1 h |
| 65,536-token cap | 14 subjects · 65,536-token cap | 2,800 | 5,787,596 | 22.3 h |
| Total | 5,600 | 8,579,318 | 30.5 h |
8.1 hours to 22.3 hours for the same 2,800 questions, at 997 → 2,067 generated tokens per answer. Engineering alone took 6.07 hours of that — one subject consuming more than a quarter of the sweep, and the second most expensive category in the entire campaign.
Throughput fell from 95.3 to 72.0 tok/s on the same machine with the same drafter, which is the expected shape: longer answers mean a larger KV cache per slot at eight-way concurrency.
What this does not tell you
200 questions per subject, and this campaign has
measured 8–10% of answers changing between identical runs.
Treat single-subject moves under about 5 points as scatter. Engineering’s +16.0,
other’s +8.5, law’s +7.0 and chemistry’s +6.5 clear that. History’s −3.5 does
not, and its answers grew 15%, so there is no mechanism for the cap to have
caused it.
The 3,072 baseline’s serving command was never recorded. That sweep predates
the fix that makes the harness record what it actually ran, so the file says
serving_command: null. The two arms are comparable on the evidence that
survives — the baseline logged speculative-decoding acceptance of 0.77–0.85 and
this arm 0.73–0.84, so both had a working drafter of the same kind — but it is
weaker than being able to diff two commands, and it is the reason the
65,536 arm’s command was reconstructed rather than copied.
Engineering has not been shown to saturate. This post says the curve was still climbing at 65,536, not that it stops there. A larger cap might buy more. It would also cost more than six hours for one subject, and this campaign chose to stop.
Nothing here is about speed, and half the argument for this model was throughput — it answered the original 2,800 questions in half the wall clock of its predecessor. That comparison is in the original post and this re-measurement does not touch it.
Setup
- NVIDIA GB10, 121 GiB unified memory, production services stopped for the duration
- vLLM
vllm/vllm-openai:v0.27.1,Qwen/Qwen3.8-27B-FP8,--max-model-len 73728,--gpu-memory-utilization 0.55 --speculative-config '{"method":"mtp","num_speculative_tokens":2}', draft acceptance 0.73–0.84 by subject- lm-evaluation-harness,
local-chat-completions, 8 concurrent requests mmlu_pro_fair: permissive answer extraction, truncating stop sequence removed- No reasoning parser. Leaving it on costs 18 points on this model
- 5-shot CoT, greedy: lm-eval sends
temperature: 0andseed: 1234on every request. 200 questions per subject, fourteen subjects