Google's Gemini 3.8 Flash, released 2 September 2026, is quoted at 89.4% on Terminal-Bench and at 19.1% on Terminal-Bench. Both numbers are correct. They are different versions of the benchmark, and they are not interchangeable.
The two numbers
Terminal-Bench 4.0 launched in early September 2026 and is substantially harder than 2.1. Across the models we track, 4.0 scores run from 19% to 58%, while the same models sit in the high eighties on 2.1. A model quoted on 2.1 will always look dramatically stronger than one quoted on 4.0, and a table mixing the two is not a comparison at all.
This is the same trap that once had a column of SWE-bench *Verified* scores published under a SWE-bench *Pro* heading. The fix is not to pick the friendlier number; it is to record which version produced it and refuse to mix scales.
What we publish, and why
Our coding recipe names Terminal-Bench 4.0. So 19.1% is the figure that enters Gemini 3.8 Flash's coding rating, alongside its DeepSWE v1.1 result of 73.7% — a strong score — and that combination is why its published coding rating sits below models with a narrower but more flattering evidence base.
That is the honest outcome rather than a flaw. Gemini 3.8 Flash is measured on both the generous benchmark and the brutal one; several competitors publish only the generous one. Having more evidence should not be penalised, so we hold every category to a minimum number of contributing benchmark families before it rates at all.
The rest of the model
It accepts text, image, audio, video and file input, which is unusual at this price. On Artificial Analysis' Intelligence Index v4.3 it scores 41.
What to do with a benchmark number
Before comparing two models on any published score, check three things: the benchmark version, the reasoning effort, and who ran it. A vendor's launch table and an independent harness frequently disagree, and a version bump can move a score by seventy points without the model changing at all.
- Full specification sheet: Gemini 3.8 Flash
- Our rating methodology: /methodology
- Compare against a model with complete coverage: Gemini 3.8 Flash vs GPT-6 Astra




