DeepSeek released V4.1 Flash on 10 September 2026 under the MIT licence: a 552-billion-parameter sparse mixture-of-experts model that activates 8B parameters on input and 16B on output, built on the company's Causal Encoder-Decoder architecture — a 40-layer transformer arranged as a 20-layer causal encoder followed by a 20-layer decoder.
The headline is the price-to-capability ratio. It is the cheapest model in our tracked text catalog and it is not the weakest.
What it costs
Every figure above is read from DeepSeek's own served endpoint rather than the cheapest reseller, so the price and the speed describe the same configuration you would actually buy. Other providers serving this model publish different numbers.
Where it ranks
DeepSeek publishes 74.2% on DeepSWE v1.1, the end-to-end agentic software-engineering benchmark, and 90.9% on GPQA Diamond. The DeepSWE figure puts it within half a point of models costing eighty times more per output token.
That comparison needs one caveat, and it is the important one. A single benchmark is not a category. DeepSeek publishes no Terminal-Bench 4.0 result, and Terminal-Bench 4.0 is where the spread between models is widest — scores across our catalog run from 19% to 58% on it, against 57-75% on DeepSWE. A model measured only on the generous benchmark will always look better than one measured on both. We therefore hold its coding rating to the evidence that exists rather than extrapolating from one strong result.
On Artificial Analysis' Intelligence Index v4.3, a ten-evaluation composite, DeepSeek V4.1 Flash scores 40.
Running it yourself
The weights are published on Hugging Face under MIT, which permits modification, self-hosting and redistribution without restriction. The practical constraint is memory, not licensing:
That is a multi-GPU node or a Mac Studio Ultra cluster, not a single consumer card. An RTX 5090 at 32GB will not hold it at any quantization. Our hardware calculator sizes it against specific machines.
Who it is for
If you are running high-volume classification, bulk summarization or agentic coding loops where per-token cost dominates, this is the model to price against. If you need a rating backed by the full benchmark set rather than one result, see the leaderboard for models with complete coverage.
- Full specification sheet: DeepSeek V4.1 Flash
- Head to head: DeepSeek V4.1 Flash vs GPT-6 Astra




