Mistral's volume workhorse — now cheaper on input than the Small 3.1 it replaces, and still the cheapest European-hosted general model on our ledger.
Mistral Small 3.1 was officially retired on November 30, 2025; Small 4, released March 16, 2026, takes over the same slot at a lower input price — $0.15 versus $0.20 per 1M tokens — while output stays at $0.60 and the context window doubles to 256K. It's Mistral's answer for the 90% of API traffic that doesn't need a flagship, and it undercuts GPT-5.4 mini on both rates. For teams with EU data-residency requirements, it's often the only model in this price class that ticks the compliance box without a US or Chinese provider in the loop.
The other thing no rival here offers: an open-weights sibling. If your volume grows to where per-token pricing hurts, you can move the same family onto your own GPUs — the API becomes a prototyping stage rather than a permanent bill.
Small 4 is a competent generalist with vision support: classification, extraction, summarization, routine drafting, and solid multilingual coverage across European languages. Tool calling works, simple agent loops work, and the doubled 256K window means fewer documents need chunking than they did on Small 3.1.
The honest weakness: hard reasoning is still out of scope — that's Large 3 territory, or a different provider entirely.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Mistral Small 4 | $0.15 | $0.60 | 256K |
| Mistral Large 3 | $0.50 | $1.50 | 256K |
| GPT-5.4 mini | $0.75 | $4.50 | 400K |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M |
| DeepSeek V4-Flash | $0.14 | $0.28 | 1M |
On pure price only DeepSeek V4-Flash beats it, and that means Chinese-hosted infrastructure. Against GPT-5.4 mini and Gemini Flash-Lite, Small 4 trades context window for cheaper output and EU jurisdiction. If sovereignty matters, this is the budget pick; if window size matters, look at the 1M rivals.