← Back to API models

Mistral Small 4

by Mistral AI · budget tier · released March 16, 2026

Mistral's volume workhorse — now cheaper on input than the Small 3.1 it replaces, and still the cheapest European-hosted general model on our ledger.

Input $0.15 / 1M
Output $0.60 / 1M
Context 256K tokens
Open weights Yes
La Plateforme ↗ Updated July 27, 2026
§ API pricing

Per-token rates.

Input
$0.15/1M tokens
Prompt tokens
  • 25% cheaper input than Small 3.1's $0.20
  • A fifth of GPT-5.4 mini's $0.75
  • Vision inputs supported
Output
$0.60/1M tokens
Completion tokens
  • Unchanged from Small 3.1
  • An eighth of GPT-5.4 mini's $4.50
  • Cheapest EU-hosted output here
Context
256Ktokens
Window
  • 2× Small 3.1's 128K window
  • ~190K words of practical input
  • Same window as Mistral Large 3
Open weights
Apachelicensed sibling
Self-hosting
  • Downloadable open-weights release
  • Run on your own GPUs
  • Same family, no per-token fees

Why Small 4 exists

Mistral Small 3.1 was officially retired on November 30, 2025; Small 4, released March 16, 2026, takes over the same slot at a lower input price — $0.15 versus $0.20 per 1M tokens — while output stays at $0.60 and the context window doubles to 256K. It's Mistral's answer for the 90% of API traffic that doesn't need a flagship, and it undercuts GPT-5.4 mini on both rates. For teams with EU data-residency requirements, it's often the only model in this price class that ticks the compliance box without a US or Chinese provider in the loop.

The other thing no rival here offers: an open-weights sibling. If your volume grows to where per-token pricing hurts, you can move the same family onto your own GPUs — the API becomes a prototyping stage rather than a permanent bill.

Capabilities

Small 4 is a competent generalist with vision support: classification, extraction, summarization, routine drafting, and solid multilingual coverage across European languages. Tool calling works, simple agent loops work, and the doubled 256K window means fewer documents need chunking than they did on Small 3.1.

The honest weakness: hard reasoning is still out of scope — that's Large 3 territory, or a different provider entirely.

Typical use cases

  • EU data-residency-compliant volume pipelines
  • Classification, extraction, and summarization at scale
  • Multilingual European-language products
  • Prototype on API, graduate to self-hosted open weights
  • Cost-floor chat features

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
Mistral Small 4$0.15$0.60256K
Mistral Large 3$0.50$1.50256K
GPT-5.4 mini$0.75$4.50400K
Gemini 3.1 Flash-Lite$0.25$1.501M
DeepSeek V4-Flash$0.14$0.281M

On pure price only DeepSeek V4-Flash beats it, and that means Chinese-hosted infrastructure. Against GPT-5.4 mini and Gemini Flash-Lite, Small 4 trades context window for cheaper output and EU jurisdiction. If sovereignty matters, this is the budget pick; if window size matters, look at the 1M rivals.

← See the full Mistral lineup