← Back to API models

Ministral 3 14B

by Mistral AI · budget vision · replaced Pixtral 12B

A 14B-parameter multimodal model at a flat $0.20 per 1M tokens both directions — pricier than Pixtral 12B, but with double the context and performance Mistral compares to its own larger Small 3.2 24B.

Input $0.20 / 1M
Output $0.20 / 1M
Context 256K tokens
Specialty Vision
La Plateforme ↗ Updated July 27, 2026
§ API pricing

Per-token rates.

Input
$0.20/1M tokens
Prompt tokens
  • Up from Pixtral 12B's $0.15
  • Images billed as tokens
  • Text input at the same rate
Output
$0.20/1M tokens
Completion tokens
  • Same rate both directions
  • No output-premium math
  • Simple, symmetric pricing
Context
256Ktokens
Window
  • 2× Pixtral 12B's 128K window
  • ~190K words of text alongside images
  • Dozens of images per call
Size
14Bparameters
Open weights
  • Comparable to Mistral's Small 3.2 24B, per Mistral
  • Self-hostable, single-GPU friendly
  • Apache-licensed release

Why Ministral 3 14B exists

Pixtral 12B was officially retired on December 2, 2025. Ministral 3 14B takes its place as Mistral's budget vision option, moving to a different model family in the process — Ministral rather than Pixtral — while keeping the same "cheap way to look at images" mission. Pricing moves to a flat $0.20 per 1M tokens, up from Pixtral's $0.15, but the context window doubles to 256K and Mistral's own model card places its performance in the same range as the much larger Small 3.2 24B.

It's a 14-billion-parameter open-weights model, which sets expectations correctly: this is a tool for volume vision tasks at a modest step up in capability from its predecessor, not a frontier brain that happens to see. The open release also means you can self-host it on a single GPU when API economics stop making sense.

Capabilities

Solid image understanding — captioning, OCR-ish reading, chart and screenshot description, content tagging — plus ordinary text chat in the same call, with headroom over Pixtral 12B thanks to the larger parameter count and Mistral's claimed Small 3.2-class performance. The 256K window fits more images and more surrounding text per request than Pixtral's 128K did, useful for batch processing.

The honest weakness: detail and reasoning. Fine-grained chart analysis, dense document layouts, and visual reasoning chains still belong to Opus 5, Gemini 3.1 Pro, or GPT-5.5 — at many times the price.

Typical use cases

  • Bulk image captioning and alt-text generation
  • Content moderation on image streams
  • Product-photo tagging for e-commerce catalogs
  • Screenshot triage and routing
  • Self-hosted vision pipelines on a single GPU

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
Ministral 3 14B$0.20$0.20256K
Gemini 3.1 Flash-Lite$0.25$1.501M
GPT-5.4 mini$0.75$4.50400K
Claude Haiku 4.5$1$5200K
Mistral Small 4$0.15$0.60256K

Every multimodal rival still charges more on output — Gemini Flash-Lite 7.5×, GPT-5.4 mini 22.5×. For pure "describe/tag/filter this image" volume, Ministral 3 14B remains near the price floor even after the increase from Pixtral 12B. The moment the task becomes "reason about this image", spend up.

← See the full Mistral lineup