Mistral Large 3 was priced at $0.50 per million input tokens and $1.50 per million output tokens when we read the API pricing page on 14 August 2026. That puts it below every US frontier and mid-tier model we track, above OpenAI’s small model, and above the cheapest Chinese grids. Mistral publishes one card per model rather than one grid, so the comparison takes longer to assemble than it should. The comparator below does it for you, against whatever model your workload sits on today.
What Mistral would cost you, against the model you run today
Pick the model your workload sits on, enter your monthly volume, and see the same month billed on Mistral. Every rate was read from the provider's own pricing page on 14 August 2026, per million tokens, standard tier.
Mistral Large 3 carried no cache discount on the day we read the page, so raising the cache share works against it and in favour of the models that do discount repeated input. Long context, batch and priority tiers are billed on separate grids and are not covered here.
What Mistral charges per million tokens
Two model cards matter for a general-purpose text workload, and both were on the page on 14 August 2026. One is Mistral’s own flagship, the other is a partner model served from the same API.
The two cards we read on 14 August 2026
Mistral Large 3 sits at $0.50 in and $1.50 out. GLM 5.2, served through the same endpoint, sits at $1.40 in and $4.40 out, with cached input at $0.14. That cache line is the one difference that changes the arithmetic later on, because Mistral Large 3 had no equivalent on the day we looked.
| What we read on the API pricing page | Rate per million tokens |
|---|---|
| Mistral Large 3, input | $0.50 |
| Mistral Large 3, output | $1.50 |
| Mistral Large 3, cached input | Not offered |
| GLM 5.2, input | $1.40 |
| GLM 5.2, output | $4.40 |
| GLM 5.2, cached input | $0.14 |

What the pricing page leaves out
The card format hides the shape of the offer. You read a rate in three seconds and still do not know how the model is billed once your workload stops being a plain chat loop. Here is what the two cards do not tell you, and what we therefore did not verify:
- Whether a long-context request is billed on the same rate as a short one.
- Whether a batch or off-peak tier exists, and at what discount.
- What the non-text products cost, since we captured only the two general-purpose text cards.
- What an enterprise or self-hosted contract looks like, which is negotiated rather than published.
Treat the two rates above as the published standard tier for text, nothing wider than that. Every other line on Mistral’s own API pricing page is outside what we read.
Where Mistral lands against the US models
The interesting comparison is not against the frontier, which Mistral does not try to price against. It is against the mid tier, where the same class of task gets done for very different money.

The mid tier is where the gap is widest
Take a workload of 20 million input tokens and 5 million output tokens a month, with no cache, which is a busy production assistant. On Mistral Large 3 that bill is $17.50. On gpt-5.6-terra it is $100.00, and on Claude Sonnet 5 it is $90.00.
The reason is output. Mistral Large 3 charges $1.50 per million output tokens against $12.00 for gpt-5.6-terra, an eightfold difference on the side of the bill that actually moves.
| Same month, 20M input and 5M output, no cache | Monthly bill |
|---|---|
| DeepSeek V4-Flash | $4.20 |
| gpt-5.6-luna | $10.00 |
| DeepSeek V4-Pro | $13.05 |
| Mistral Large 3 | $17.50 |
| GLM 5.2 on Mistral | $50.00 |
| grok-4.6 | $70.00 |
| Claude Sonnet 5 | $90.00 |
| gpt-5.6-terra | $100.00 |
Same month, same tokens: $17.50 on Mistral Large 3 against $100.00 on gpt-5.6-terra. The output rate creates that gap, not the input rate.
Three grids sit below Mistral Large 3 on that list: both DeepSeek models and OpenAI’s small model. None of them is doing the same job as a mid-tier assistant, which is the comparison this page is about, but pretending they are not cheaper would be a way of losing your trust.
The frontier tier is not a price fight
Against gpt-5.6-cyber at $75.00 per million output tokens, Mistral Large 3 is fifty times cheaper, and the number tells you nothing useful. Frontier models are bought for the hardest slice of a workload, not for volume, so Mistral competes with the tier below and the frontier is a different purchase. Our full LLM API pricing table puts all of them on one grid.
The cache discount, where Mistral gives ground
Mistral Large 3 had no cached input rate on 14 August 2026, and that is a real cost on repetitive workloads. Cache pricing applies to input you send again, typically a long system prompt or a document you keep in context, and the US grids discount it heavily.
Flip the workload to 200 million input tokens and 2 million output tokens a month, which is what retrieval looks like. Mistral Large 3 bills $103.00 on that month.
gpt-5.6-terra with 95 percent of its input served from cache bills $82.00, and Claude Sonnet 5 on the same basis bills $78.00. Without that cache, terra would cost $424.00 for the same tokens. The discount is doing all of the work, and we took Anthropic’s cache pricing apart on its own page.
So the ranking is not stable. It depends on whether your input repeats, and the comparator above lets you move that slider yourself.
The European argument, and what it actually covers
The second reason people look at Mistral has nothing to do with the rate card. It is that the vendor is European, and for some buyers that decides the question before price does.
What is on the record
Mistral AI is a French company selling from the European Union, which puts the contract, the corporate entity and the dispute resolution inside EU law rather than outside it. That is a structural fact about the vendor, not a claim about any particular server.
There is a second, more concrete point. Microsoft’s announcement of Mistral Large 3 on Foundry describes the model as shipping under an Apache 2.0 licence, which means the weights can be run on infrastructure you control instead of on anybody’s API. For a team whose blocker is data leaving its own network, that is a different answer to the same problem, and it does not appear on any pricing page.
What we did not verify
We read a pricing page, not a contract. The following were outside what we checked, and you should not take the European argument as covering them until you have read the paperwork yourself:
- Which region a given API call is actually served from.
- The sub-processor list, and whether any of it sits outside the European Union.
- Data retention and training commitments on the standard tier against the enterprise tier.
- Any certification claim, which we did not audit.
Stated plainly: buying from a European vendor changes who you have a contract with. Whether it changes where your tokens are processed is a question for the terms, and the terms are not what this page measured.
What the launch numbers say about Mistral Large 3
A rate card says what a model costs, not whether it does your job. Mistral Large 3 shipped in December 2025, and we have not benched it ourselves, so the honest substitute is what was published at launch and what a public leaderboard measured at the time.
The breakdown Rohan Paul posted on X after launch covers the architecture and the leaderboard placement rather than an impression. It reports a sparse mixture of experts with 41 billion active parameters out of 675 billion total, image and text input, base and instruct versions shipped together, and a placement on the LMArena text leaderboard at sixth among open models and twenty-eighth overall. Those are figures from the post reporting the launch numbers, not from our own bench, and a leaderboard position set in December 2025 is not one to assume still holds.
Read the leaderboard placement for what it is. Twenty-eighth overall on a public arena is a long way from the frontier models it undercuts on price, and sixth among open models is a strong result in the category that matters if you intend to self-host.
How we checked these prices
Every rate on this page was read from the provider’s own pricing page on 14 August 2026 and captured as a screenshot on the same day. We do not copy rates from aggregators, and we do not carry a number forward from a previous article without reading it again.
The Mistral figures come from the API pricing page and are the two general-purpose text cards. The OpenAI, Anthropic, xAI and DeepSeek figures used for comparison come from those vendors’ own pages, read on the same date, and are the same set we published in our full LLM API pricing comparison. Our testing protocol is written down if you want to judge whether it applies to your case, and the editorial policy behind it is public too.
When a provider changes a grid, we reread it and change the date on the page. A price with no date on it is not a price, it is a memory.
Frequently asked questions
Is Mistral cheaper than OpenAI?
Cheaper than the OpenAI models most teams run in production, yes. On 14 August 2026, gpt-5.6-terra charged $12.00 per million output tokens against $1.50 for Mistral Large 3. Cheaper than all of them, no: gpt-5.6-luna was at $0.20 in and $1.20 out on the same date, which undercuts Mistral Large 3 on both sides, and a heavily cached workload can reverse the mid-tier gap as well.
What does Mistral Large 3 cost per million tokens?
$0.50 for input and $1.50 for output, read from Mistral’s API pricing page on 14 August 2026. No cached input rate was published for that model on that date.
Why is GLM 5.2 listed on Mistral’s pricing page?
Mistral serves models other than its own through the same API, and GLM 5.2 was one of the cards on the page when we read it. It is priced separately at $1.40 in and $4.40 out, with cached input at $0.14, so it is not a substitute for the Mistral Large 3 rate.
Does the European hosting argument justify a higher price?
It does not have to, because Mistral is not asking for one at the tier we compared. On the rates above the European option is also the cheaper option against the US mid tier, which makes the question easier than it usually is.