Frontier LLM API providers

Anthropic, OpenAI and Google lead on capability; DeepSeek, Moonshot, Z.ai and Qwen undercut them by roughly an order of magnitude on price. Routers like OpenRouter, Groq and Cerebras trade first-party features for reach or speed. Pick on continuity and tool-use fidelity, not headline token price — price moves monthly, migrations do not.

Last verified
2026-06-24 (29d ago)
Source confidence
27%
Re-verified
every 7 days
Tools
18
Fields
27
18 tools · verified 29d ago
Google GeminiGemini via AI Studio for prototyping or Vertex AI for production79.9

Capability

Flagship model
Gemini 3 Pro
Context window (tokens)
1,000,000 tokens
Max output (tokens)
Image input
Yes
Audio in/out
Yes
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$2 /M tok
$/M output (/M tok)
$12 /M tok
$/M cache read (/M tok)
$0.2 /M tok
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Partial
MCP support
Partial
Cache TTL
Explicit caches, default 1 h TTL; implicit caching too

Governance & continuity

Continuity policy
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
Partial
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Long context and native multimodality across Google's cloud
OpenAIGPT models plus audio, images and embeddings on one bill79.8

Capability

Flagship model
GPT-5.1
Context window (tokens)
272,000 tokens
Max output (tokens)
128,000 tokens
Image input
Yes
Audio in/out
Yes
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$1.25 /M tok
$/M output (/M tok)
$10 /M tok
$/M cache read (/M tok)
$0.125 /M tok
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Partial
MCP support
Yes
Cache TTL
Automatic, ~5–60 min, no configuration

Governance & continuity

Continuity policy
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's — closer to a quarter than half a year for API snapshots — and preview models have been pulled faster still.
Notice period (days)
90 days
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Unknown

Traction

API since
2020

Positioning

Positioning
One API for text, reasoning, audio, images and embeddings
Amazon BedrockMulti-vendor model access inside your existing AWS account78.1

Capability

Flagship model
Multi-vendor — Claude Opus 4.8, Llama 4, Mistral, Nova Premier
Context window (tokens)
1,000,000 tokens
Max output (tokens)
128,000 tokens
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
No

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Yes
MCP support
No
Cache TTL
5 min default, 1 h option; explicit breakpoints only

Governance & continuity

Continuity policy
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.
Notice period (days)
Pinnable versions
Yes
Zero retention
Yes
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Frontier models inside your existing AWS security perimeter
Mistral AIEuropean lab with an open-weight lineage and EU-resident hosting72.8

Capability

Flagship model
Mistral Large 3
Context window (tokens)
Max output (tokens)
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Partial
Computer use
No
MCP support
Partial
Cache TTL

Governance & continuity

Continuity policy
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
European models with open weights and deployable anywhere
QwenAlibaba's model family — huge open-weight range, closed flagship67.4

Capability

Flagship model
Qwen3-Max
Context window (tokens)
262,144 tokens
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$1.2 /M tok
$/M output (/M tok)
$6 /M tok
$/M cache read (/M tok)
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
Partial
Cache TTL

Governance & continuity

Continuity policy
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
Partial
AU region
Partial

Traction

API since
2023

Positioning

Positioning
The widest open-weight family, plus a closed flagship tier
Meta LlamaOpen-weight Llama models, hosted almost everywhere but Meta66.6

Capability

Flagship model
Llama 4 Maverick
Context window (tokens)
1,000,000 tokens
Max output (tokens)
Image input
Yes
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL
Host-dependent

Governance & continuity

Continuity policy
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you — if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Open-weight models you can host anywhere, forever
Self-hosted (vLLM)Run open weights on your own GPUs behind an OpenAI-shaped API65.1

Capability

Flagship model
Whatever you deploy — DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)
0 %

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL
Automatic prefix cache in VRAM, no TTL

Governance & continuity

Continuity policy
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.
Notice period (days)
Pinnable versions
Yes
Zero retention
Yes
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Your weights, your GPUs, your OpenAI-compatible endpoint
Together AIServerless and dedicated hosting for open-weight models64.2

Capability

Flagship model
Open-weight catalogue — DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)
50 %

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
Open-weight models with a Western contract and dedicated capacity
xAIGrok models with an OpenAI-shaped API and live X data access63.6

Capability

Flagship model
Grok 4.1
Context window (tokens)
Max output (tokens)
Image input
Yes
Audio in/out
No
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
Partial
AU region
No

Traction

API since
2024

Positioning

Positioning
Fast, cheap frontier models with live access to X
Fireworks AIFast open-weight inference with strong structured-output support63.2

Capability

Flagship model
Open-weight catalogue — DeepSeek, Kimi K2, Qwen3, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
Low-latency open-weight inference with strict structured output
CohereEnterprise-focused models built for RAG and private deployment62.9

Capability

Flagship model
Command A (command-a-03-2025)
Context window (tokens)
256,000 tokens
Max output (tokens)
Image input
Partial
Audio in/out
No
Open weights
Partial
OpenAI-compat API
Partial

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$2.5 /M tok
$/M output (/M tok)
$10 /M tok
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Partial

Traction

API since
2021

Positioning

Positioning
Enterprise RAG models you can deploy inside your own network
AnthropicClaude models, built around long agentic runs and tool use62.7

Capability

Flagship model
Claude Fable 5 (claude-fable-5)
Context window (tokens)
1,000,000 tokens
Max output (tokens)
128,000 tokens
Image input
Yes
Audio in/out
No
Open weights
No
OpenAI-compat API
Partial

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$10 /M tok
$/M output (/M tok)
$50 /M tok
$/M cache read (/M tok)
$1 /M tok
Batch discount (%)
50 %

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Yes
MCP support
Yes
Cache TTL
5 min default, 1 h option; automatic or explicit breakpoints

Governance & continuity

Continuity policy
Published deprecation page lists a retirement date per model, typically months ahead — Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.
Notice period (days)
180 days
Pinnable versions
Partial
Zero retention
Partial
Trains on your data
No
AU region
Partial

Traction

API since
2023

Positioning

Positioning
Frontier models for long-horizon agentic work and code
OpenRouterOne OpenAI-shaped key in front of hundreds of models60.9

Capability

Flagship model
Whichever upstream model you route to (400+ available)
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)
0 %

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Passed through where the upstream supports caching

Governance & continuity

Continuity policy
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself — a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.
Notice period (days)
Pinnable versions
Partial
Zero retention
Partial
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
One key, one API shape, every model worth calling
Z.ai (GLM)GLM models with an unusually cheap flat-rate coding plan59.9

Capability

Flagship model
GLM-4.6
Context window (tokens)
200,000 tokens
Max output (tokens)
Image input
Partial
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$0.6 /M tok
$/M output (/M tok)
$2.2 /M tok
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.
Notice period (days)
Pinnable versions
Yes
Zero retention
No
Trains on your data
Partial
AU region
No

Traction

API since
2023

Positioning

Positioning
Coding-focused GLM models with a flat-rate subscription
Moonshot AIKimi models — open-weight agentic performance at low cost58.8

Capability

Flagship model
Kimi K2 Thinking
Context window (tokens)
256,000 tokens
Max output (tokens)
Image input
No
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$0.6 /M tok
$/M output (/M tok)
$2.5 /M tok
$/M cache read (/M tok)
$0.15 /M tok
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Explicit context caching, paid per storage hour

Governance & continuity

Continuity policy
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence — download the checkpoint and continuity is your problem, not theirs.
Notice period (days)
Pinnable versions
Yes
Zero retention
No
Trains on your data
Partial
AU region
No

Traction

API since
2023

Positioning

Positioning
Open-weight agentic performance at open-weight prices
GroqCustom LPU silicon serving open-weight models at extreme speed56.6

Capability

Flagship model
Curated open-weight models (Kimi K2, Llama 4, Qwen3, GPT-OSS)
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
250 ms
Output tok/s (tok/s)
400 tok/s

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.
Notice period (days)
Pinnable versions
Partial
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2024

Positioning

Positioning
Custom LPU silicon for open-weight models at very low latency
CerebrasWafer-scale inference — the fastest tokens per second available56.5

Capability

Flagship model
Curated open-weight models (Qwen3, GLM, Llama, GPT-OSS)
Context window (tokens)
Max output (tokens)
Image input
No
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
200 ms
Output tok/s (tok/s)
2,000 tok/s

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
Batch discount (%)

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL

Governance & continuity

Continuity policy
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.
Notice period (days)
Pinnable versions
Partial
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2024

Positioning

Positioning
Wafer-scale inference — the highest tokens per second available
DeepSeekFrontier-adjacent models at a small fraction of Western prices46.2

Capability

Flagship model
DeepSeek-V3.2 (deepseek-chat / deepseek-reasoner)
Context window (tokens)
128,000 tokens
Max output (tokens)
Image input
No
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$0.28 /M tok
$/M output (/M tok)
$0.42 /M tok
$/M cache read (/M tok)
$0.028 /M tok
Batch discount (%)
0 %

Agentic

Tool use
Partial
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Automatic disk cache, no configuration

Governance & continuity

Continuity policy
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases — deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.
Notice period (days)
Pinnable versions
No
Zero retention
No
Trains on your data
Yes
AU region
No

Traction

API since
2023

Positioning

Positioning
Near-frontier quality at a fraction of the token price
yespartialnounknown◆ measured · ▸ vendor-claimed · ▪ community · · inferredExpand a row for the source and date behind every cell.
Frontier LLM APIs compared — price, context, tool use and deprecation policy pricing, compared line by line

Which frontier LLM API should I build on?

Every provider on this page will quote you a $/M token figure. Almost none of them will tell you how long the model behind it lives. That asymmetry is the single biggest hidden cost in this category, and it is why this page leads with a continuity column rather than a price column.

For most production work in 2026 the shortlist is short. Anthropic if the workload is agentic — tool calls, long multi-step runs, code — because effort control, prompt caching and MCP are first-class rather than bolted on. OpenAI if you need audio, image generation and text behind one billing relationship, or if your team is already fluent in the Responses API. Google if you need million-token context cheaply, native audio in and out, or an AU-resident deployment through Vertex. Those three are not interchangeable at the prompt level: a prompt tuned on one lands differently on the others, and the effort/thinking knobs have no common vocabulary.

The Chinese open-weight labs — DeepSeek, Moonshot, Qwen, Z.ai — have changed what the floor looks like. DeepSeek's flagship bills $0.42 per million output tokens against $50 for Claude Fable 5 and $15 for Claude Sonnet 5 — roughly a hundred and twenty times cheaper than Anthropic's top-end model, and still around thirty-five times cheaper than the volume tier most teams would actually reach for. For classification, extraction, summarisation and bulk rewriting the quality gap does not justify the price gap. The catch is governance, not capability: retention defaults are permissive, zero-data-retention is generally not on offer, and several expose only a rolling alias that silently upgrades under you. If the data is sensitive, run the weights yourself or route through a host that is contractually clean rather than calling the origin API.

Routers deserve more credit than they get. OpenRouter is the cheapest insurance policy in the category — one integration, hundreds of models, and a real chance that a retired model stays reachable somewhere because an open-weight copy is still hosted. Groq and Cerebras are not general-purpose substitutes; they are latency instruments. If your product's differentiator is that the answer appears instantly — voice, autocomplete, interactive search — they are worth an entire architecture. If it is not, their model menu will frustrate you within a quarter.

Where this is heading: prices keep falling, context windows have stopped being the differentiator, and the fight has moved to agentic fidelity — parallel tool calls that do not degrade, structured output that never breaks schema, caching that survives a long agent loop. Assume the model you ship on today will be retired inside two years. Build the abstraction layer now, keep an evaluation set you can re-run in an afternoon, and treat every provider on this page as replaceable.

The pricing traps nobody puts on the pricing page

Headline $/M is the least useful number in this category. Four things distort it.

First, output tokens dominate. A reasoning model can burn tens of thousands of thinking tokens before it writes a word you see, and those bill at the output rate. A provider with a cheap input rate and an expensive output rate is expensive for reasoning work and cheap for retrieval-augmented answering. Model your actual input:output ratio before comparing.

Second, cached input is where the real money is. Cache reads bill at roughly a tenth of the input rate at Anthropic and OpenAI, and an agent loop resending its history every turn is mostly cache reads. But caching is a prefix match — a timestamp in the system prompt, a non-deterministic JSON serialisation, or a tool list that varies per user destroys it silently. Watch the cache-read token counter, not the invoice.

Third, long-context surcharges. Google charges a higher rate above 200K tokens on its flagship; others quietly meter cache storage per hour. A million-token window advertised at the short-context rate is not what you will pay.

Fourth, batch — but check the column before you budget for it. Where an asynchronous endpoint exists the discount is usually a flat 50%, and Anthropic, OpenAI, Google, Mistral and Qwen all publish that rate, as do Bedrock and Together. It is not universal. DeepSeek offers no batch endpoint at all — its list price already sits below most competitors' batch rates — and neither does OpenRouter, which is a real cost if half your workload could tolerate async. xAI, Moonshot, Z.ai, Cohere, Fireworks and Cerebras publish nothing we could confirm, and Groq's batch discount is documented as existing without a percentage we could stand behind. Where the halving is there and part of your workload tolerates a few hours of latency, it is the one saving no negotiation will beat; where it is not, no amount of committed spend conjures it.

How to make a migration survivable

Assume an eighteen-month life for whatever you ship on. The providers that publish retirement dates are doing you a favour; the ones that don't are the ones that will break you.

Pin a dated snapshot wherever the provider offers one, and record the pin in configuration rather than code. Note that this is getting harder, not easier: Anthropic's current line-up ships as bare aliases with no dated model ID behind them, so 'pin the snapshot' is not always an available move any more, and you fall back on the published retirement date instead.

Keep the provider behind one seam. Not an abstraction that pretends every model is the same — that fails the moment you touch thinking blocks or tool-result shapes — but a single module that owns request construction, retry policy and response parsing per provider. Two implementations behind one interface is the cheap insurance; a universal adapter is the expensive fantasy.

Maintain an evaluation set you can re-run in an afternoon: fifty to two hundred real inputs with graded expected outputs. Without it, a forced migration becomes a multi-week vibe check. With it, a migration is a morning's work and a number you can defend.

Finally, know your escape hatch before you need it. For open-weight models it is total — the weights are on disk and vLLM will serve them for as long as you have GPUs. For closed models it is a router that may still have an upstream, or nothing at all.

Deciding: three workloads, three answers

Bulk text processing — classification, extraction, tagging, translation, summarising a firehose. Price dominates and capability differences barely register. DeepSeek, Qwen, Moonshot and Z.ai are the rational choices, or Groq and Together if you want open weights with a Western contract. Add the batch endpoint if latency is negotiable. Do not put sensitive data through a provider whose retention terms you haven't read.

Agent harnesses — a loop that plans, calls tools, reads results and iterates for minutes at a time. This is where the frontier labs earn their price. You need tool calls that stay well-formed at depth, structured output that is enforced rather than requested, an effort or thinking control so you can trade cost against quality per route, and prompt caching that survives the loop. Anthropic and OpenAI are the credible options; Google is close and cheaper on long context. Budget for the fact that a long agentic turn can now run for minutes on a single request.

Interactive, latency-critical UI — voice, live search, inline completion. Time to first token is the product. Cerebras and Groq are in a different class here, an order of magnitude ahead of first-party endpoints on tokens per second, and the constraint is that you take whichever open-weight models they have chosen to host. Design the feature around the model menu, not the other way around.

What is frontier llm apis compared — price, context, tool use and deprecation policy?

Frontier LLM APIs compared — price, context, tool use and deprecation policy

On toolweight, Frontier LLM APIs compared — price, context, tool use and deprecation policy means the 18 tools benchmarked on this page — Anthropic, OpenAI, Google Gemini, xAI, Meta Llama, Mistral AI, DeepSeek, Qwen, Moonshot AI, Z.ai (GLM), Amazon Bedrock, Cohere, OpenRouter, Together AI, Fireworks AI, Groq, Cerebras, Self-hosted (vLLM) — judged on the same 27 fields, from the same sources, on the same date. The question it exists to answer: Which frontier LLM API should I build on?

How does toolweight compare these?

Pricing on this page is stale within a fortnight. That is not hedging — this category re-prices faster than any other on toolweight, so the page carries a seven-day verification cadence and every price cell is dated. Treat any cell whose date is more than a month old as indicative and click through to the vendor's pricing page before you build a forecast on it. Figures are pay-as-you-go list rates in USD per million tokens, before batch, cache or committed-spend discounts, and before any provider-specific long-context surcharge. One convention worth stating: where a vendor is running an unexpired introductory rate, the column carries the list price and the cell note carries the introductory rate and its expiry date. An introductory rate is not a discount you negotiate — it is what you actually pay today — so read the note before modelling a bill.

Entries here are providers, not models. Because a provider ships a dozen models at a dozen prices, every price, context and capability figure on a row describes that provider's current flagship model, named explicitly in the Flagship model column so no number is orphaned from the thing it measures. Compare rows knowing that a cheaper flagship often sits beside a cheaper mid-tier that would serve your workload better. The killer column, deprecation and continuity policy, records two things: how much notice you get before a model is retired, and whether you can pin a dated snapshot that keeps serving after the alias moves on. Continuity is judged from published deprecation pages and observed retirements, not from marketing claims, and where a provider publishes no policy at all we say so rather than guessing a number. Latency figures are community-reported from third-party benchmarks; toolweight has not run its own TTFT harness yet, so most of that column is honestly blank.

Full methodology and sourcing policy →

Frequently asked questions

What actually happens when a model I depend on is deprecated?

The alias stops resolving and requests 404. If you pinned a dated snapshot you keep serving until that snapshot's published retirement date; if you called a bare alias you are migrating today. Anthropic and OpenAI publish retirement dates months ahead on dedicated deprecation pages. Several providers publish nothing, and rolling aliases can change behaviour underneath you with no version bump at all.

Is DeepSeek or Kimi really good enough to replace Claude or GPT?

For extraction, classification, summarisation, translation and bulk rewriting, yes — and at roughly a thirty-fifth of the output-token cost of Anthropic's volume model, or about a hundred and twentieth of its top-end one. Which multiple you should care about depends on which model you would otherwise have used; the honest comparison for bulk work is against the volume tier, not the flagship. For long agentic runs with many tool calls, sustained instruction following, and code that has to compile first time, the frontier labs are still measurably ahead. Split the workload rather than picking one provider for everything.

Does an OpenAI-compatible endpoint actually make switching easy?

It makes the transport identical and the semantics different. Chat completions port cleanly. Tool-call formats, reasoning-effort parameters, cached-token accounting, thinking blocks and structured-output enforcement do not, and prompts tuned against one model's instruction-following behaviour regress on another. Budget a re-evaluation, not a config change.

How much does prompt caching really save on an agent loop?

More than any other lever. Cache reads typically bill at 10–25% of the input rate, and an agent loop resends the entire conversation on every turn, so a long run can be 80–90% cache reads. The trap is that caching is a prefix match: a timestamp or a reordered tool list at the front of the prompt silently invalidates everything after it.

Which providers can serve inference from an Australian region?

Through the hyperscalers, reliably: Amazon Bedrock in ap-southeast-2 and Google Vertex AI in australia-southeast1 both serve frontier models in-country. First-party endpoints from Anthropic, OpenAI and xAI route to US or EU infrastructure by default, with data-residency controls that vary by plan. Self-hosting is the only answer that is unambiguously AU-resident.

Should I go through a router like OpenRouter instead of direct?

Use a router when model choice is a runtime decision, when you want a fallback that survives an upstream outage, or when you are still benchmarking. Go direct when you depend on provider-specific features — prompt caching breakpoints, computer use, batch endpoints, enterprise retention terms — because routers expose the intersection of what upstreams support, not the union.

Why is the time-to-first-token column mostly empty?

Because almost nobody publishes it, and self-reported latency is meaningless without a specified prompt length, region and concurrency. The figures shown are community measurements from third-party benchmarks for providers whose entire pitch is speed. toolweight will fill this column when it runs its own harness with published methodology, not before.