Anthropic, OpenAI and Google lead on capability; DeepSeek, Moonshot, Z.ai and Qwen undercut them by roughly an order of magnitude on price. Routers like OpenRouter, Groq and Cerebras trade first-party features for reach or speed. Pick on continuity and tool-use fidelity, not headline token price — price moves monthly, migrations do not.
Last verified
2026-06-24 (29d ago)
Source confidence
27%
Re-verified
every 7 days
Tools
18
Fields
27
18 tools · verified 29d ago
Frontier LLM API providers — 18 tools compared across 27 of 27 fields.
Tool
Weighted
w5
w3
w4
w3
w4
w5
w4
w4
w8
w9
w6
w4
w8
w7
w6
w3
w5
w7
w7
w6
w7
w4
w1
Google GeminiGemini via AI Studio for prototyping or Vertex AI for production
79.9
Gemini 3 Pro·
1,000,000 tokens▸
–
●Yes▸
●Yes▸
◐Partial▸
●Yes▸
–
–
$2 /M tok▸
$12 /M tok▸
$0.2 /M tok·
50 %▸
●Yes▸
●Yes▸
●Yes▸
◐Partial·
◐Partial·
Explicit caches, default 1 h TTL; implicit caching too·
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.▪
–
●Yes▸
◐Partial·
◐Partial·
●Yes·
2023·
Long context and native multimodality across Google's cloud·
OpenAIGPT models plus audio, images and embeddings on one bill
79.8
GPT-5.1·
272,000 tokens·
128,000 tokens·
●Yes▸
●Yes▸
◐Partial▸
●Yes·
–
–
$1.25 /M tok▸
$10 /M tok▸
$0.125 /M tok·
50 %▸
●Yes▸
●Yes▸
●Yes▸
◐Partial·
●Yes▸
Automatic, ~5–60 min, no configuration·
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's — closer to a quarter than half a year for API snapshots — and preview models have been pulled faster still.·
90 days·
●Yes▸
◐Partial·
○No▸
–Unknown
2020·
One API for text, reasoning, audio, images and embeddings·
Amazon BedrockMulti-vendor model access inside your existing AWS account
78.1
Multi-vendor — Claude Opus 4.8, Llama 4, Mistral, Nova Premier·
1,000,000 tokens·
128,000 tokens·
●Yes▸
◐Partial·
◐Partial·
○No·
–
–
–
–
–
50 %·
●Yes▸
●Yes▸
●Yes·
●Yes▸
○No·
5 min default, 1 h option; explicit breakpoints only▸
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.▸
–
●Yes▸
●Yes▸
○No▸
●Yes·
2023·
Frontier models inside your existing AWS security perimeter·
Mistral AIEuropean lab with an open-weight lineage and EU-resident hosting
72.8
Mistral Large 3·
–
–
●Yes▸
◐Partial·
◐Partial▸
●Yes·
–
–
–
–
–
50 %·
●Yes▸
●Yes·
◐Partial·
○No·
◐Partial·
–
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.·
–
●Yes▸
◐Partial·
○No·
○No·
2023·
European models with open weights and deployable anywhere·
QwenAlibaba's model family — huge open-weight range, closed flagship
67.4
Qwen3-Max·
262,144 tokens·
–
◐Partial·
◐Partial·
◐Partial▸
●Yes▸
–
–
$1.2 /M tok·
$6 /M tok·
–
50 %·
●Yes·
◐Partial·
◐Partial·
○No·
◐Partial·
–
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.·
–
●Yes·
–Unknown
◐Partial·
◐Partial·
2023·
The widest open-weight family, plus a closed flagship tier·
Meta LlamaOpen-weight Llama models, hosted almost everywhere but Meta
66.6
Llama 4 Maverick·
1,000,000 tokens▸
–
●Yes▸
○No·
●Yes▸
●Yes·
–
–
–
–
–
–
◐Partial▪
◐Partial·
○No·
○No·
○No·
Host-dependent·
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you — if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.·
–
●Yes·
◐Partial·
○No·
●Yes·
2023·
Open-weight models you can host anywhere, forever·
SHSelf-hosted (vLLM)baseline · Run open weights on your own GPUs behind an OpenAI-shaped API
65.1
Whatever you deploy — DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4·
–
–
◐Partial·
○No·
●Yes·
●Yes▸
–
–
–
–
–
0 %·
◐Partial·
●Yes▸
○No·
○No·
○No·
Automatic prefix cache in VRAM, no TTL▸
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.·
–
●Yes·
●Yes·
○No·
●Yes·
2023·
Your weights, your GPUs, your OpenAI-compatible endpoint·
Together AIServerless and dedicated hosting for open-weight models
64.2
Open-weight catalogue — DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4▸
–
–
◐Partial·
◐Partial·
●Yes·
●Yes▸
–
–
–
–
–
50 %·
◐Partial·
●Yes·
○No·
○No·
○No·
–
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.·
–
●Yes·
–Unknown
○No·
○No·
2023·
Open-weight models with a Western contract and dedicated capacity·
xAIGrok models with an OpenAI-shaped API and live X data access
63.6
Grok 4.1·
–
–
●Yes·
○No·
◐Partial▪
●Yes·
–
–
–
–
–
–
●Yes▸
●Yes·
◐Partial·
○No·
○No·
–
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.·
–
●Yes·
–Unknown
◐Partial·
○No·
2024·
Fast, cheap frontier models with live access to X·
FAFireworks AIFast open-weight inference with strong structured-output support
63.2
Open-weight catalogue — DeepSeek, Kimi K2, Qwen3, Llama 4▸
–
–
◐Partial·
◐Partial·
●Yes·
●Yes▸
–
–
–
–
–
–
◐Partial·
●Yes▸
○No·
○No·
○No·
–
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.·
–
●Yes·
–Unknown
○No·
○No·
2023·
Low-latency open-weight inference with strict structured output·
CohereEnterprise-focused models built for RAG and private deployment
62.9
Command A (command-a-03-2025)·
256,000 tokens·
–
◐Partial·
○No·
◐Partial·
◐Partial·
–
–
$2.5 /M tok·
$10 /M tok·
–
–
●Yes▸
●Yes·
○No·
○No·
○No·
–
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.·
–
●Yes▸
◐Partial·
○No·
◐Partial·
2021·
Enterprise RAG models you can deploy inside your own network·
AnthropicClaude models, built around long agentic runs and tool use
62.7
Claude Fable 5 (claude-fable-5)▸
1,000,000 tokens▸
128,000 tokens▸
●Yes▸
○No▸
○No·
◐Partial·
–
–
$10 /M tok▸
$50 /M tok▸
$1 /M tok·
50 %▸
●Yes▸
●Yes▸
●Yes▸
●Yes▸
●Yes▸
5 min default, 1 h option; automatic or explicit breakpoints▸
Published deprecation page lists a retirement date per model, typically months ahead — Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.·
180 days·
◐Partial▸
◐Partial·
○No·
◐Partial·
2023·
Frontier models for long-horizon agentic work and code·
OpenRouterOne OpenAI-shaped key in front of hundreds of models
60.9
Whichever upstream model you route to (400+ available)▸
–
–
◐Partial·
◐Partial·
◐Partial·
●Yes▸
–
–
–
–
–
0 %·
●Yes·
◐Partial·
◐Partial·
○No·
○No·
Passed through where the upstream supports caching·
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself — a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.·
–
◐Partial·
◐Partial·
○No·
○No·
2023·
One key, one API shape, every model worth calling·
ZAZ.ai (GLM)GLM models with an unusually cheap flat-rate coding plan
59.9
GLM-4.6·
200,000 tokens·
–
◐Partial·
○No·
●Yes▸
●Yes·
–
–
$0.6 /M tok·
$2.2 /M tok·
–
–
●Yes▪
◐Partial·
◐Partial·
○No·
○No·
–
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.·
–
●Yes·
○No·
◐Partial·
○No·
2023·
Coding-focused GLM models with a flat-rate subscription·
MAMoonshot AIKimi models — open-weight agentic performance at low cost
58.8
Kimi K2 Thinking·
256,000 tokens·
–
○No·
○No·
●Yes▸
●Yes▸
–
–
$0.6 /M tok·
$2.5 /M tok·
$0.15 /M tok·
–
●Yes▪
◐Partial·
◐Partial·
○No·
○No·
Explicit context caching, paid per storage hour·
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence — download the checkpoint and continuity is your problem, not theirs.·
–
●Yes·
○No·
◐Partial·
○No·
2023·
Open-weight agentic performance at open-weight prices·
GroqCustom LPU silicon serving open-weight models at extreme speed
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.▪
–
◐Partial·
–Unknown
○No·
○No·
2024·
Custom LPU silicon for open-weight models at very low latency·
CerebrasWafer-scale inference — the fastest tokens per second available
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.▪
–
◐Partial·
–Unknown
○No·
○No·
2024·
Wafer-scale inference — the highest tokens per second available·
DeepSeekFrontier-adjacent models at a small fraction of Western prices
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases — deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.▪
–
○No▪
○No·
●Yes▪
○No·
2023·
Near-frontier quality at a fraction of the token price·
Google GeminiGemini via AI Studio for prototyping or Vertex AI for production79.9
Capability
Flagship model
Gemini 3 Pro·
Context window (tokens)
1,000,000 tokens▸
Max output (tokens)
–
Image input
●Yes▸
Audio in/out
●Yes▸
Open weights
◐Partial▸
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$2 /M tok▸
$/M output (/M tok)
$12 /M tok▸
$/M cache read (/M tok)
$0.2 /M tok·
Batch discount (%)
50 %▸
Agentic
Tool use
●Yes▸
Schema output
●Yes▸
Effort control
●Yes▸
Computer use
◐Partial·
MCP support
◐Partial·
Cache TTL
Explicit caches, default 1 h TTL; implicit caching too·
Governance & continuity
Continuity policy
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.▪
Notice period (days)
–
Pinnable versions
●Yes▸
Zero retention
◐Partial·
Trains on your data
◐Partial·
AU region
●Yes·
Traction
API since
2023·
Positioning
Positioning
Long context and native multimodality across Google's cloud·
OpenAIGPT models plus audio, images and embeddings on one bill79.8
Capability
Flagship model
GPT-5.1·
Context window (tokens)
272,000 tokens·
Max output (tokens)
128,000 tokens·
Image input
●Yes▸
Audio in/out
●Yes▸
Open weights
◐Partial▸
OpenAI-compat API
●Yes·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$1.25 /M tok▸
$/M output (/M tok)
$10 /M tok▸
$/M cache read (/M tok)
$0.125 /M tok·
Batch discount (%)
50 %▸
Agentic
Tool use
●Yes▸
Schema output
●Yes▸
Effort control
●Yes▸
Computer use
◐Partial·
MCP support
●Yes▸
Cache TTL
Automatic, ~5–60 min, no configuration·
Governance & continuity
Continuity policy
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's — closer to a quarter than half a year for API snapshots — and preview models have been pulled faster still.·
Notice period (days)
90 days·
Pinnable versions
●Yes▸
Zero retention
◐Partial·
Trains on your data
○No▸
AU region
–Unknown
Traction
API since
2020·
Positioning
Positioning
One API for text, reasoning, audio, images and embeddings·
Amazon BedrockMulti-vendor model access inside your existing AWS account78.1
Capability
Flagship model
Multi-vendor — Claude Opus 4.8, Llama 4, Mistral, Nova Premier·
Context window (tokens)
1,000,000 tokens·
Max output (tokens)
128,000 tokens·
Image input
●Yes▸
Audio in/out
◐Partial·
Open weights
◐Partial·
OpenAI-compat API
○No·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
50 %·
Agentic
Tool use
●Yes▸
Schema output
●Yes▸
Effort control
●Yes·
Computer use
●Yes▸
MCP support
○No·
Cache TTL
5 min default, 1 h option; explicit breakpoints only▸
Governance & continuity
Continuity policy
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.▸
Notice period (days)
–
Pinnable versions
●Yes▸
Zero retention
●Yes▸
Trains on your data
○No▸
AU region
●Yes·
Traction
API since
2023·
Positioning
Positioning
Frontier models inside your existing AWS security perimeter·
Mistral AIEuropean lab with an open-weight lineage and EU-resident hosting72.8
Capability
Flagship model
Mistral Large 3·
Context window (tokens)
–
Max output (tokens)
–
Image input
●Yes▸
Audio in/out
◐Partial·
Open weights
◐Partial▸
OpenAI-compat API
●Yes·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
50 %·
Agentic
Tool use
●Yes▸
Schema output
●Yes·
Effort control
◐Partial·
Computer use
○No·
MCP support
◐Partial·
Cache TTL
–
Governance & continuity
Continuity policy
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.·
Notice period (days)
–
Pinnable versions
●Yes▸
Zero retention
◐Partial·
Trains on your data
○No·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
European models with open weights and deployable anywhere·
QwenAlibaba's model family — huge open-weight range, closed flagship67.4
Capability
Flagship model
Qwen3-Max·
Context window (tokens)
262,144 tokens·
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
◐Partial·
Open weights
◐Partial▸
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$1.2 /M tok·
$/M output (/M tok)
$6 /M tok·
$/M cache read (/M tok)
–
Batch discount (%)
50 %·
Agentic
Tool use
●Yes·
Schema output
◐Partial·
Effort control
◐Partial·
Computer use
○No·
MCP support
◐Partial·
Cache TTL
–
Governance & continuity
Continuity policy
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
–Unknown
Trains on your data
◐Partial·
AU region
◐Partial·
Traction
API since
2023·
Positioning
Positioning
The widest open-weight family, plus a closed flagship tier·
Meta LlamaOpen-weight Llama models, hosted almost everywhere but Meta66.6
Capability
Flagship model
Llama 4 Maverick·
Context window (tokens)
1,000,000 tokens▸
Max output (tokens)
–
Image input
●Yes▸
Audio in/out
○No·
Open weights
●Yes▸
OpenAI-compat API
●Yes·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
–
Agentic
Tool use
◐Partial▪
Schema output
◐Partial·
Effort control
○No·
Computer use
○No·
MCP support
○No·
Cache TTL
Host-dependent·
Governance & continuity
Continuity policy
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you — if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
◐Partial·
Trains on your data
○No·
AU region
●Yes·
Traction
API since
2023·
Positioning
Positioning
Open-weight models you can host anywhere, forever·
SHSelf-hosted (vLLM)Run open weights on your own GPUs behind an OpenAI-shaped API65.1
Capability
Flagship model
Whatever you deploy — DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4·
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
○No·
Open weights
●Yes·
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
0 %·
Agentic
Tool use
◐Partial·
Schema output
●Yes▸
Effort control
○No·
Computer use
○No·
MCP support
○No·
Cache TTL
Automatic prefix cache in VRAM, no TTL▸
Governance & continuity
Continuity policy
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
●Yes·
Trains on your data
○No·
AU region
●Yes·
Traction
API since
2023·
Positioning
Positioning
Your weights, your GPUs, your OpenAI-compatible endpoint·
Together AIServerless and dedicated hosting for open-weight models64.2
Capability
Flagship model
Open-weight catalogue — DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4▸
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
◐Partial·
Open weights
●Yes·
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
50 %·
Agentic
Tool use
◐Partial·
Schema output
●Yes·
Effort control
○No·
Computer use
○No·
MCP support
○No·
Cache TTL
–
Governance & continuity
Continuity policy
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
–Unknown
Trains on your data
○No·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
Open-weight models with a Western contract and dedicated capacity·
xAIGrok models with an OpenAI-shaped API and live X data access63.6
Capability
Flagship model
Grok 4.1·
Context window (tokens)
–
Max output (tokens)
–
Image input
●Yes·
Audio in/out
○No·
Open weights
◐Partial▪
OpenAI-compat API
●Yes·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
–
Agentic
Tool use
●Yes▸
Schema output
●Yes·
Effort control
◐Partial·
Computer use
○No·
MCP support
○No·
Cache TTL
–
Governance & continuity
Continuity policy
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
–Unknown
Trains on your data
◐Partial·
AU region
○No·
Traction
API since
2024·
Positioning
Positioning
Fast, cheap frontier models with live access to X·
FAFireworks AIFast open-weight inference with strong structured-output support63.2
Capability
Flagship model
Open-weight catalogue — DeepSeek, Kimi K2, Qwen3, Llama 4▸
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
◐Partial·
Open weights
●Yes·
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
–
Agentic
Tool use
◐Partial·
Schema output
●Yes▸
Effort control
○No·
Computer use
○No·
MCP support
○No·
Cache TTL
–
Governance & continuity
Continuity policy
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
–Unknown
Trains on your data
○No·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
Low-latency open-weight inference with strict structured output·
CohereEnterprise-focused models built for RAG and private deployment62.9
Capability
Flagship model
Command A (command-a-03-2025)·
Context window (tokens)
256,000 tokens·
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
○No·
Open weights
◐Partial·
OpenAI-compat API
◐Partial·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$2.5 /M tok·
$/M output (/M tok)
$10 /M tok·
$/M cache read (/M tok)
–
Batch discount (%)
–
Agentic
Tool use
●Yes▸
Schema output
●Yes·
Effort control
○No·
Computer use
○No·
MCP support
○No·
Cache TTL
–
Governance & continuity
Continuity policy
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.·
Notice period (days)
–
Pinnable versions
●Yes▸
Zero retention
◐Partial·
Trains on your data
○No·
AU region
◐Partial·
Traction
API since
2021·
Positioning
Positioning
Enterprise RAG models you can deploy inside your own network·
AnthropicClaude models, built around long agentic runs and tool use62.7
Capability
Flagship model
Claude Fable 5 (claude-fable-5)▸
Context window (tokens)
1,000,000 tokens▸
Max output (tokens)
128,000 tokens▸
Image input
●Yes▸
Audio in/out
○No▸
Open weights
○No·
OpenAI-compat API
◐Partial·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$10 /M tok▸
$/M output (/M tok)
$50 /M tok▸
$/M cache read (/M tok)
$1 /M tok·
Batch discount (%)
50 %▸
Agentic
Tool use
●Yes▸
Schema output
●Yes▸
Effort control
●Yes▸
Computer use
●Yes▸
MCP support
●Yes▸
Cache TTL
5 min default, 1 h option; automatic or explicit breakpoints▸
Governance & continuity
Continuity policy
Published deprecation page lists a retirement date per model, typically months ahead — Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.·
Notice period (days)
180 days·
Pinnable versions
◐Partial▸
Zero retention
◐Partial·
Trains on your data
○No·
AU region
◐Partial·
Traction
API since
2023·
Positioning
Positioning
Frontier models for long-horizon agentic work and code·
OpenRouterOne OpenAI-shaped key in front of hundreds of models60.9
Capability
Flagship model
Whichever upstream model you route to (400+ available)▸
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
◐Partial·
Open weights
◐Partial·
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
–
Batch discount (%)
0 %·
Agentic
Tool use
●Yes·
Schema output
◐Partial·
Effort control
◐Partial·
Computer use
○No·
MCP support
○No·
Cache TTL
Passed through where the upstream supports caching·
Governance & continuity
Continuity policy
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself — a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.·
Notice period (days)
–
Pinnable versions
◐Partial·
Zero retention
◐Partial·
Trains on your data
○No·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
One key, one API shape, every model worth calling·
ZAZ.ai (GLM)GLM models with an unusually cheap flat-rate coding plan59.9
Capability
Flagship model
GLM-4.6·
Context window (tokens)
200,000 tokens·
Max output (tokens)
–
Image input
◐Partial·
Audio in/out
○No·
Open weights
●Yes▸
OpenAI-compat API
●Yes·
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$0.6 /M tok·
$/M output (/M tok)
$2.2 /M tok·
$/M cache read (/M tok)
–
Batch discount (%)
–
Agentic
Tool use
●Yes▪
Schema output
◐Partial·
Effort control
◐Partial·
Computer use
○No·
MCP support
○No·
Cache TTL
–
Governance & continuity
Continuity policy
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
○No·
Trains on your data
◐Partial·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
Coding-focused GLM models with a flat-rate subscription·
MAMoonshot AIKimi models — open-weight agentic performance at low cost58.8
Capability
Flagship model
Kimi K2 Thinking·
Context window (tokens)
256,000 tokens·
Max output (tokens)
–
Image input
○No·
Audio in/out
○No·
Open weights
●Yes▸
OpenAI-compat API
●Yes▸
Performance
TTFT p50
–
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$0.6 /M tok·
$/M output (/M tok)
$2.5 /M tok·
$/M cache read (/M tok)
$0.15 /M tok·
Batch discount (%)
–
Agentic
Tool use
●Yes▪
Schema output
◐Partial·
Effort control
◐Partial·
Computer use
○No·
MCP support
○No·
Cache TTL
Explicit context caching, paid per storage hour·
Governance & continuity
Continuity policy
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence — download the checkpoint and continuity is your problem, not theirs.·
Notice period (days)
–
Pinnable versions
●Yes·
Zero retention
○No·
Trains on your data
◐Partial·
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
Open-weight agentic performance at open-weight prices·
GroqCustom LPU silicon serving open-weight models at extreme speed56.6
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.▪
Notice period (days)
–
Pinnable versions
◐Partial·
Zero retention
–Unknown
Trains on your data
○No·
AU region
○No·
Traction
API since
2024·
Positioning
Positioning
Custom LPU silicon for open-weight models at very low latency·
CerebrasWafer-scale inference — the fastest tokens per second available56.5
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.▪
Notice period (days)
–
Pinnable versions
◐Partial·
Zero retention
–Unknown
Trains on your data
○No·
AU region
○No·
Traction
API since
2024·
Positioning
Positioning
Wafer-scale inference — the highest tokens per second available·
DeepSeekFrontier-adjacent models at a small fraction of Western prices46.2
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases — deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.▪
Notice period (days)
–
Pinnable versions
○No▪
Zero retention
○No·
Trains on your data
●Yes▪
AU region
○No·
Traction
API since
2023·
Positioning
Positioning
Near-frontier quality at a fraction of the token price·
●yes◐partial○no–unknown◆ measured · ▸ vendor-claimed · ▪ community · · inferredExpand a row for the source and date behind every cell.
Every provider on this page will quote you a $/M token figure. Almost none of them will tell you how long the model behind it lives. That asymmetry is the single biggest hidden cost in this category, and it is why this page leads with a continuity column rather than a price column.
For most production work in 2026 the shortlist is short. Anthropic if the workload is agentic — tool calls, long multi-step runs, code — because effort control, prompt caching and MCP are first-class rather than bolted on. OpenAI if you need audio, image generation and text behind one billing relationship, or if your team is already fluent in the Responses API. Google if you need million-token context cheaply, native audio in and out, or an AU-resident deployment through Vertex. Those three are not interchangeable at the prompt level: a prompt tuned on one lands differently on the others, and the effort/thinking knobs have no common vocabulary.
The Chinese open-weight labs — DeepSeek, Moonshot, Qwen, Z.ai — have changed what the floor looks like. DeepSeek's flagship bills $0.42 per million output tokens against $50 for Claude Fable 5 and $15 for Claude Sonnet 5 — roughly a hundred and twenty times cheaper than Anthropic's top-end model, and still around thirty-five times cheaper than the volume tier most teams would actually reach for. For classification, extraction, summarisation and bulk rewriting the quality gap does not justify the price gap. The catch is governance, not capability: retention defaults are permissive, zero-data-retention is generally not on offer, and several expose only a rolling alias that silently upgrades under you. If the data is sensitive, run the weights yourself or route through a host that is contractually clean rather than calling the origin API.
Routers deserve more credit than they get. OpenRouter is the cheapest insurance policy in the category — one integration, hundreds of models, and a real chance that a retired model stays reachable somewhere because an open-weight copy is still hosted. Groq and Cerebras are not general-purpose substitutes; they are latency instruments. If your product's differentiator is that the answer appears instantly — voice, autocomplete, interactive search — they are worth an entire architecture. If it is not, their model menu will frustrate you within a quarter.
Where this is heading: prices keep falling, context windows have stopped being the differentiator, and the fight has moved to agentic fidelity — parallel tool calls that do not degrade, structured output that never breaks schema, caching that survives a long agent loop. Assume the model you ship on today will be retired inside two years. Build the abstraction layer now, keep an evaluation set you can re-run in an afternoon, and treat every provider on this page as replaceable.
The pricing traps nobody puts on the pricing page
Headline $/M is the least useful number in this category. Four things distort it.
First, output tokens dominate. A reasoning model can burn tens of thousands of thinking tokens before it writes a word you see, and those bill at the output rate. A provider with a cheap input rate and an expensive output rate is expensive for reasoning work and cheap for retrieval-augmented answering. Model your actual input:output ratio before comparing.
Second, cached input is where the real money is. Cache reads bill at roughly a tenth of the input rate at Anthropic and OpenAI, and an agent loop resending its history every turn is mostly cache reads. But caching is a prefix match — a timestamp in the system prompt, a non-deterministic JSON serialisation, or a tool list that varies per user destroys it silently. Watch the cache-read token counter, not the invoice.
Third, long-context surcharges. Google charges a higher rate above 200K tokens on its flagship; others quietly meter cache storage per hour. A million-token window advertised at the short-context rate is not what you will pay.
Fourth, batch — but check the column before you budget for it. Where an asynchronous endpoint exists the discount is usually a flat 50%, and Anthropic, OpenAI, Google, Mistral and Qwen all publish that rate, as do Bedrock and Together. It is not universal. DeepSeek offers no batch endpoint at all — its list price already sits below most competitors' batch rates — and neither does OpenRouter, which is a real cost if half your workload could tolerate async. xAI, Moonshot, Z.ai, Cohere, Fireworks and Cerebras publish nothing we could confirm, and Groq's batch discount is documented as existing without a percentage we could stand behind. Where the halving is there and part of your workload tolerates a few hours of latency, it is the one saving no negotiation will beat; where it is not, no amount of committed spend conjures it.
How to make a migration survivable
Assume an eighteen-month life for whatever you ship on. The providers that publish retirement dates are doing you a favour; the ones that don't are the ones that will break you.
Pin a dated snapshot wherever the provider offers one, and record the pin in configuration rather than code. Note that this is getting harder, not easier: Anthropic's current line-up ships as bare aliases with no dated model ID behind them, so 'pin the snapshot' is not always an available move any more, and you fall back on the published retirement date instead.
Keep the provider behind one seam. Not an abstraction that pretends every model is the same — that fails the moment you touch thinking blocks or tool-result shapes — but a single module that owns request construction, retry policy and response parsing per provider. Two implementations behind one interface is the cheap insurance; a universal adapter is the expensive fantasy.
Maintain an evaluation set you can re-run in an afternoon: fifty to two hundred real inputs with graded expected outputs. Without it, a forced migration becomes a multi-week vibe check. With it, a migration is a morning's work and a number you can defend.
Finally, know your escape hatch before you need it. For open-weight models it is total — the weights are on disk and vLLM will serve them for as long as you have GPUs. For closed models it is a router that may still have an upstream, or nothing at all.
Deciding: three workloads, three answers
Bulk text processing — classification, extraction, tagging, translation, summarising a firehose. Price dominates and capability differences barely register. DeepSeek, Qwen, Moonshot and Z.ai are the rational choices, or Groq and Together if you want open weights with a Western contract. Add the batch endpoint if latency is negotiable. Do not put sensitive data through a provider whose retention terms you haven't read.
Agent harnesses — a loop that plans, calls tools, reads results and iterates for minutes at a time. This is where the frontier labs earn their price. You need tool calls that stay well-formed at depth, structured output that is enforced rather than requested, an effort or thinking control so you can trade cost against quality per route, and prompt caching that survives the loop. Anthropic and OpenAI are the credible options; Google is close and cheaper on long context. Budget for the fact that a long agentic turn can now run for minutes on a single request.
Interactive, latency-critical UI — voice, live search, inline completion. Time to first token is the product. Cerebras and Groq are in a different class here, an order of magnitude ahead of first-party endpoints on tokens per second, and the constraint is that you take whichever open-weight models they have chosen to host. Design the feature around the model menu, not the other way around.
What is frontier llm apis compared — price, context, tool use and deprecation policy?
Frontier LLM APIs compared — price, context, tool use and deprecation policy
On toolweight, Frontier LLM APIs compared — price, context, tool use and deprecation policy means the 18 tools benchmarked on this page — Anthropic, OpenAI, Google Gemini, xAI, Meta Llama, Mistral AI, DeepSeek, Qwen, Moonshot AI, Z.ai (GLM), Amazon Bedrock, Cohere, OpenRouter, Together AI, Fireworks AI, Groq, Cerebras, Self-hosted (vLLM) — judged on the same 27 fields, from the same sources, on the same date. The question it exists to answer: Which frontier LLM API should I build on?
How does toolweight compare these?
Pricing on this page is stale within a fortnight. That is not hedging — this category re-prices faster than any other on toolweight, so the page carries a seven-day verification cadence and every price cell is dated. Treat any cell whose date is more than a month old as indicative and click through to the vendor's pricing page before you build a forecast on it. Figures are pay-as-you-go list rates in USD per million tokens, before batch, cache or committed-spend discounts, and before any provider-specific long-context surcharge. One convention worth stating: where a vendor is running an unexpired introductory rate, the column carries the list price and the cell note carries the introductory rate and its expiry date. An introductory rate is not a discount you negotiate — it is what you actually pay today — so read the note before modelling a bill.
Entries here are providers, not models. Because a provider ships a dozen models at a dozen prices, every price, context and capability figure on a row describes that provider's current flagship model, named explicitly in the Flagship model column so no number is orphaned from the thing it measures. Compare rows knowing that a cheaper flagship often sits beside a cheaper mid-tier that would serve your workload better. The killer column, deprecation and continuity policy, records two things: how much notice you get before a model is retired, and whether you can pin a dated snapshot that keeps serving after the alias moves on. Continuity is judged from published deprecation pages and observed retirements, not from marketing claims, and where a provider publishes no policy at all we say so rather than guessing a number. Latency figures are community-reported from third-party benchmarks; toolweight has not run its own TTFT harness yet, so most of that column is honestly blank.
What actually happens when a model I depend on is deprecated?
The alias stops resolving and requests 404. If you pinned a dated snapshot you keep serving until that snapshot's published retirement date; if you called a bare alias you are migrating today. Anthropic and OpenAI publish retirement dates months ahead on dedicated deprecation pages. Several providers publish nothing, and rolling aliases can change behaviour underneath you with no version bump at all.
Is DeepSeek or Kimi really good enough to replace Claude or GPT?
For extraction, classification, summarisation, translation and bulk rewriting, yes — and at roughly a thirty-fifth of the output-token cost of Anthropic's volume model, or about a hundred and twentieth of its top-end one. Which multiple you should care about depends on which model you would otherwise have used; the honest comparison for bulk work is against the volume tier, not the flagship. For long agentic runs with many tool calls, sustained instruction following, and code that has to compile first time, the frontier labs are still measurably ahead. Split the workload rather than picking one provider for everything.
Does an OpenAI-compatible endpoint actually make switching easy?
It makes the transport identical and the semantics different. Chat completions port cleanly. Tool-call formats, reasoning-effort parameters, cached-token accounting, thinking blocks and structured-output enforcement do not, and prompts tuned against one model's instruction-following behaviour regress on another. Budget a re-evaluation, not a config change.
How much does prompt caching really save on an agent loop?
More than any other lever. Cache reads typically bill at 10–25% of the input rate, and an agent loop resends the entire conversation on every turn, so a long run can be 80–90% cache reads. The trap is that caching is a prefix match: a timestamp or a reordered tool list at the front of the prompt silently invalidates everything after it.
Which providers can serve inference from an Australian region?
Through the hyperscalers, reliably: Amazon Bedrock in ap-southeast-2 and Google Vertex AI in australia-southeast1 both serve frontier models in-country. First-party endpoints from Anthropic, OpenAI and xAI route to US or EU infrastructure by default, with data-residency controls that vary by plan. Self-hosting is the only answer that is unambiguously AU-resident.
Should I go through a router like OpenRouter instead of direct?
Use a router when model choice is a runtime decision, when you want a fallback that survives an upstream outage, or when you are still benchmarking. Go direct when you depend on provider-specific features — prompt caching breakpoints, computer use, batch endpoints, enterprise retention terms — because routers expose the intersection of what upstreams support, not the union.
Why is the time-to-first-token column mostly empty?
Because almost nobody publishes it, and self-reported latency is meaningless without a specified prompt length, region and concurrency. The figures shown are community measurements from third-party benchmarks for providers whose entire pitch is speed. toolweight will fill this column when it runs its own harness with published methodology, not before.