Updated September 12, 2026
LLM Token Costs and Efficiency: A Practitioner's Guide (September 2026)
LLM token costs across 15+ providers: per-token pricing, caching mechanics, batch discounts, model routing, and cost optimization for September 2026.
Per-token pricing is the number on the pricing page. It is not your actual cost. The gap between the two is where most teams either save or waste thousands of dollars a month. This post covers both: the raw numbers across every major provider, and the strategies that determine what you actually pay.
Companion post: API Rate Limits Compared (September 2026) covers RPM, TPM, and throughput limits across the same providers.
Table of Contents
- Per-token pricing across providers
- The hidden cost multipliers
- Prompt caching: how it works, what it saves
- Batch APIs: the easiest 50% discount
- Model routing: the architecture that pays for itself
- Token efficiency: not all tokens are equal
- Structured outputs: fewer tokens, same information
- Self-hosting vs. API: when the math flips
- Cost modeling a real workload
- Further reading
All prices verified against official provider documentation as of September 10, 2026. Prices change frequently. Confirm on provider pricing pages before budgeting.
Per-token pricing across providers
Prices in USD per million tokens. Sorted cheapest to most expensive by output cost within each category.
Proprietary flagship models
| Model | Input/M | Output/M | Cached Input/M | Context | Notes |
|---|---|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | $0.003 | 1M | Fast/efficient; V4-Flash 0731 variant within one point of GPT-5.6 Luna on AI Index |
| GPT-5.6 Luna | $0.20 | $1.20 | $0.025 | 1.05M | Cost-sensitive tier of GPT-5.6 family |
| GPT-5.4 Nano | $0.20 | $1.25 | $0.02 | 400K | Cheapest GPT-5.4-class model |
| GPT-4.1 Nano | $0.10 | $0.40 | $0.025 | 1M | OpenAI’s cheapest model overall; budget/high-throughput tier |
| GPT-4.1 Mini | $0.40 | $1.60 | $0.10 | 1M | |
| Gemini 3.7 Flash | $0.75 | $3.75 | — | 1M | Current Google flagship; introductory rate through Dec 31, 2026, then $1.50/$7.50 |
| Gemini 3.6 Flash | $0.75 | $3.75 | — | 1M | Previous Google flagship; same introductory rate, same expiry |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | — | 1M | Cost-efficient tier |
| GPT-5.4 Mini | $0.75 | $4.50 | $0.075 | 400K | Budget GPT-5.4-class tier |
| Grok 4.3 | $1.25 | $2.50 | — | 1M | Previous xAI flagship, still available |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | 200K | Cache write: $1.25/M |
| Mistral Small 4 | $0.15 | $0.60 | — | 256K | Apache 2.0 open-weight |
| Mistral Medium 3.5 | $1.50 | $7.50 | — | 256K | Proprietary; adjustable reasoning_effort |
| Gemini 3.5 Flash | $1.50 | $9.00 | — | 1M | GA since Google I/O 2026; computer use capable |
| Grok 4.5 | $2.00 | $6.00 | — | 500K | xAI current flagship; Opus-class performance per independent evaluations |
| GPT-4.1 | $2.00 | $8.00 | $0.50 | 1M | Still active; not retired |
| GPT-5.6 Terra | $2.00 | $12.00 | — | 1.05M | Mid tier of GPT-5.6 family |
| GPT-5.4 | $2.50 | $15.00 | $0.25 | 1M | |
| Qwen3.7-Max | $2.50 | $7.50 | — | 1M | Previous Alibaba flagship, still available |
| Qwen3.8-Max | not published | not published | — | 1M | 2.4T params; announced Aug 3, 2026; weights release scheduled |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | 1M | $2/$10 is the permanent standard rate — the increase to $3/$15 was cancelled; cache write: $2.50/M |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $0.30 | 1M | Previous balanced tier, still available; cache write: $3.75/M |
| DeepSeek V4-Pro | $0.435 | $0.87 | $0.004 | 1M | |
| Gemini 3.1 Pro | $2–4 | $12–18 | — | 1M | $2/$12 under 200K; $4/$18 above |
| GPT-5.5 Instant | $0.20 | $1.00 | — | 128K | Default ChatGPT model since May 5 |
| GPT-5.5 | $5.00 | $30.00 | $0.50 | 1M | Previous flagship |
| GPT-5.6 Sol | $5.00 | $30.00 | — | 1.05M | Current OpenAI flagship; default in Microsoft 365 Copilot |
| Claude Opus 5 | $5.00 | $25.00 | $0.50 | 1M | Current Anthropic flagship; cache write: $6.25/M |
| Claude Opus 4.8 | $5.00 | $25.00 | $0.50 | 1M | Previous Anthropic flagship; cache write: $6.25/M |
| Mistral Large 3 | $0.50 | $1.50 | — | 256K | 675B MoE / 41B active; Apache 2.0 |
| Muse Spark 1.1 | — | $4.25 | — | — | Meta; $0.26/task; 71.3 coding score |
| Claude Fable 5 | $10.00 | $50.00 | $1.00 | 1M | Permanent tier; 95% SWE-bench; cache write: $12.50/M |
| GPT-5.5 Thinking | $6.00 | $36.00 | — | 1M | Thinking tokens billed at output rate |
| GPT-5.5 Pro | Contact sales | — | — | — | Enterprise only |
Several notes on this table.
Claude Sonnet 5 stays at $2/$10 — the increase was cancelled. Anthropic confirmed on August 10, 2026 that the rise to $3/$15 scheduled for September 1 would not happen, and that $2/$10 per million input/output tokens is now the permanent standard rate. Cost models built on $2/$10 remain correct. The ~30% tokenizer inflation relative to Sonnet 4.6 still applies, so the effective per-character cost sits closer to Sonnet 4.6 than the list-price gap suggests: Sonnet 5 lists a third cheaper than Sonnet 4.6’s $3/$15 but emits roughly 30% more tokens for equivalent text.
Gemini 3.7 Flash (GA, August 2026) is Google’s current flagship, replacing Gemini 3.6 Flash. It adds tunable thinking levels and is positioned for complex coding and agentic workflows. Both 3.7 and 3.6 Flash run on the same $0.75/$3.75 introductory pricing through December 31, 2026, doubling to $1.50/$7.50 on January 1, 2027. New work should target 3.7 rather than 3.6. Gemini 3.5 Flash-Lite ($0.30/$2.50) launched as the cost-efficient tier.
GPT-5.6 Sol remains at $5/$30 per million tokens — identical to GPT-5.5, not cheaper. The “up to 80%” cost reduction OpenAI has cited applies to the lower tiers: Terra at $2/$12 and Luna at $0.20/$1.20. Sol is the default in Microsoft 365 Copilot. The GPT-4.1 family (4.1, Mini, Nano) remains active and is not retired — GPT-4.1 Nano at $0.10/$0.40 is OpenAI’s cheapest model per token, cheaper than GPT-5.4 Nano at $0.20/$1.25.
Claude Opus 5 (July 2026) is Anthropic’s current reasoning flagship at $5/$25. It matches Fable 5-level coding performance while costing approximately 50% less per task on knowledge-intensive workloads, and scored 30.2% on ARC-AGI-3 — nearly four times the previous record. On redeployment after the Fable 5 export-control suspension, Anthropic reroutes flagged Fable 5 requests to Opus 5 rather than Opus 4.8.
Claude Fable 5 confirmed as a permanent tier on July 18, 2026. Available on Claude Platform, Claude.ai, Claude Code, and Cowork (AWS/Google Cloud/Microsoft Foundry). Its 95% SWE-bench score remains the highest published figure for any generally available model.
Grok 4.5 launched approximately July 9, 2026, at $2.00/$6.00 per million tokens on the base sub-200K tier. xAI doubles rates above 200K tokens — confirm against x.ai/api before budgeting long-context workloads.
Qwen3.8-Max was announced August 3, 2026, with 2.4T parameters and a 1M context window. Alibaba has not published per-token pricing. Figures circulating in press coverage do not appear on an official pricing page. Check Alibaba Cloud Model Studio for your account’s actual rate.
DeepSeek V4-Flash cache-hit pricing: the verified figure is $0.003625/M for the Pro variant cache hit; V4-Flash cache hit is $0.0028/M. These are substantially lower than any standard cached input rate from Anthropic or OpenAI and reflect DeepSeek’s aggressive caching discount structure.
Access-gated models: GPT-5.4-Cyber (defensive security, vetted teams only), Claude Mythos 5 (Project Glasswing, vetted US organizations for critical infrastructure and defensive security). Neither has publicly documented per-token pricing.
Reasoning/thinking models
| Model | Input/M | Output/M | Context | Notes |
|---|---|---|---|---|
| DeepSeek R1 | $0.55 | $2.19 | 128K | Cache hits: $0.14/M |
| o4-mini | $1.10 | $4.40 | 200K | Thinking tokens: $1.10/M |
| o3 | $2.00 | $8.00 | 200K | Thinking tokens: $2.00/M |
| GPT-5.4 Thinking | $3.00 | $12.00 | 1M | Thinking tokens billed at output rate |
| GPT-5.5 Thinking | $6.00 | $36.00 | 1M | Thinking tokens billed at output rate |
Thinking tokens are not free. A request that triggers 10,000 thinking tokens before generating a 500-token response costs 10,500 output tokens. Anthropic’s budget_tokens parameter caps thinking tokens per request. OpenAI’s reasoning_effort parameter (low/medium/high) controls thinking depth. Use these to bound costs on reasoning-heavy workloads.
Claude Opus 5’s ARC-AGI-3 score of 30.2% likely reflects extended internal reasoning on novel problem types. Benchmark scores derived at unconstrained thinking budgets will not reflect production costs under budget_tokens limits.
Open-weight models (hosted pricing)
Self-hosted = free. These are representative hosted prices from providers like Groq, Together AI, Fireworks, and Alibaba Cloud.
| Model | Hosted Input/M | Hosted Output/M | Context | License |
|---|---|---|---|---|
| Tencent Hy3 (295B) | ~$0.06 | ~$0.06 | — | Open |
| Llama 4 Scout (109B/17B active) | ~$0.10 | ~$0.40 | 10M | Meta |
| Qwen3.6-27B | ~$0.12 | ~$0.50 | 256K | Apache 2.0 |
| Mistral Small 4 | $0.15 | $0.60 | 256K | Apache 2.0 |
| Llama 4 Maverick | ~$0.15 | ~$0.60 | 1M | Meta |
| Qwen3.6-35B-A3B | ~$0.15 | ~$0.60 | 256K | Apache 2.0 |
| Devstral Small 2 (24B) | $0.10 | $0.30 | 128K | Apache 2.0 |
| Gemma 3 27B | ~$0.20 | ~$0.20 | 128K | Apache 2.0 |
| DeepSeek R1 (hosted) | $0.55 | $2.19 | 128K | MIT |
| Mistral Large 3 (675B MoE, hosted) | $0.50 | $1.50 | 256K | Apache 2.0 |
| GLM-5.2 | ~$0.50 | ~$2.00 | — | Open-source |
| Command A+ (218B MoE, 25B active) | ~$0.50 | ~$1.50 | 256K | Apache 2.0 |
| Devstral 2 (123B) | $0.40 | $2.00 | 256K | Modified MIT |
| Poolside Laguna S 2.1 | ~$0.30 | ~$1.00 | — | Open-weight |
| Kimi K3 (2.8T) | ~$1.00 | ~$3.00 | 1M | Open |
| Thinky Inkling (975B) | ~$0.80 | ~$2.50 | — | Apache 2.0 |
Kimi K3 (Moonshot AI) is the largest open model released to date at 2.8T parameters, with a 1M context window and Opus-class performance at approximately Sonnet-level pricing. Demand was high enough that Moonshot paused new subscriptions within 48 hours of launch.
Thinky Inkling (975B, multimodal, Apache 2.0) is reported to lead US labs on certain benchmarks. Hosted pricing is from third-party inference providers; self-hosted requires substantial multi-GPU capacity.
Tencent Hy3 (295B, open-weights) at roughly $0.06/M tokens is among the cheapest capable options for workloads it handles well. Available on several third-party hosting platforms.
GLM-5.2 is a notable open-source Chinese model in production use — Databricks reported it matched Claude Opus 4.8 on their internal codebase at lower cost.
Poolside Laguna S 2.1 is a compact open-weight coding model with reported strength in self-correction and persistence on long agentic tasks. Hosted pricing is approximate from available inference providers.
Devstral 2 (123B, 72.2% SWE-bench, Modified MIT) holds as the dedicated open-weight coding tier. Devstral Small 2 (24B, 68.0% SWE-bench, Apache 2.0) runs locally and is listed on Mistral’s API at $0.10/$0.30 — the cheapest dedicated coding model in this table.
Qwen3.8-Max is API-only on Alibaba Cloud Model Studio today, with a weights release scheduled but not yet shipped. Until it lands, the open-weight Qwen3.6 family (including Qwen3.6-35B-A3B with 3B active parameters) remains the primary self-hosting target within that family.
Embeddings and other modalities
| Model | Price | Unit | Notes |
|---|---|---|---|
| OpenAI text-embedding-3-small | $0.02/M tokens | tokens | |
| OpenAI text-embedding-3-large | $0.13/M tokens | tokens | |
| Gemini Embedding 2 | Free (AI Studio) | tokens | Multimodal (text, image, video, audio) |
| Cohere Embed v3 | $0.10/M tokens | tokens | Multilingual |
| Whisper (OpenAI) | $0.006/min | audio minutes | |
| GPT Transcribe / GPT Live Transcribe | not published | audio minutes | Now available via API; independent benchmarks show higher error rates than ElevenLabs, Google, and Mistral |
| Whisper (Groq, free tier) | $0.00 | audio minutes | 8 hrs/day limit |
| Gemini 3.1 Flash TTS | Contact Google | characters | Dedicated TTS model; 70+ languages |
| Gemini Omni Flash | Contact Google | — | Multimodal any-input; video generation/editing; launched Google I/O 2026 |
| DeepMind Lyria 3.5 | Contact Google | — | Music generation; integrated into Google Flow Music |
| OpenAI TTS | $15.00/M chars | characters | |
| OpenAI TTS HD | $30.00/M chars | characters | |
| GPT Image 2 | $0.02–$0.17 | per image | Varies by quality/size |
The hidden cost multipliers
The pricing table above is the starting point, not the answer. Several factors multiply your actual spend beyond the per-token rate.
1. Output tokens cost 2–6x more than input tokens
Every major provider charges more for output than input. The ratio varies:
| Provider/Model | Typical Output:Input Ratio |
|---|---|
| DeepSeek V4-Flash | 2:1 |
| Grok 4.5 | 3:1 |
| Grok 4.3 | 2:1 |
| Qwen3.7-Max | 3:1 |
| GPT-5.4 | 6:1 |
| GPT-5.6 Terra | 6:1 |
| Claude Sonnet 5 | 5:1 |
| Claude Opus 5 | 5:1 |
| GPT-5.5 | 6:1 |
| GPT-5.6 Sol | 6:1 |
Prompt engineering that reduces output length has an outsized cost impact. A system prompt that gets the model to return 200 tokens instead of 500 tokens saves more than a system prompt that reduces input by the same 300 tokens.
2. Context length pricing tiers
Google charges different rates based on how much of the context window you use. Gemini 3.1 Pro costs $2/$12 per million tokens under 200K, but $4/$18 above — a 50% price increase for long-context workloads. Gemini 3.7 Flash and 3.6 Flash’s flat pricing regardless of context depth is a practical advantage for workloads that push toward the 1M token window. xAI doubles rates above 200K tokens on Grok 4.5 — the $2/$6 figure applies only to the sub-200K tier.
3. Thinking token overhead
Extended thinking is not free. A request to Claude Opus 5 that triggers 20,000 thinking tokens at $25/M output adds $0.50 to that single request before the actual response tokens are counted. For complex reasoning tasks, thinking tokens can exceed response tokens by 10–50x. Anthropic, OpenAI, and Google all bill thinking tokens at the same rate as standard output tokens for their respective models.
Opus 5’s ARC-AGI-3 performance almost certainly involves substantial thinking-token use — benchmark scores at unconstrained thinking budgets will not reflect production costs under budget_tokens limits.
4. Cache write costs
Anthropic charges a 25% premium on tokens written to cache versus standard input. Claude Opus 5: $6.25/M for cache writes versus $5.00/M for standard input. The write premium pays for itself after the second cache read ($0.50/M), so any cached prefix used three or more times is net positive. Single-use prompts should skip caching. Claude Sonnet 5 follows the same pattern at $2.50/M cache write versus $2.00/M standard input.
5. Claude Sonnet 5 tokenizer inflation — at a price that did not rise
Sonnet 5’s tokenizer emits approximately 30% more tokens for the same text than the tokenizer used in Sonnet 4.6. Its $2/$10 rate is permanent — the September 1 increase to $3/$15 was cancelled — so Sonnet 5 lists a third below Sonnet 4.6 while emitting roughly 30% more tokens for the same text. Those two effects very nearly cancel: expect rough per-task parity with Sonnet 4.6, not the 33% saving the list prices imply. Whether it comes out ahead depends on whether quality improvements allow shorter outputs or fewer retries on your workload. Measure rather than assume.
6. Image and multimodal token costs
Images are converted to tokens before processing. A single high-resolution image can consume 1,000–5,000+ tokens depending on the model and resolution settings. For vision-heavy workloads — document OCR, screenshot analysis, diagram interpretation — image token costs can dominate the bill. Resize images to the minimum resolution needed before sending them.
Gemini 3.7 Flash’s agentic and multimodal positioning, combined with Gemini 3.5 Flash’s computer-use capability (screenshots, clicking, UI navigation), creates a pattern where a single agentic session accumulates dozens of screenshot tokens across a task sequence. Budgeting per-session rather than per-request is more accurate for those workloads.
7. Autonomous operation and agentic loops
An agent that reads its own output history accumulates input tokens roughly quadratically over a long session. Qwen3.7-Max’s 35-hour autonomous operation demonstration illustrates the ceiling — a session of that length, even at $2.50/M input, could accumulate millions of context tokens across iterative steps. Claude Code Auto Mode (now the default for Claude Code) introduces similar dynamics: its safety classifier processes each command, and multi-step coding tasks generate context that grows with each tool call. Claude Code instances running in parallel on macOS and Linux can now message each other and share context across terminals, which creates multi-agent context accumulation patterns that are difficult to budget without session-level token caps. Budget token limits at the task level, not just the request level.
8. GPT-5.6 pricing: the savings come from tiering, not from the flagship
Sol lists at $5/$30 per million tokens — identical to GPT-5.5. The savings live in the lower tiers: Terra at $2/$12 is 60% below GPT-5.5, and Luna at $0.20/$1.20 is 96% below it. Budget from the specific variant you plan to call, not from the family-level marketing claim.
9. Gemini introductory pricing has a hard expiry
Both Gemini 3.7 Flash and Gemini 3.6 Flash run at $0.75/$3.75 through December 31, 2026, then double to $1.50/$7.50. A cost model built on the current number has a known expiry date four months out. Any multi-year or calendar-year budget for Google’s Flash tier should incorporate both rates.
10. Access-gated models
GPT-5.4-Cyber is available only to vetted teams working on defensive security. Claude Mythos 5 is restricted to vetted US organizations via Project Glasswing. Neither has publicly documented per-token pricing. Workload planning that assumes access to either model requires factoring in the vetting process separately from cost modeling.
Prompt caching: how it works, what it saves
Prompt caching stores the processed representation of input tokens on the provider side so subsequent requests with the same prefix skip re-processing.
Provider comparison
| Provider | Cache Read Discount | Cache Write Premium | Min Cacheable | TTL | Rate Limit Benefit |
|---|---|---|---|---|---|
| Anthropic | 90% off input | 25% over input | 1,024 tokens | 5 min (auto-extends) | Cached reads don’t count toward ITPM |
| OpenAI | 90% off input | None (automatic) | Varies | Automatic | No rate limit benefit |
| 75% off input | None (usage-based) | 4,096 tokens | Configurable | No rate limit benefit | |
| DeepSeek | ~99% off input (V4-Pro cache hit: $0.004/M vs $0.435/M standard) | None | Automatic | Varies | No rate limit benefit |
Anthropic’s caching has the most generous rate limit interaction: cached read tokens are excluded from ITPM limits entirely. With 80% cache hits, a 2M ITPM limit effectively handles 10M total input tokens per minute. This is a throughput multiplier, not just a cost reduction.
DeepSeek’s cache hit pricing deserves attention: V4-Pro cache hits cost $0.003625/M input versus $0.435/M standard — a reduction of over 99%. V4-Flash hits $0.0028/M versus $0.14/M standard. For workloads with stable prefixes and high request volume, DeepSeek’s cache economics are the most aggressive in the market.
Claude Opus 5 cache pricing ($0.50/M reads, $6.25/M writes) follows the same structural ratios — 90% off input for reads, 25% premium for writes — scaled to the $5.00/M list price. Re-benchmark cached prefix sizes on actual prompts when moving between Anthropic models; Sonnet 5’s tokenizer inflation means the same text uses more tokens, which affects cache write costs.
When caching pays off
The break-even math for Anthropic: cache write costs 25% more than standard input. Cache reads cost 90% less. After two cache reads, the cost is net positive even accounting for the write premium. For a system prompt reused across thousands of requests, the savings approach 90% on that portion of every request.
Practical caching patterns
RAG with stable context: Cache the system prompt and retrieved documents as a prefix. Each follow-up query appends to the cached prefix. Works well when the same document set is queried multiple times within the TTL window.
Multi-turn conversations: Cache the conversation history prefix. Each new turn extends the cache. Anthropic’s 5-minute auto-extending TTL means active conversations stay cached indefinitely.
Batch evaluation: Cache the evaluation rubric and few-shot examples. Run hundreds of evaluations against the cached prefix. The rubric tokens (often 2,000–5,000) are paid once at write price, then read at 90% discount for every subsequent evaluation.
Agentic task context: For long-running autonomous tasks, cache the task description, tool definitions, and accumulated working memory up to the stable prefix boundary. Claude Code Auto Mode sessions, Gemini 3.7 Flash agentic tasks via Managed Agents, and long-horizon Qwen3.8-Max runs all benefit from this pattern. For Gemini Flash computer-use sessions, caching the stable task description and tool schema before the screenshot stream begins can meaningfully reduce costs on long UI-navigation tasks. Claude Code instances running parallel on multiple terminals compound the benefit — a cached shared context is read multiple times simultaneously.
Batch APIs: the easiest 50% discount
OpenAI, Anthropic, and Google all offer batch processing with lower pricing and higher throughput limits.
| Provider | Batch Discount | Turnaround | Queue Limits |
|---|---|---|---|
| OpenAI | 50% off standard pricing | Up to 24 hours | 3x–100x higher than real-time TPM |
| Anthropic | 50% off standard pricing | Up to 24 hours | Higher than real-time |
| 50% off standard pricing | Up to 24 hours | Available for Gemini 3.1+ models |
OpenAI’s batch queue limits are notably generous: at Tier 5, the batch queue reaches 15 billion tokens. The 50% discount applies to whichever pricing tier is in effect at submission time.
For Claude Sonnet 5, batch pricing is $1.00/$5.00 per million input/output — the most accessible entry point to Sonnet 5-class quality, and half the permanent $2/$10 real-time rate. For Claude Opus 5 at $5.00/$25.00 standard, batch pricing of $2.50/$12.50 positions it closer to Sonnet-tier costs on a per-token basis, which changes the routing calculus for non-real-time complex workloads.
Workloads that fit batch processing: content generation pipelines, bulk classification, evaluation suites, data extraction from document archives, nightly report generation, synthetic data creation. Any workload where the output is consumed minutes or hours after generation rather than in real time.
Model routing: the architecture that pays for itself
The most impactful cost optimization is not using one model for everything. A routing layer that sends requests to different models based on complexity can cut spend 60–85%.
Cost comparison: routed vs. single-model
Assume 100,000 requests/day, average 500 output tokens per request.
Single model (Claude Sonnet 5):
- 100K requests × 500 tokens = 50M output tokens/day
- 50M × $10/M = $500/day
Note: Sonnet 5’s tokenizer produces ~30% more tokens for equivalent output compared to Sonnet 4.6. Actual token counts per request should be measured on real workload data rather than assumed from prior benchmarks.
Three-tier routing (70/25/5 split):
- 70K requests × 500 tokens × $5/M (Haiku 4.5) = $175.00
- 25K requests × 500 tokens × $10/M (Sonnet 5) = $125.00
- 5K requests × 500 tokens × $25/M (Opus 5) = $62.50
- Total: $362.50/day (28% savings over Sonnet 5 alone)
The classifier itself costs almost nothing. Running GPT-4.1 Nano ($0.10/M input) or Haiku 4.5 ($1.00/M input) to classify intent on a 100-token input across 100K daily requests: roughly 10M input tokens per day, or $1–10/day.
For coding-specific routing, Fable 5 remains the clearest complex-tier target at 95% SWE-bench. Opus 5 matches its coding performance while costing 50% less per task on knowledge tasks, making it a reasonable alternative depending on the mix of coding versus reasoning in the complex tier.
For factual accuracy-constrained workloads, Grok 4.5 at $6.00/M output deserves evaluation as a medium-tier option. Its independent evaluation at Opus-class quality sits between Haiku 4.5 and Sonnet 5 on output cost, making it worth benchmarking before committing to the routing split.
Gemini 3.7 Flash at $3.75/M output (introductory) slots naturally into the medium tier for teams already using Google infrastructure. The rate doubles January 1, 2027 — any routing architecture built around the introductory number should have a migration plan for the standard rate.
Open-weight routing for cost floors
For teams with GPU infrastructure, routing the simple tier to a self-hosted model (Llama 4 Scout, Qwen3.6-35B-A3B, Tencent Hy3, Command A+, Devstral Small 2) eliminates the per-token cost entirely on the majority of requests. Kimi K3’s 2.8T parameter scale makes it impractical for typical self-hosting, but hosted pricing at approximately Sonnet-level cost makes it a candidate for the medium tier in routing pipelines. Command A+‘s Apache 2.0 license and 2×H100 inference footprint (or single B200) remain practical for the medium tier at near-zero marginal cost above infrastructure. Devstral 2’s 72.2% SWE-bench score holds as the dedicated open-weight coding tier — the gap to Fable 5 on hard coding problems is substantial, but Devstral 2 remains competitive with mid-tier proprietary models on most code tasks at self-hosted inference costs.
Token efficiency: not all tokens are equal
The same task produces different token counts on different models. This is the most underappreciated cost variable.
GPT-5.6’s efficiency and tiering
OpenAI attributes GPT-5.6’s cost reductions to recursive self-improvement and distillation. GPT-5.5 was reported as ~40% more token-efficient than GPT-5.4 on coding tasks; whether GPT-5.6 compounds that further is not yet documented in third-party evaluations. The DeepSeek V4-Flash 0731 result — within one point of GPT-5.6 Luna on the Artificial Analysis Intelligence Index at substantially lower cost per task — suggests that per-token price cuts alone do not guarantee task-level cost leadership.
Claude Opus 5 task-level economics
Anthropic’s claim that Opus 5 costs roughly 50% less per task than Fable 5 on knowledge-intensive workloads is directly supported by list pricing: at $5/$25 per million tokens versus Fable 5’s $10/$50, Opus 5 is half the price per token. Any additional saving from Opus 5 reaching the correct answer in fewer reasoning steps or fewer retry loops still needs to be measured on representative workloads before it is assumed in a budget.
Claude Sonnet 5 tokenizer effects
Sonnet 5 lists at $2/$10 against Sonnet 4.6’s $3/$15 — a third cheaper — but emits ~30% more tokens for equivalent text. The two effects roughly cancel, landing near per-task parity rather than the 33% saving the list prices suggest. Benchmark on representative workloads before assuming either a saving or a penalty.
Grok 4.5’s per-task cost position
At $2.00/$6.00 per million tokens with a 3:1 output-to-input ratio, Grok 4.5 sits below the frontier flagships on the cost curve for a model with Opus-class performance claims. For workloads where factual accuracy or instruction following is the binding quality constraint, benchmark it on a per-task basis rather than relying on per-token comparisons alone.
Tokenizer differences matter across providers
The same English sentence might be 15 tokens on OpenAI’s tiktoken, 18 tokens on Anthropic’s Sonnet 4.6 tokenizer, and approximately 23 tokens on Sonnet 5’s tokenizer. For pricing comparisons across providers, convert to a common unit (characters or words) rather than comparing raw per-token prices.
A rough heuristic: 1 token ≈ 4 characters ≈ 0.75 words in English. This varies by language (CJK characters often consume 1–2 tokens each) and by content type (code is tokenized differently than prose). Sonnet 5’s tokenizer shifts this ratio to approximately 1 token ≈ 3 characters in English.
Prompt engineering for token efficiency
Small changes to prompts can produce large token savings at scale:
- “Respond in JSON” vs. natural language: structured responses are 30–60% shorter.
- “Be concise” in the system prompt: an explicit brevity instruction reduces output tokens by 20–40% on average.
- Few-shot examples: Adding 2–3 examples of desired output length calibrates verbosity. This costs a few hundred extra input tokens but saves thousands of output tokens across a batch.
- Max tokens parameter: Set a hard cap. Prevents runaway generation on edge cases where the model would otherwise produce 10x the expected output.
Structured outputs: fewer tokens, same information
JSON mode and structured outputs (GA on OpenAI, Anthropic, and Google) constrain the model to return data in a predefined schema. This is both a reliability and a cost optimization.
A sentiment classification task:
Unstructured response (typical): “The sentiment of this review is positive. The customer expresses satisfaction with the product quality and delivery speed, though they note minor concerns about packaging.” (~35 tokens)
Structured response: {"sentiment": "positive", "confidence": 0.92} (~12 tokens)
That is a 65% reduction in output tokens per request. At $10/M output (Sonnet 5), across 1M daily classifications, the savings are: (35−12) × 1M / 1M × $10 = $230/day.
For any pipeline where the downstream consumer is code rather than a human reading prose, structured outputs should be the default.
Self-hosting vs. API: when the math flips
Open-weight models (Llama 4, Qwen3.6, Mistral, DeepSeek, Command A+, Devstral 2, Kimi K3, Thinky Inkling, Tencent Hy3) are free to run. The cost is infrastructure.
Break-even estimation
A single H100 (80GB) rents for roughly $2–3/hour on cloud providers. It can serve a 70B model (quantized to 4-bit) at approximately 30–50 tokens/second.
Monthly H100 cost: ~$2,000
Monthly tokens at 40 tok/s sustained: ~100B tokens
Equivalent API cost for 100B output tokens:
- Claude Haiku 4.5: $500,000
- Claude Sonnet 5: $1,000,000
- GPT-4.1 Nano: $40,000
- DeepSeek V4-Flash: $28,000
- Grok 4.5: $600,000
Self-hosting a 70B open model at $2,000/month breaks even against DeepSeek V4-Flash at roughly 7B output tokens/month. Against Haiku 4.5, break-even is under 500M tokens/month.
Kimi K3 (2.8T parameters) is too large for most self-hosting setups. Hosted pricing at approximately $3/M output positions it as an API option rather than a self-hosting target for most teams.
Thinky Inkling (975B, Apache 2.0) is more feasible to self-host than Kimi K3 but still requires significant multi-GPU infrastructure. For teams already operating H100 clusters at scale, its Apache 2.0 license and reported benchmark leadership make it worth evaluating.
Command A+ changes the calculus meaningfully. At 218B total parameters but only 25B active, it runs on 2×H100 (~$4,000/month) while delivering quality closer to mid-tier proprietary models than typical 70B open-weights. The B200 option reduces infrastructure cost further for teams with access to newer hardware. Its Apache 2.0 license and no published per-token API price (free on API until rate limits) make it unusual in this market.
Devstral 2 occupies a narrow but useful niche: a 123B coding-focused model with 72.2% SWE-bench and 256K context. At self-hosted inference costs, its performance is difficult to match with API alternatives at comparable volume. Devstral Small 2 (24B, Apache 2.0) runs locally with 68.0% SWE-bench — a credible option for coding workloads on single-GPU consumer hardware.
When APIs still win
- Low volume: Under ~5B tokens/month, API costs are lower than maintaining infrastructure.
- Burst traffic: APIs absorb spikes. Self-hosted GPUs sit idle during troughs and cannot handle peaks without over-provisioning.
- Frontier quality: No open model approaches Claude Fable 5 or Opus 5 on complex coding, or matches Opus 5’s ARC-AGI-3 score of 30.2% on novel reasoning. Tasks requiring the best available performance leave no self-hosted alternative.
- Compliance and support: Enterprise API tiers include SLAs, BAAs (HIPAA), data residency controls, and zero-retention options that would require substantial additional investment to replicate on self-hosted infrastructure.
When self-hosting wins
- High sustained volume: Above 10B+ tokens/month, owned or reserved GPU capacity typically beats API pricing.
- Data sovereignty: Air-gapped or on-premise requirements.
- Fine-tuning: Self-hosted models can be fine-tuned on proprietary data with full control. API-based fine-tuning is available from some providers but with less flexibility.
- Latency control: No network round-trip. Co-located inference can achieve sub-10ms TTFT.
Cost modeling a real workload
Pricing tables are abstractions. A concrete example of modeling costs for a production RAG pipeline follows.
Scenario: customer support agent
- 50,000 queries/day
- Each query: 200 tokens input (user question), 2,000 tokens retrieved context, 300 tokens system prompt, 400 tokens output
- System prompt + retrieval context is 80% cacheable (same FAQ docs across queries)
Important: These token counts assume a Sonnet 4.6-tokenized baseline. Running on Sonnet 5, the ~30% tokenizer inflation means 2,500 input tokens likely becomes ~3,250 input tokens for the same text, and the output column is similarly affected. Measure on actual workload data before committing to a cost model.
Option A: Claude Sonnet 5, no optimization
| Component | Tokens/request | Daily tokens | Cost/M | Daily cost |
|---|---|---|---|---|
| Input (system + context + query) | 2,500 (est. ~3,250 actual on Sonnet 5) | ~162.5M | $2.00 | $325.00 |
| Output | 400 (est. ~520 actual on Sonnet 5) | ~26M | $10.00 | $260.00 |
| Total | $585.00/day |
Monthly: ~$17,550. The $2/$10 rate the prior edition used still stands — Anthropic cancelled the September 1 increase — so the rise from that edition’s $450/day is tokenizer inflation alone, not a price change.
Option B: Claude Sonnet 5, with prompt caching
| Component | Tokens/request | Daily tokens | Cost/M | Daily cost |
|---|---|---|---|---|
| Cache write (system + context, first request/5min) | 2,300 (est. ~2,990 actual) | ~0.9M | $2.50 | $2.25 |
| Cache read (system + context, subsequent) | 2,300 (est. ~2,990 actual) | ~143.5M | $0.20 | $28.70 |
| Uncached input (user query) | 200 (est. ~260 actual) | ~13M | $2.00 | $26.00 |
| Output | 400 (est. ~520 actual) | ~26M | $10.00 | $260.00 |
| Total | $316.95/day |
Monthly: ~$9,509. 46% savings from caching alone. The ratio is unchanged by the price level — caching savings scale with the rate, not against it.
Option C: Routed, with caching
Route 70% of queries (simple FAQ-type) to Haiku 4.5, 30% (complex) to Sonnet 5. Both with caching.
| Component | Daily cost |
|---|---|
| Haiku tier (35K queries, cached) | ~$52.50 |
| Sonnet tier (15K queries, cached) | ~$101.00 |
| Classifier (50K queries, Haiku 4.5 at $1.00/M input) | ~$0.50 |
| Total | ~$154.00/day |
Monthly: ~$4,620. 74% savings vs. Option A. (The routing gain is smaller than it looks at a higher Sonnet price: Haiku’s cost is unchanged, so a cheaper Sonnet tier narrows the gap routing is closing.)
Option D: Add batch processing for non-real-time queries
If 20% of queries (overnight analytics, report generation) can tolerate 24-hour turnaround:
| Component | Daily cost |
|---|---|
| Real-time routed (40K queries) | ~$123.20 |
| Batch (10K queries, 50% discount on Sonnet 5 = $1.00/$5.00) | ~$13.63 |
| Total | ~$136.83/day |
Monthly: ~$4,105. 77% savings vs. Option A.
The progression from ~$17,550/month to ~$4,105/month uses no different models, no degradation in output quality, and no self-hosting. It is entirely a function of caching, routing, and batching.
Quick reference: cost per 1M output tokens
Sorted cheapest to most expensive. Input costs excluded.
| Model | Output $/M | Type |
|---|---|---|
| Tencent Hy3 (295B) | ~$0.06 | Open-weight |
| DeepSeek V4-Flash | $0.28 | Proprietary |
| Devstral Small 2 (hosted) | $0.30 | Open-weight |
| GPT-4.1 Nano | $0.40 | Proprietary |
| Llama 4 Scout (hosted) | ~$0.40 | Open-weight |
| Mistral Small 4 | $0.60 | Open-weight |
| DeepSeek V4-Pro | $0.87 | Proprietary |
| GPT-5.5 Instant | $1.00 | Proprietary |
| GPT-5.6 Luna | $1.20 | Proprietary |
| GPT-5.4 Nano | $1.25 | Proprietary |
| GPT-4.1 Mini | $1.60 | Proprietary |
| DeepSeek R1 | $2.19 | Open-weight |
| Gemini 3.5 Flash-Lite | $2.50 | Proprietary |
| Grok 4.3 | $2.50 | Proprietary |
| Gemini 3.7 Flash (intro) | $3.75 | Proprietary |
| Gemini 3.6 Flash (intro) | $3.75 | Proprietary |
| Muse Spark 1.1 | $4.25 | Proprietary |
| GPT-5.4 Mini | $4.50 | Proprietary |
| Claude Haiku 4.5 | $5.00 | Proprietary |
| Grok 4.5 | $6.00 | Proprietary |
| Qwen3.7-Max | $7.50 | Proprietary |
| Mistral Medium 3.5 | $7.50 | Proprietary |
| GPT-4.1 | $8.00 | Proprietary |
| Gemini 3.5 Flash | $9.00 | Proprietary |
| Claude Sonnet 5 | $10.00 | Proprietary |
| GPT-5.6 Terra | $12.00 | Proprietary |
| GPT-5.4 Thinking | $12.00 | Proprietary |
| Gemini 3.1 Pro | $12–18 | Proprietary |
| GPT-5.4 | $15.00 | Proprietary |
| Claude Sonnet 4.6 | $15.00 | Proprietary |
| Claude Opus 4.8 | $25.00 | Proprietary |
| Claude Opus 5 | $25.00 | Proprietary |
| GPT-5.5 | $30.00 | Proprietary |
| GPT-5.6 Sol | $30.00 | Proprietary |
| GPT-5.5 Thinking | $36.00 | Proprietary |
| Claude Fable 5 | $50.00 | Proprietary |
Further reading
OpenAI
- Pricing: https://openai.com/api/pricing
- Batch API: https://platform.openai.com/docs/guides/batch
- Prompt caching: https://platform.openai.com/docs/guides/prompt-caching
- Tokenizer (tiktoken): https://github.com/openai/tiktoken
Anthropic
- Pricing: https://www.anthropic.com/pricing
- Prompt caching: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
- Message Batches API: https://docs.anthropic.com/en/docs/build-with-claude/message-batches
- Token counting: https://docs.anthropic.com/en/docs/build-with-claude/token-counting
- Claude Opus 5 announcement: https://www.anthropic.com/news/claude-opus-5
- Claude Fable 5 announcement: https://www.anthropic.com/news/claude-fable-5-mythos-5
- Claude Sonnet 5 announcement: https://www.anthropic.com/news/claude-sonnet-5
- Gemini API pricing: https://ai.google.dev/pricing
- Context caching: https://ai.google.dev/gemini-api/docs/caching
- Vertex AI pricing: https://cloud.google.com/vertex-ai/generative-ai/pricing
- Gemini API model docs: https://ai.google.dev/gemini-api/docs/models
- Gemini Managed Agents: https://ai.google.dev/gemini-api/docs/agents
xAI
- Grok API pricing: https://x.ai/api
DeepSeek
- API pricing: https://api-docs.deepseek.com/quick_start/pricing
Mistral
- Pricing: https://mistral.ai/technology/#pricing
- Devstral 2 announcement: https://mistral.ai/news/devstral
Cohere
- Pricing: https://cohere.com/pricing
- Command A+ announcement: https://cohere.com/blog/command-a
Alibaba / Qwen
- Alibaba Cloud model pricing: https://www.alibabacloud.com/en/product/bailian/pricing
- Qwen model announcements: https://qwenlm.github.io/blog/
Moonshot / Kimi
- Kimi K3 model page: https://kimi.moonshot.cn
Poolside
- Laguna S 2.1: https://poolside.ai
Meta
- Muse Spark 1.1: https://ai.meta.com
Third-party trackers
- Live pricing across 300+ models: https://pricepertoken.com
- Model benchmarks with cost context: https://artificialanalysis.ai
Prices verified September 10, 2026. The LLM pricing landscape shifts every 2–4 weeks. Use pricepertoken.com for live tracking. Confirm against official provider documentation before production budgeting.