API Rate Limits Compared: Every Major LLM Provider (August 2026)
API rate limits for every major LLM provider — August 10, 2026. Side-by-side tables for OpenAI, Anthropic, Google, Groq, xAI, DeepSeek, Mistral, Cerebras, SambaNova, and more.
Ninth edition. The eighth (July 10, 2026) covered 17 providers. This update reflects the four-week window since July 10. GPT-5.6 Sol is now broadly available, replacing GPT-5.5 as OpenAI’s flagship. Claude Opus 5 is Anthropic’s new flagship, replacing Opus 4.8. Google’s current flagship is now Gemini 3.6 Flash. xAI’s current flagship is Grok 4.5, replacing Grok 4.3. Alibaba’s current flagship is Qwen3.8-Max. The DeepSeek V3.2 alias deprecation passed on July 24 — those endpoints are gone. Kimi K3 is now the largest open-weight model released to date.
Table of Contents
- What changed since July 10
- Free Tier — RPM & RPD
- Free Tier — TPM & TPD
- Entry Paid Tier — RPM & RPD
- Entry Paid Tier — TPM
- Scaled Tier — RPM & TPM
- Cerebras & SambaNova
- Audio Models — ASH & ASD
- Cloud Aggregators — Azure AI & AWS Bedrock
- More Providers — Perplexity, Alibaba, Moonshot
- Provider Notes
- Tips for managing rate limits
- Further Reading
Last updated: August 10, 2026.
What changed since July 10
GPT-5.6 Sol is now broadly available. Previewed June 27 and restricted to trusted partners at the time of the July 10 edition, GPT-5.6 Sol began launching broadly in the July 9–11 window and is now OpenAI’s current flagship. Variants Sol, Terra, and Luna are rolling out in tiers. GPT-5.6 is now the default model in Microsoft 365 Copilot. Pricing has been cut by up to 80% relative to GPT-5.5 through recursive self-improvement and distillation — roughly a 13x cost reduction over four months. Rate limit tier structure for GPT-5.6 Sol is not yet fully published in OpenAI’s public docs; tables below carry forward GPT-5.5 figures as a provisional floor and note where Sol-specific figures are absent. GPT-5.5 remains available. GPT-5.3 Codex has been consolidated into ChatGPT as a unified superapp interface and is no longer a standalone API endpoint. GPT-4.1 Nano replaces GPT-5.4 Nano as the recommended budget/high-throughput model. OpenAI also launched GPT Transcribe and GPT Live Transcribe via API, though independent benchmarks show both lag ElevenLabs, Google, and Mistral on error rates.
Claude Opus 5 is Anthropic’s new flagship. Released July 2026, Opus 5 matches Fable 5-level performance on coding and knowledge tasks at 50% lower token cost and scored 30.2% on ARC-AGI-3 — nearly 4x the previous record on that benchmark. Rate limit tier structure for Opus 5 is not yet fully published; tables carry forward Opus 4.8 figures as a provisional floor. Opus 4.8 remains available. Claude Fable 5 is now confirmed as a permanent production tier (July 18, 2026); rate limit figures remain unpublished. The Sonnet 5 introductory pricing ($2/$10 per M input/output) runs through August 31, 2026 — three weeks out — then shifts to standard $3/$15. The Fable 5 classifier added on July 1 redeployment now routes flagged requests to Opus 5 rather than Opus 4.8.
Google’s current flagship is Gemini 3.6 Flash. Launched July 2026, Gemini 3.6 Flash is now listed first in the official Gemini API model docs and balances speed with intelligence for agentic and multimodal tasks. Gemini API Managed Agents support it with production-grade hooks. Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber (cybersecurity domain) also launched July 2026. Rate limit figures for Gemini 3.6 Flash are not yet published in public documentation — check AI Studio (aistudio.google.com/rate-limit) for your project’s actual limits. Gemini 3.5 Flash (GA since May 19) remains available as the previous flagship. Gemini 3.5 Pro remains in limited preview.
xAI flagship is now Grok 4.5. Launched approximately July 9, 2026, Grok 4.5 delivers Opus-class performance per independent coverage. Grok 4.3 (April 30, 2026) is the previous flagship, still available. xAI continues to gate numerical RPM/TPM behind the console. Tables updated to reference Grok 4.5.
DeepSeek V3.2 aliases are gone. The July 24 deprecation deadline passed. Any pipeline still referencing a V3.2 alias is now receiving errors. V4-Flash and V4-Pro are the current endpoints. DeepSeek V4-Flash 0731 is a new fast/efficient variant that reached within one point of GPT-5.6 Luna on the Artificial Analysis Intelligence Index at approximately 60% lower cost per task.
Alibaba’s current flagship is Qwen3.8-Max. Announced August 3, 2026. 2.4T parameters, 1M context, ranks 5th on Text Arena and 2nd on Vision Arena. Available via API on Alibaba Cloud Model Studio; weights are scheduled for release. Published rate limit figures for Qwen3.8-Max are not yet available — Qwen3.7-Max figures remain as a provisional floor where shown.
Kimi K3 is now the largest open-weight model released. 2.8T parameters, 1M context, matches Opus-class performance at Sonnet-level pricing. Demand was high enough that Moonshot paused new subscriptions within 48 hours of launch. Rate limit figures for K3 are not yet confirmed in Moonshot’s public documentation — K2.5 figures remain in tables as a provisional floor.
GPT-5.3 Codex is no longer a standalone API endpoint. Consolidated into ChatGPT. Removed from entry paid and scaled tables.
Anthropic Sonnet 5 introductory pricing expires August 31. Three weeks out. Pricing shifts from $2/$10 to $3/$15 per M input/output on September 1. TPM budget planning should account for the price increase now if cost modeling is tied to per-token rates.
Still pending formal documentation:
- OpenAI: GPT-5.6 Sol rate limit tier tables not yet published. GPT-5.5 figures carried forward as a provisional floor.
- Anthropic: Claude Opus 5 rate limit figures not yet published. Claude Fable 5 rate limit figures not yet published. Fable 5 confirmed as a permanent tier July 18 but absent from rate limit tables until Anthropic publishes figures.
- Google: Gemini 3.6 Flash rate limit figures not yet published. Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.5 Flash Cyber, and Gemini Omni figures also absent. Gemini 3.5 Pro remains in limited preview.
- Alibaba: Qwen3.8-Max (August 3) rate limit tables not yet published. Qwen3.7-Max limits carried forward as a provisional floor.
- Moonshot: Kimi K3 rate limit figures not yet confirmed. K2.5 figures carried forward as a provisional floor.
Deprecation calendar:
- DeepSeek V3.2 aliases deprecated July 24, 2026 — passed. V4-Flash and V4-Pro are the current endpoints. If you haven’t migrated, your pipeline is broken.
- OpenAI Assistants API deprecates August 2026 — this month. Replaced by the Responses API. Migrate now.
- Anthropic Sonnet 5 introductory pricing ($2/$10 per M) expires August 31, 2026 — three weeks out. Standard pricing ($3/$15) applies from September 1.
What are rate limits?
Rate limits cap how much you can use an API within a given time window. Providers enforce them to manage capacity and ensure fair access. Exceeding a limit returns a 429 Too Many Requests error.
Most providers use the token bucket algorithm: capacity refills continuously up to your maximum, rather than resetting at fixed intervals. A 60 RPM limit means roughly 1 request per second with steady refill, not 60 requests then a hard stop.
The metrics
- RPM (Requests per minute): API calls per minute, regardless of size. A 10-token request and a 100K-token request both count as one.
- RPD (Requests per day): Daily cap on total API calls. Some providers use this instead of (or alongside) RPM, especially on free tiers.
- TPM (Tokens per minute): Total tokens (input + output) processed per minute. Usually the binding constraint in production.
- TPD (Tokens per day): Daily token cap. More common on free tiers.
- ITPM/OTPM (Input/Output tokens per minute): Anthropic separates input and output token limits, giving finer control. Cached input tokens don’t count toward ITPM on current Claude 4.x and 5.x models.
- ASH/ASD (Audio seconds per hour/day): For speech models like Whisper.
RPM limits set your concurrency ceiling. TPM limits set your throughput. For batch workloads, RPD and TPD matter more. For real-time apps, RPM and TPM hit first.
Free Tier — Requests per Minute (RPM) & Requests per Day (RPD)
| Provider / Model | RPM | RPD |
|---|---|---|
| Google — Gemini 3.1 Flash-Lite | 15 | 1,000 |
| Groq — Llama 4 Scout | 30 | 1,000 |
| Cerebras — Llama 4 Scout | 30 | — |
| SambaNova — open models | 10–30 | — |
| Fireworks AI — OS models | 10 | — |
| Cohere — Command R+* | 20 | ~33† |
*Cohere’s free tier figures reference Command R+. Command A+ free tier limits not yet published. †Cohere trial keys get 1,000 calls/month across all endpoints, roughly ~33/day.
Note: Gemini 3.5 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber, and 3.6 Flash free tier rate limit figures are not yet published in public documentation. Gemini 3.1 Flash-Lite figures shown above are the last confirmed published numbers. OpenAI, Anthropic, xAI, DeepSeek, and Mistral have no free API tier for current flagship models. OpenAI’s free tier does not support any GPT-5.x model via the API. Anthropic’s lowest entry is $5 (Tier 1). xAI and Mistral gate numerical limits behind their respective consoles.
Migration note: Gemini 2.0 Flash and Flash-Lite shut down June 1, 2026. Those endpoints no longer accept requests. DeepSeek V3.2 aliases deprecated July 24, 2026 — those endpoints are gone.
Free Tier — Tokens per Minute (TPM) & Tokens per Day (TPD)
| Provider / Model | TPM | TPD |
|---|---|---|
| Google — Gemini 3.1 Flash-Lite | 250,000 | — |
| Groq — Llama 4 Scout | 30,000 | 500,000 |
| Groq — other open models | 6,000–30,000 | 500,000 |
| Cerebras — Llama 4 Scout | 60,000 | 1,000,000 |
| SambaNova (free tier) | — | — |
Standout: Google leads with 250K TPM on the free tier for Gemini 3.1 Flash-Lite, but RPM/RPD caps (10–15 RPM) mean it’s best suited to fewer, larger requests. Gemini 3.6 Flash, 3.5 Flash, and 3.5 Flash-Lite free tier TPM figures are not yet published. Cerebras offers 60K TPM with 1M TPD and approximately 2,100 tokens/second inference speed on Llama 4 Scout — the fastest free option by throughput. SambaNova publishes RPM but not TPM/TPD specifics.
Entry Paid Tier — Requests per Minute (RPM) & Requests per Day (RPD)
The tier most developers start with. OpenAI Tier 1 ($5), Anthropic Tier 1 ($5), Google Tier 1 (pay-as-you-go), and others at their lowest paid level.
| Provider / Model | RPM | RPD |
|---|---|---|
| OpenAI — GPT-5.6 Sol* | 500* | — |
| OpenAI — GPT-4.1 Nano | 500 | — |
| Anthropic — Opus 5† | 50† | — |
| Anthropic — Sonnet 5 | 50 | — |
| Anthropic — Haiku 4.5 | 50 | — |
| Google — Gemini 3.1 Pro | 150 | 1,000 |
| Google — Gemini 3.1 Flash-Lite | 300 | 1,500 |
| xAI — Grok 4.5 | Console‡ | — |
| DeepSeek — V4-Flash | Dynamic§ | — |
| DeepSeek — V4-Flash 0731 | Dynamic§ | — |
| Fireworks AI — OS models | ≤6,000 | — |
| Cohere — Command A+** | 500 | — |
*GPT-5.6 Sol rate limit tier tables not yet published. 500 RPM carried forward from GPT-5.5 as a provisional floor. †Opus 5 rate limit tier tables not yet published. 50 RPM carried forward from Opus 4.8 as a provisional floor. ‡xAI publishes tier thresholds ($0/$50/$250/$1K/$5K) but numerical RPM/TPM are only visible in the xAI Console after login. §DeepSeek uses fully dynamic concurrency limits based on server load. No fixed RPM/TPM published. **Command A+ entry paid tier limits not yet published. 500 RPM is the Command R+ figure, carried forward as a reference.
Gemini 3.6 Flash note: Gemini 3.6 Flash is Google’s current flagship but has no published rate limit figures in public documentation. Gemini 3.5 Flash (GA) is the previous flagship — also without published rate limit figures. Gemini 3.1 Pro and 3.1 Flash-Lite remain available as the previous generation, with the last confirmed published figures shown above. Check AI Studio for your project’s actual limits.
Mistral note: Mistral moved to console-only limits. Actual numbers require login at admin.mistral.ai/plateforme/limits. Mistral is excluded from RPM columns where no public figure exists.
Standout: Fireworks can spike to 6,000 RPM but it’s a dynamic ceiling, not guaranteed (soft limit starts at approximately 1 RPS and doubles hourly). Anthropic’s 50 RPM at Tier 1 is the lowest here, but jumps 20x to 1,000 RPM at Tier 2 ($40). Sonnet 5’s tokenizer emits approximately 30% more tokens per request than Sonnet 4.6 for identical prompts — TPM headroom shrinks faster than it did on 4.6 at the same RPM. Opus 5 tokenizer behavior relative to Opus 4.8 is not yet formally documented.
Entry Paid Tier — Tokens per Minute (TPM)
| Provider / Model | TPM |
|---|---|
| OpenAI — GPT-5.6 Sol* | 500K* |
| OpenAI — GPT-4.1 Nano | 200K |
| Anthropic — Opus 5† | 30K in / 8K out† |
| Anthropic — Sonnet 5 | 30K in / 8K out |
| Anthropic — Haiku 4.5 | 50K in / 10K out |
| Google — Gemini 3.1 Pro | 1M |
| Google — Gemini 3.1 Flash-Lite | 2M |
| DeepSeek — V4-Flash | Dynamic |
*GPT-5.6 Sol TPM figures not yet published. 500K carried forward from GPT-5.5 as a provisional floor. †Opus 5 TPM figures not yet published. Opus 4.8 figures carried forward as a provisional floor.
Standout: Google’s 3.1 Flash-Lite leads confirmed figures at 2M TPM. Anthropic’s Sonnet 5 shows 30K ITPM — the tokenizer change means identical prompts now consume more tokens against that limit than they did on Sonnet 4.6. Cached input tokens don’t count toward ITPM on Claude 4.x and 5.x models — with an 80% cache hit rate, effective throughput remains substantially higher than the raw ITPM figure.
Scaled Tier — RPM & TPM at Higher Spend
For teams past entry-level. OpenAI Tier 3 ($100+), Anthropic Tier 3 ($200+), Google Tier 2 ($250+).
| Provider (Tier) / Model | RPM | TPM |
|---|---|---|
| OpenAI (Tier 3) — GPT-5.6 Sol* | 5,000* | 2M* |
| OpenAI (Tier 3) — GPT-4.1 Nano | 5,000 | 4M |
| Anthropic (Tier 3) — Opus 5† | 2,000† | 800K in / 160K out† |
| Anthropic (Tier 3) — Sonnet 5 | 2,000 | 800K in / 160K out |
| Anthropic (Tier 3) — Haiku 4.5 | 2,000 | 1M in / 200K out |
| Google (Tier 2) — Gemini 3.1 Pro | 1,000 | 2M |
| Google (Tier 2) — Gemini 3.1 Flash-Lite | 2,000 | 4M |
*GPT-5.6 Sol Tier 3 figures not yet published. GPT-5.5 figures carried forward as a provisional floor. †Opus 5 Tier 3 figures not yet published. Opus 4.8 figures carried forward as a provisional floor.
At the highest standard tiers (OpenAI Tier 5 at $1,000+, Anthropic Tier 4 at $400+):
| Provider (Tier) / Model | RPM | TPM |
|---|---|---|
| OpenAI (Tier 5) — GPT-5.6 Sol* | 15,000* | 40M* |
| OpenAI (Tier 5) — GPT-4.1 Nano | 30,000 | 180M |
| Anthropic (Tier 4) — Opus 5† | 4,000† | 2M in / 400K out† |
| Anthropic (Tier 4) — Sonnet 5 | 4,000 | 2M in / 400K out |
| Anthropic (Tier 4) — Haiku 4.5 | 4,000 | 4M in / 800K out |
*GPT-5.6 Sol Tier 5 figures not yet published. GPT-5.5 figures carried forward as a provisional floor. †Opus 5 Tier 4 figures not yet published. Opus 4.8 figures carried forward as a provisional floor. Confirm current limits in the Anthropic console.
Standout: OpenAI’s Tier 5 numbers remain the highest of any provider with published or provisionally carried figures. GPT-4.1 Nano at 180M TPM and 30K RPM is built for high-volume classification and routing. Anthropic’s Tier 4 caps at 4K RPM, but cached tokens don’t count toward ITPM — effective throughput can be substantially higher with good cache hit rates.
Model pool notes:
- Opus 5, Opus 4.8, 4.7, and 4.6 pool assignment is not yet formally documented for Opus 5. Opus 4.8, 4.7, and 4.6 shared one rate limit bucket; check your Anthropic dashboard for how Opus 5 is bucketed.
- Sonnet 5 and Sonnet 4.6 likely share a pool — not formally confirmed; check your dashboard.
- Sending traffic to multiple model versions within the same pool draws from one shared bucket, not separate ones.
Gemini 3.6 Flash scaled tier note: Gemini 3.6 Flash is Google’s current flagship with no published rate limit figures. Gemini 3.5 Flash (GA) is also without published rate limit figures. The 3.1 Flash-Lite figures shown above remain the last confirmed published numbers.
Cerebras & SambaNova
Both specialize in custom silicon for inference speed.
Cerebras
Hardware: Wafer-Scale Engine (WSE-3). Fastest published inference: approximately 2,100 tokens/second on Llama 4 Scout.
| Tier | RPM | TPM | TPD |
|---|---|---|---|
| Free | 30 | 60,000 | 1,000,000 |
| Paid | Higher (contact sales) | Higher | Higher |
Hosts Llama 4 Scout and other open models. The speed advantage is real: tasks that take 30 seconds on GPU-based providers finish in under 5 seconds on the WSE-3. Paid tier limits are not publicly documented.
SambaNova
Hardware: Custom RDU (Reconfigurable Dataflow Unit). Best time-to-first-token (TTFT): approximately 0.2 seconds.
| Tier | RPM | Notes |
|---|---|---|
| Free | 10–30 (varies by model) | Hosts up to Llama 4 Maverick for free |
SambaNova offers Llama 4 Maverick (1M context) on the free tier — most other free-tier providers cap out at smaller models. TPM and TPD limits are not publicly documented.
Audio Models — ASH & ASD
Groq and Fireworks both publish audio-specific limits.
| Provider / Model | RPM | RPD | ASH | ASD | Audio min/min |
|---|---|---|---|---|---|
| Groq (Free) — Whisper Large v3 | 20 | 2,000 | 7,200 | 28,800 | — |
| Groq (Free) — Whisper Large v3 Turbo | 20 | 2,000 | 7,200 | 28,800 | — |
| Fireworks AI — Whisper v3-large | — | — | — | — | 200 |
| Fireworks AI — Whisper v3-turbo | — | — | — | — | 400 |
Groq: 7,200 ASH = 2 hours of audio per hour of wall time. 28,800 ASD = 8 hours per day. Adequate for a podcast transcription pipeline on the free tier.
Fireworks: 200 min/min for Whisper v3-large, 400 min/min for v3-turbo. Concurrent streaming capped at 10 connections.
OpenAI launched GPT Transcribe and GPT Live Transcribe via API since the last edition. Independent benchmarks show both lag ElevenLabs, Google, and Mistral on error rates. OpenAI does not yet publish audio-specific ASH/ASD figures for these endpoints in its public rate limit documentation.
Cloud Aggregators — Azure AI & AWS Bedrock
Both use configurable per-deployment limits, not fixed tiers. The shared ratio is 6 RPM per 1,000 TPM.
Azure AI (Microsoft)
| Deployment Type | How Limits Work |
|---|---|
| Pay-as-you-go (Standard) | TPM quota per model per region, auto-scales with usage |
| Provisioned (PTU) | Reserved throughput units, no per-request limits |
| Global/Data Zone | Higher default quotas, multi-region routing |
GPT-5.6 Sol is now the default in Microsoft 365 Copilot. Azure AI hosts GPT-5.6 Sol, GPT-5.5, GPT-5.4, Claude Opus 5/Sonnet 5, and Gemini 3.1 Pro (verify availability per region — not all models are deployed globally; Gemini 3.6 Flash and 3.5 Flash availability on Azure is not yet confirmed). Azure AI Quotas & Limits
AWS Bedrock
Hosts Claude (Opus 5, Opus 4.8, Sonnet 5, Haiku 4.5), Llama 4, Mistral, and Nova models with per-model, per-region quotas. Claude Fable 5 is being restored on AWS following the July 1 global redeployment and July 18 permanent-tier confirmation — check the Bedrock model catalog for current status. Amazon is scaling back most Nova AI models (Premier, Omni, Reel, Canvas into “keep the lights on” mode) while a new Frontier research team works on a foundation model expected at re:Invent. Default quotas vary by model and can be increased via AWS Service Quotas console. Provisioned Throughput deployments remove rate limits entirely. AWS Bedrock Quotas
More Providers — Perplexity, Alibaba (Qwen), Moonshot (Kimi)
| Provider / Model | RPM (T0) | RPM (T1) | RPM (T3+) | TPM | TPD (T0) |
|---|---|---|---|---|---|
| Perplexity — Sonar Pro | 50 | 150 | 1,000 | — | — |
| Perplexity — Sonar | 50 | 150 | 1,000 | — | — |
| Perplexity — Deep Research | 5 | 10 | 40 | — | — |
| Alibaba (Qwen) — Qwen3.8-Max* | — | 600* | 600* | 1M* | — |
| Alibaba (Qwen) — Qwen3.7-Max† | — | 600† | 600† | 1M† | — |
| Alibaba (Qwen) — Qwen3.5 Plus | — | 15,000 | 30,000 | 5M | — |
| Alibaba (Qwen) — Qwen3.5 Flash | — | 15,000 | 30,000 | 10M | — |
| Moonshot — Kimi K3‡ | 3‡ | 200‡ | 5,000‡ | — | 1.5M‡ |
| Moonshot — Kimi K2.5 | 3 | 200 | 5,000 | — | 1.5M |
*Qwen3.8-Max (August 3) rate limit figures not yet published. Qwen3.7-Max figures carried forward as a provisional floor. †Qwen3.7-Max (May 20) rate limit figures not yet formally published. Qwen3 Max limits carried forward as a provisional floor. ‡Kimi K3 rate limit figures not yet confirmed in Moonshot’s public documentation. K2.5 figures carried forward as a provisional floor.
Perplexity tiers: T0 = new account, T1 = $50+, T3 = $500+. Alibaba limits shown for international (Singapore) deployment; Beijing deployment is higher (Qwen3.5 Flash at 30K RPM — Qwen3.8-Max and 3.7-Max Beijing limits not yet confirmed). Moonshot T1 = $10+ cumulative recharge; T0 has 1.5M TPD cap, T1+ is unlimited.
Standout: Alibaba’s Qwen3.5 Flash at 30K RPM / 10M TPM (Beijing) remains the highest confirmed throughput of any provider in this post. Moonshot’s T5 tier (10K RPM, 1,000 concurrent connections, $3,000+ recharge) is competitive at the high end. Kimi K3’s 2.8T parameters and Opus-class performance at Sonnet-level pricing make it a meaningful entrant for teams running high-volume open-weight workloads — but wait for confirmed rate limit figures before sizing pipelines.
Provider Notes
- OpenAI: GPT-5.6 Sol is the current flagship, broadly available as of July 9–11, 2026. Variants Sol, Terra, and Luna are rolling out in tiers. Pricing cut up to 80% relative to GPT-5.5 through recursive self-improvement and distillation. GPT-5.6 Sol rate limit tier tables not yet published in public docs — GPT-5.5 figures carried forward as a provisional floor in all tables above. GPT-5.5 remains available (variants: GPT-5.5 Thinking, GPT-5.5 Pro). GPT-5.5 Instant remains available. GPT-5.4 remains available (variants: GPT-5.4 Thinking, GPT-5.4 Pro, GPT-5.4-Cyber for vetted security teams). GPT-5.3 Codex consolidated into ChatGPT — no longer a standalone API endpoint. GPT-4.1 Nano is the budget/high-throughput model. GPT-5.2 being phased out. GPT-4 and GPT-4o retired February 2026. GPT Transcribe and GPT Live Transcribe now available via API; audio-specific rate limit figures not yet published. Assistants API deprecates August 2026 — this month; replaced by the Responses API. Batch API offers 50% discount with higher queue limits. Docs
- Anthropic: Claude Opus 5 (July 2026) is the current flagship — scored 30.2% on ARC-AGI-3, nearly 4x the previous record; matches Fable 5-level performance on coding and knowledge tasks at 50% lower token cost. Rate limit figures for Opus 5 not yet published; Opus 4.8 figures carried forward as a provisional floor. Opus 4.8 (May 2026) is the previous flagship, still available. Claude Sonnet 5 (June 30, 2026) is the current balanced tier — tokenizer emits approximately 30% more tokens per request than Sonnet 4.6 for identical text. Introductory pricing ($2/$10 per M) runs through August 31, 2026; standard $3/$15 applies from September 1. Claude Fable 5 (restored July 1, confirmed permanent July 18) is now a committed production tier; rate limit figures not yet published. On redeployment, Anthropic added a classifier that blocks the reported exploit technique and reroutes flagged requests to Opus 5. Claude Mythos 5 (defensive-cybersecurity focus) had access restored June 26 for vetted US organizations via Project Glasswing — not generally available. Opus 5 pool assignment not yet formally documented; Opus 4.8, 4.7, and 4.6 shared a combined pool — check your Anthropic dashboard for current bucketing. Sonnet 5 and Sonnet 4.6 likely share a pool — not yet formally documented. Cached input tokens don’t count toward ITPM on current 4.x and 5.x models. No free tier; lowest entry is $5. Claude 3 Haiku retired April 2026. Docs
- Google: Static rate limit tables removed from public docs in Q1 2026. Actual limits only visible in AI Studio dashboard (
aistudio.google.com/rate-limit). Gemini 3.6 Flash (July 2026) is the current flagship — listed first in official Gemini API model docs; Gemini API Managed Agents support it with production-grade hooks. Rate limit figures not yet published. Gemini 3.5 Flash (GA since May 19, Google I/O 2026) is the previous flagship — approximately 4x faster output than other frontier models, leads on agentic and coding benchmarks, includes computer use capability; rate limit figures not yet published. Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber (cybersecurity domain) launched July 2026; rate limit figures not yet published. Gemini 3.5 Pro remains in limited preview (2M context, Deep Think). Gemini 3.1 Pro remains available as the previous-generation Pro model. Gemini Omni (multimodal) rate limit figures also absent. Gemini 2.5 Pro/Flash are previous-generation models no longer recommended for new deployments. Gemini 2.0 Flash/Flash-Lite shut down June 1, 2026. Google is planning Gemini 4 and has raised CapEx to $205B. Docs - Groq: Free tier publishes exact numbers (30 RPM, 6K–30K TPM depending on model). Daily token budgets (TPD) alongside TPM. Cached tokens don’t count. Hosts Llama 4 Scout and other open models on free tier. Runs on custom LPU hardware. Docs
- xAI: Current flagship is Grok 4.5 (approximately July 9, 2026; Opus-class performance per independent coverage). Grok 4.3 (April 30, 2026; 1M context; $1.25/$2.50 per M input/output) is the previous flagship, still available. grok-build-0.1 is the agentic coding model. 5-tier structure ($0/$50/$250/$1K/$5K thresholds) but numerical RPM/TPM only visible in xAI Console. Docs
- Mistral: European data residency (GDPR). 5-tier structure (Free through Tier 4 at $500+). Limits enforced per RPS (not RPM), TPM, and tokens/month. Actual numbers require login at admin.mistral.ai. Mistral Large 3 (675B MoE) and Mistral Small 4 are available. Devstral 2 (72.2% SWE-bench) is available for coding workloads. Docs
- DeepSeek: Fully dynamic concurrency limits based on server load. No published RPM/TPM. V4-Flash 0731 is a new fast/efficient variant that reached within one point of GPT-5.6 Luna on Artificial Analysis Intelligence Index at approximately 60% lower cost per task. V4-Flash and V4-Pro are the current endpoints. V3.2 aliases deprecated July 24, 2026 — those endpoints are gone. V4-Flash pricing: $0.14/M input (cache miss), $0.028/M (cache hit), $0.28/M output. Docs
- Cerebras: Fastest inference (approximately 2,100 tok/s on Llama 4 Scout). Free tier: 30 RPM, 60K TPM, 1M TPD. Paid tier limits not publicly documented. Custom WSE-3 silicon. Docs
- SambaNova: Best TTFT (approximately 0.2s). Free tier: 10–30 RPM depending on model size. Hosts up to Llama 4 Maverick for free. TPM and TPD limits not publicly documented. Custom RDU hardware. Docs
- Together AI: Dynamic rate limits since January 2026. No fixed tiers or published numbers. Limits grow with sustained usage and are returned in API response headers. Docs
- Fireworks AI: Dynamic ceiling up to 6,000 RPM (soft limit starts approximately 1 RPS, doubles hourly). On-demand GPU deployments remove limits. Spending tier caps: $50–$50K/month by tier. Docs
- Cohere: Command A+ (May 20, Apache 2.0, 218B MoE / 25B active; runs on 2x H100 or 1x B200; native citations) is the flagship. Self-hosting is viable on 2x H100 or 1x B200 and removes API rate limits entirely for teams with the hardware. API rate limit figures for Command A+ not yet published — 500 RPM (production chat) shown in tables is the Command R+ figure. Trial keys: 1,000 calls/month. Docs
- Perplexity: 6 tiers (T0–T5) based on cumulative spend. Leaky bucket algorithm. Deep Research model has very low limits (5–100 RPM). Agent API: 50–2,000 RPM by tier. Docs
- Alibaba (Qwen): Qwen3.8-Max (August 3, 2026; 2.4T parameters; 1M context; ranks 5th on Text Arena, 2nd on Vision Arena) is the current flagship; published rate limits not yet confirmed. Qwen3.7-Max (May 20; 1M context; $2.50/$7.50 per M) is the previous flagship. Open-weight family is Qwen3.6 (Apache 2.0 — e.g., Qwen3.6-35B-A3B, Qwen3.6-27B). Region-specific limits apply — Beijing deployment is more generous than Singapore/Global. Qwen3.5 Flash at 30K RPM / 10M TPM (Beijing) is the highest confirmed throughput of any provider listed. Docs
- Moonshot (Kimi): 6 tiers (T0–T5). T0 ($1 recharge): 3 RPM, 1 concurrent, 1.5M TPD. T5 ($3,000): 10K RPM, 1,000 concurrent, 5M TPM, unlimited TPD. Kimi K3 (2.8T parameters, 1M context, Opus-class performance at Sonnet-level pricing) is now available — rate limit figures not yet confirmed in public documentation; K2.5 figures carried forward as a provisional floor. Demand for K3 was high enough that Moonshot paused new subscriptions within 48 hours of launch. Automatic 75% caching discount applied without opt-in. Docs
- AWS Bedrock: Hosts Claude (Opus 5, Opus 4.8, Sonnet 5, Haiku 4.5), Llama 4, Mistral with per-model per-region quotas. Claude Fable 5 restoration on Bedrock is in progress following the July 18 permanent-tier confirmation — check the Bedrock model catalog. Amazon scaling back most Nova AI models (Premier, Omni, Reel, Canvas) while a new Frontier team works on a foundation model expected at re:Invent. 6 RPM per 1K TPM ratio. Provisioned Throughput removes limits. Docs
- Azure AI: Hosts GPT-5.6 Sol (default in Microsoft 365 Copilot), GPT-5.5, GPT-5.4, Claude Opus 5/Sonnet 5, and Gemini 3.1 Pro across regions (verify model and region availability before planning deployments — not all models are deployed globally; Gemini 3.6 Flash availability on Azure not yet confirmed). PTU deployments remove per-request limits. Docs
- NVIDIA NIM: Hosted API is for prototyping (40 RPM). Self-hosted NIM containers have no limits. Docs
Tips for managing rate limits
For a deep dive on token pricing, cost optimization strategies, caching mechanics, model routing architectures, and production cost modeling, see LLM Token Costs and Efficiency. This section covers rate limit management specifically.
Handling 429s
Use exponential backoff when you hit 429s. The Anthropic SDK, OpenAI SDK, and most third-party clients handle this automatically. Don’t build your own retry loop unless you need custom jitter or circuit-breaking behavior.
Check response headers before you hit the wall. Together AI and Fireworks return your current rate limit state in every API response. DeepSeek adjusts concurrency limits dynamically based on server load. Reading these headers in production enables proactive throttling rather than reactive retries.
GPT-5.6 Sol: treat published limits as provisional
GPT-5.6 Sol is now broadly available, but its rate limit tier tables are not yet published. The figures in this post are carried forward from GPT-5.5. Sol’s pricing is cut by up to 80% relative to GPT-5.5, which may or may not be reflected in rate limit changes when official figures publish. Check platform.openai.com/account/limits for your account’s actual current limits.
Sonnet 5 tokenizer: revise your throughput math
Claude Sonnet 5’s tokenizer emits approximately 30% more tokens per request than Sonnet 4.6 for identical prompts. A workload that consumed 20K ITPM on Sonnet 4.6 will consume approximately 26K ITPM on Sonnet 5 with the same inputs. Recheck your limits against actual Sonnet 5 token counts, particularly at Tier 1 where ITPM is already constrained at 30K. Opus 5 tokenizer behavior relative to Opus 4.8 is not yet formally documented — treat Opus 5 throughput calculations as provisional until Anthropic publishes figures.
Sonnet 5 introductory pricing expires August 31
Three weeks out. If your cost modeling is tied to the $2/$10 introductory rate, revise for $3/$15 from September 1. The rate limit structure does not change at that date — only pricing.
Caching for throughput (not just cost)
Prompt caching has a rate limit benefit separate from its cost benefit. The cost savings are covered in the token costs post. The rate limit implications:
- Anthropic: Cached input tokens don’t count toward ITPM limits at all on Claude 4.x and 5.x models. With 80% cache hit rate, a 2M ITPM limit effectively handles approximately 10M total input tokens per minute. This is a throughput multiplier, not just a cost reduction. Given Sonnet 5’s higher per-token counts, caching is more important for staying within ITPM limits than it was on Sonnet 4.6.
- Groq: Cached tokens don’t count toward TPM/TPD limits on the free tier, effectively expanding your daily token budget.
- Moonshot (Kimi): Automatic 75% caching discount applied without opt-in. No cache management required.
DeepSeek V3.2 is gone
The July 24 deprecation deadline passed. If anything in your pipeline is still referencing a V3.2 alias, it is receiving errors now, not rate limit responses. V4-Flash and V4-Flash 0731 are the current fast endpoints; V4-Pro for higher capability. Audit your model endpoint references if you haven’t already.
OpenAI Assistants API deprecates this month
The Assistants API deprecates August 2026. The Responses API is the replacement. Migration guidance is in OpenAI’s migration docs.
Batch APIs: higher queue limits
OpenAI, Anthropic, and Google all offer batch endpoints with queue limits that far exceed real-time TPM. OpenAI’s GPT-5.6 Sol (provisional) inherits GPT-5.5’s 1.5M batch queue tokens at Tier 1 vs. 500K real-time TPM. For workloads that don’t need sub-second responses, batch endpoints let you move more tokens through tighter rate limits. See batch API cost savings for pricing details.
Model routing for rate limit distribution
Routing requests to different models by complexity distributes rate limit pressure across multiple buckets instead of concentrating it on your most constrained tier. A three-tier architecture (budget for triage, mid-tier for generation, flagship for reasoning) also cuts costs 60–85%. Full routing architecture with cost modeling is in the token costs post.
GPT-4.1 Nano (30K RPM / 180M TPM at Tier 5) and DeepSeek V4-Flash 0731 (dynamic, approximately 60% lower cost than GPT-5.6 Luna per task) are strong candidates for the triage and generation layers of a routing stack. Kimi K3 (Opus-class at Sonnet-level pricing) is worth evaluating as a reasoning layer once rate limit figures are confirmed.
Custom silicon for fewer concurrent connections
Cerebras (approximately 2,100 tok/s) and SambaNova (approximately 0.2s TTFT) are faster than GPU-based providers by a large margin. A task that requires 10 parallel GPU-based API calls to meet a latency SLA might need 2–3 calls on Cerebras. Fewer concurrent calls means less rate limit pressure per unit of output.
Self-hosting as an escape hatch
Cohere’s Command A+ (Apache 2.0, 218B MoE / 25B active) runs on 2x H100 or a single B200. Kimi K3 (2.8T parameters, Apache 2.0 per Moonshot’s licensing) and Thinky Inkling (975B, Apache 2.0) are also open-weight options, though hardware requirements at that scale are substantially higher. For teams already operating that hardware, self-hosting removes API rate limits entirely. The economics only work at scale — hardware cost needs to be weighed against API spend — but the viable hardware tier for self-hosting has expanded as model weights become more accessible.
Watch for shared pools and hidden constraints
- Anthropic model pools: Opus 5 pool assignment is not yet formally documented — check your dashboard. Opus 4.8, 4.7, and 4.6 shared a combined pool. Sonnet 5 and Sonnet 4.6 likely share another pool — not yet formally confirmed. Sending traffic to multiple model versions within the same pool draws from one shared bucket.
- Google’s invisible limits: With static rate limit tables removed from public docs, check AI Studio (
aistudio.google.com/rate-limit) to see your actual per-project limits. Gemini 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber all have no published rate limit figures. The 3.1 Flash-Lite and 3.1 Pro figures in this post are the last confirmed published numbers and may be stale for current deployments. - DeepSeek’s dynamic ceiling: No published RPM/TPM means no guaranteed minimum. Under heavy load, effective concurrency can drop without warning. V3.2 aliases are gone — if you haven’t migrated, you’re receiving errors. Build fallback routing to a second provider if DeepSeek availability is critical to your pipeline.
- New model lag: Qwen3.8-Max (August 3), Kimi K3, Claude Opus 5, GPT-5.6 Sol, and Gemini 3.6 Flash all launched without formal rate limit tables. Figures in this post for those models are either provisional floors based on predecessors or absent entirely. Plan around confirmed figures where possible, and monitor provider dashboards for when official limits publish.
Further Reading
- LLM Token Costs and Efficiency — per-token pricing, caching mechanics, model routing, batch discounts, production cost modeling
- OpenAI rate limits — tiers, usage tracking, batch API, Responses API migration
- Anthropic rate limits — build tier system, prompt caching behavior
- Google Gemini rate limits — free vs paid, AI Studio dashboard
- Groq rate limits — real-time dashboard, daily token caps
- xAI Grok API — rate limits and pricing
- DeepSeek API — dynamic limits, V4 pricing, V3.2 deprecation
- Mistral rate limits — tier structure, RPS enforcement
- Cerebras API — WSE-3 inference, free tier
- SambaNova API — RDU inference, free tier
- Perplexity rate limits — search-augmented API tiers
- Alibaba Qwen rate limits — region-specific limits, Qwen3.8-Max
- Moonshot Kimi rate limits — tier structure, caching, K3 availability
- AWS Bedrock quotas — per-model per-region, Fable 5 status
- Azure AI quotas — PTU capacity, global deployment, GPT-5.6 Sol availability
- Cohere rate limits — production vs trial, Command A+ self-hosting
- NVIDIA NIM — hosted prototyping limits, self-hosted containers