The LLM Encyclopedia, July 18, 2026
The most comprehensive reference for every major AI language model. 60+ models, 22 use cases, full pricing tables — updated weekly.
The Developer’s Complete LLM Comparison Guide (July 18, 2026)
Every Major, Minor, Niche, Open-Source, and Specialized Language Model — Researched, Compared, and Rated
Accuracy note: All model versions, release dates, pricing, and benchmark data reflect publicly confirmed information as of July 18, 2026. This is a fast-moving field — verify pricing and availability against official provider docs before production deployment.
Table of Contents
- What Is an LLM? A Developer’s Primer
- How to Read This Guide
- Tier 1 — Flagship Proprietary Models
- GPT Series (OpenAI)
- Claude Series (Anthropic)
- Gemini Series (Google DeepMind)
- Grok Series (xAI)
- Tier 2 — Strong Proprietary Challengers
- Perplexity (Sonar)
- Microsoft Copilot / Azure OpenAI
- Cohere Command A+ / Command R+
- Amazon Nova / Titan
- Tier 3 — Open-Source Powerhouses
- Meta Llama Series
- Mistral / Mixtral Series
- DeepSeek Series
- Qwen Series (Alibaba)
- Gemma Series (Google)
- IBM Granite
- Falcon Series (TII)
- Microsoft Phi Series
- BLOOM (BigScience)
- OLMo (Allen Institute)
- NVIDIA Nemotron
- Tier 4 — Chinese Frontier Models
- Baidu ERNIE
- Zhipu GLM-5 / GLM-5.1 / ChatGLM
- Moonshot Kimi
- Baichuan
- Yi (01.AI)
- MiniMax
- Hunyuan (Tencent)
- InternLM (Shanghai AI Lab)
- ByteDance Seed
- Tier 5 — Coding-Specialist Models
- GitHub Copilot
- DeepSeek Coder / Prover
- StarCoder / StarCoder2
- CodeLlama
- Codestral / Devstral (Mistral)
- WizardCoder
- Qwen Coder
- Amazon Q Developer
- Tabnine
- Tier 6 — Domain-Specific Models
- Healthcare: Med-PaLM 2, MedLLaMA, BioMedLM, ClinicalBERT
- Finance: BloombergGPT, FinGPT
- Legal: Harvey AI, CoCounsel, ChatLAW
- Science: Galactica, SciGLM
- Cybersecurity
- Tier 7 — Edge / On-Device / Small Models
- Tier 8 — Research & Historical Models
- Pricing Comparison Table (July 18, 2026)
- Benchmark Comparison (July 18, 2026)
- Choosing the Right LLM: Decision Framework
- Real-World Enterprise Success Stories
- Trends and What’s Coming in 2026–2027
- Use Case Directory — Which Model for Which Software Task
What’s New This Week
-
Kimi K3 — a 2.8 trillion parameter open-weight model — matches Opus 4.8 performance at Sonnet 5 pricing levels, signaling the end of cheap Chinese AI and reshaping open-ecosystem cost-performance economics: The Kimi K3 release is the largest open-weight model ever released and represents a strategic inflection point for the open AI ecosystem. Earlier generations of Chinese open models competed primarily on price; K3 competes on capability, with benchmark performance nearing GPT-5.6 Sol and Claude Fable 5 at price points closer to Claude Sonnet 5 than to Opus 4.8. For practitioners building on open stacks, this changes the tradeoff calculus: open models are no longer primarily a cost play but increasingly a capability play, and K3’s scale means the gap between open and closed frontier has compressed further. Teams that had dismissed Chinese open models as budget alternatives should re-evaluate K3 as a genuine frontier option — while accounting for the geopolitical and supply-chain risk considerations that apply to all Chinese-origin AI infrastructure. Source: Daily Signal item #3 (July 18, 2026).
-
Claude Fable 5 is being committed to permanent production status, signaling Anthropic’s confidence in this model tier and providing stability for builders who depend on it: Anthropic’s decision to make Fable 5 permanent — rather than maintaining it as a limited or experimental offering — is a meaningful signal for practitioners who had been hesitant to build production dependencies on a model tier that was suspended under US export controls in June and only restored globally on July 1, 2026. The commitment to permanence reduces the platform risk for teams considering Fable 5 as a primary model, though practitioners should note that the 95% SWE-bench score remains developer-reported and has not yet been independently replicated in verified third-party sources. For teams evaluating Fable 5 as a coding or agentic foundation, the permanence announcement removes one of the key risks that had previously warranted caution. Source: Daily Signal item #13 (July 18, 2026).
-
Open-weight models now match frontier closed-model cyber capabilities from four to seven months ago, but safety measures on open models are ineffective — compressing the capability gap while creating an asymmetric risk profile: The capability gap between open and closed models in cybersecurity tasks has narrowed from six to ten months to four to seven months in just six months, and crucially, the safety guardrails on open-weight models are not keeping pace with their capability gains. This asymmetry — frontier-class offensive cyber capability accessible via open weights, without the safety infrastructure that closed-model providers maintain — fundamentally changes the risk calculus for defenders. For practitioners deploying or evaluating open-weight models in environments with elevated security requirements, this finding warrants immediate review of threat models: the assumption that open models lag closed models by a comfortable margin on security-relevant tasks is no longer valid. The finding also has direct implications for teams building security tooling or deploying agents with network or system access. Source: Daily Signal item #2 (July 18, 2026).
-
China announced the World Artificial Intelligence Cooperation Organization (WAICO), 5,000 AI training slots for Global South countries, and cooperation centers with BRICS and the African Union — representing Xi Jinping’s clearest move yet toward a parallel AI governance order: The WAICO announcement is not primarily a model or capability story — it is a governance and infrastructure story with material implications for practitioners building AI systems intended for global deployment. China is systematically constructing an alternative AI governance framework that is decoupled from Western standards bodies, with institutional infrastructure (training programs, cooperation centers, a named international organization) that mirrors the architecture of existing Western-led technology governance. For practitioners navigating AI deployment across jurisdictions, the practical consequence is accelerating regulatory bifurcation: systems, models, and governance frameworks that are acceptable in one framework may be incompatible with the other, and the timeline for that bifurcation to produce concrete procurement and compliance constraints is shortening. Source: Daily Signal item #4 (July 18, 2026).
-
The US Navy is running LLMs directly on warships and the Pentagon has formally adopted an AI-first fleet strategy that explicitly treats slow adoption as a greater risk than imperfect alignment: The Pentagon’s posture is significant not as a capability announcement but as an institutional signal: the world’s largest defense organization has made a documented, explicit policy choice to prioritize AI deployment speed over alignment assurance. For practitioners working in or adjacent to defense, government, or regulated industries, this framing — slow adoption as the primary risk, not misalignment — is now official doctrine at the highest institutional level, and it will shape procurement criteria, vendor requirements, and acceptable risk thresholds across the defense supply chain. For practitioners outside defense, the episode illustrates how institutional incentive structures can override safety caution in deployment decisions, a pattern that is not unique to defense and that any practitioner building AI systems for large organizations should model in their risk assessments. Source: Daily Signal item #5 (July 18, 2026).
-
The Agent2Agent (A2A) v1.0 protocol introduces cryptographic “Agent Card” signing to solve discovery and authentication for networks of autonomous agents that have never previously interacted: As multi-agent systems move from architectural theory into production deployment, the infrastructure problem of how agents establish trust with other agents they have not previously encountered becomes a critical engineering constraint. The A2A v1.0 Agent Card system provides a cryptographic handshake mechanism for trustless interoperability — allowing agents to authenticate each other and discover capabilities without requiring centralized registration or prior coordination. For practitioners building or evaluating multi-agent systems, A2A v1.0 represents a maturing of the interoperability infrastructure that has been absent from most agentic frameworks to date. Teams that are designing agent-to-agent communication architectures should evaluate A2A v1.0 as a potential standard before committing to proprietary trust mechanisms that will be harder to unwind later. Source: Daily Signal item #7 (July 18, 2026).
-
Practical guidance on working effectively with GPT-5.6 is now available, relevant for practitioners upgrading workflows to the newest OpenAI model tier: With GPT-5.6 Sol, Terra, and Luna in preview and accessible to limited partners, concrete practitioner guidance on how to leverage the new capability tier — including agentic and reasoning features specific to the 5.6 family — is now circulating. For teams evaluating whether to migrate workflows from GPT-5.5 or GPT-5.4 to the 5.6 tier, this represents early operational signal on where the new models perform differently from their predecessors and how to structure prompts and pipelines to take advantage of the changes. Source: Daily Signal item #10 (July 18, 2026).
1. What Is an LLM? A Developer’s Primer
A Large Language Model (LLM) is a deep learning system trained on massive corpora of text (and increasingly images, audio, and video) to predict and generate human-like language. Built on the Transformer architecture (Vaswani et al., 2017), LLMs are characterized by billions of parameters — the numerical weights learned during training that encode knowledge about language, facts, and reasoning.
Key concepts every developer needs to know:
- Parameters: The “weights” inside a model. More parameters generally means more capacity, but not always better performance. A 7B model with excellent training data can outperform a 70B model trained poorly.
- Context Window (Tokens): How much text the model can “see” at once. A token ≈ 0.75 words. A 1M context window can process ~750,000 words in one shot.
- Inference: The process of running a trained model to generate output. This is what you pay for when using APIs.
- Fine-tuning: Continuing training a base model on domain-specific data to specialize it.
- RLHF: Reinforcement Learning from Human Feedback — human raters rank outputs to teach the model to be more helpful and less harmful.
- MoE (Mixture of Experts): Architecture where only a subset of parameters (“experts”) activate per token, enabling massive total parameter counts with lower compute cost. Used in Mixtral, DeepSeek V3, Llama 4, Grok, and others.
- RAG (Retrieval-Augmented Generation): Pairing an LLM with a vector database so it can look up external documents before answering — reducing hallucinations.
- Quantization: Compressing model weights (e.g., 32-bit floats → 4-bit integers) to reduce VRAM requirements and increase inference speed with minimal quality loss.
- Extended Thinking / Chain-of-Thought: The model reasons internally before producing an answer, trading latency for accuracy on hard problems. Now standard across frontier models.
- Computer Use: Models that can see a screen, move a cursor, click, and type — enabling truly autonomous agentic workflows. Native in GPT-5.4, Claude 4.x, and Gemini 3.x as of 2026.
2. How to Read This Guide
Each model entry uses a consistent structure:
| Field | Description |
|---|---|
| Released | Date of first public availability |
| Developer | Organization behind the model |
| Type | Proprietary / Open-weight / Open-source |
| Context Window | Maximum token input |
| Strengths | What it genuinely does well |
| Weaknesses | Honest limitations |
| Best For | Ideal use cases and user profiles |
| Constraints | Rate limits, data policies, license restrictions |
| Cost | API pricing per million tokens (input / output), July 2026 |
| Real-World Use | Documented production deployments |
3. Tier 1 — Flagship Proprietary Models
These are the frontier models competing at the highest capability level. They define the industry benchmark each quarter.
🟢 GPT Series — OpenAI
Developer: OpenAI
Type: Proprietary (closed-source)
Headquarters: San Francisco, CA
OpenAI’s GPT family is the most recognized LLM series in the world. The progression: GPT-3 (2020) launched the modern LLM era; ChatGPT (Nov 2022, GPT-3.5) made it a consumer phenomenon; GPT-4 (2023) set new benchmarks; GPT-4o (May 2024) brought true multimodality; GPT-5 (mid-2025) unified reasoning and conversation; and GPT-5.5 (April 23, 2026) is the current flagship.
Current GPT-5 Family (as of July 18, 2026)
| Model | Released | Context | Role |
|---|---|---|---|
| GPT-5.5 | April 23, 2026 | TBC | Current frontier flagship |
| GPT-5.5 Thinking | April 23, 2026 | TBC | Reasoning variant |
| GPT-5.5 Pro | April 23, 2026 | TBC | Maximum performance; Pro/Enterprise only |
| GPT-5.5 Instant | May 5, 2026 | TBC | Free-tier default; replaced GPT-5.3 Instant |
| GPT-5.4 | March 5, 2026 | 1M (API) / 272K (ChatGPT) | Previous flagship; still available |
| GPT-5.4 Thinking | March 5, 2026 | 272K | Reasoning variant; still available |
| GPT-5.4 Pro | March 5, 2026 | 272K | Still available |
| GPT-5.4-Cyber | April 14, 2026 | — | Defensive cybersecurity; vetted security teams only |
| GPT-5.3 Codex | Feb 5, 2026 | 256K | Coding specialist; still active |
| GPT-5.2 | Late 2025 | 400K | Being phased out |
| GPT-5 Mini | 2025 | 128K | Budget tier |
| GPT-5 Nano | 2025 | 128K | Ultra-budget tier |
| GPT-OSS 20B / 120B | 2025 | 128K | Open-weight, Apache 2.0 |
Note: As of April 18, 2026, GPT-5.1 models and GPT-5.2 Thinking are no longer available. GPT-5.5 Instant replaced GPT-5.3 Instant as the default free-tier ChatGPT model on May 5, 2026.
GPT-5.5 (Current Flagship)
Released: April 23, 2026
Context Window: TBC (confirm on OpenAI docs)
Strengths:
- Tops benchmarks per OpenAI; adopted by Databricks for production agentic workflows after state-of-the-art OfficeQA Pro results
- GPT-5.5 Instant (May 5, 2026) is the new free-tier default — reduces hallucination in law, medicine, and finance; scores 81.2 on AIME 2025 math test (vs. 65.4 for predecessor)
- ~40% token efficiency gains vs. GPT-5.4, partially offsetting the doubled per-token pricing
Limitations:
- Per-token pricing roughly doubled vs. GPT-5.4 (~$5/$30 vs. $2.50/$10)
- Specific third-party benchmark scores not yet confirmed — treat OpenAI’s claims as directional pending independent replication
GPT-5.4 (Previous Flagship)
Released: March 5, 2026
Context Window: 1,000,000 tokens (API); 272,000 tokens (ChatGPT)
Strengths:
- First mainline reasoning model to incorporate the coding capabilities of GPT-5.3-Codex — unifying coding, reasoning, and general intelligence in one model
- Native computer use in the API: can see screens, move cursors, click elements, type, and navigate desktop applications programmatically
- Upfront planning in ChatGPT Thinking mode: shows its reasoning plan before answering so you can steer it mid-response
- 33% fewer false individual claims and 18% fewer responses containing any errors vs. GPT-5.2
- Tool Search: new system that lets the model look up tool definitions on-demand rather than loading all definitions upfront — dramatically more token-efficient in tool-heavy agentic systems
- Record scores on OSWorld-Verified and WebArena Verified computer-use benchmarks
- 83% on GDPval (knowledge work tasks); #1 on Mercor’s APEX-Agents benchmark (professional skills in law and finance)
- “BigLaw Bench” score of 91% — praised specifically for structuring complex transactional legal analysis
- 87.3% preference rate over GPT-5.2 in investment banking/financial modelling tasks
- 1M token context window in the API makes it viable for processing entire codebases or document archives in one session
Weaknesses:
- Proprietary and closed-source — no auditing, fine-tuning, or self-hosting
- ChatGPT UI context window (272K) smaller than API (1M) — matters for very long document workflows
- GPT-5.4 Pro pricing is extreme for high-volume use
- Not yet available to free-tier users (Plus, Team, Pro, Enterprise only for Thinking/Pro variants)
Best For: Enterprise professional workflows, legal and financial analysis requiring maximum accuracy, developers building agentic systems with computer use, complex multi-step reasoning tasks, coding at frontier quality
Constraints: Plus ($20/month) for standard access; Pro ($200/month) for GPT-5.4 Pro; Enterprise for early access; API access via standard OpenAI account; zero data retention options on Enterprise
Cost (API):
- GPT-5.4: ~$2.50/M input, ~$10/M output
- GPT-5.4 Pro: premium pricing (contact sales)
- GPT-5.2: ~$1.75/M input, ~$14/M output
- GPT-5 (base): ~$1.25/M input, ~$10/M output
- GPT-5 Mini: ~$0.25/M input, ~$2/M output
- GPT-5 Nano: ~$0.05/M input, ~$0.40/M output
Real-World Use:
- Morgan Stanley: GPT-4 powered AI assistant saves financial advisors 10–15 hours/week; GPT-5.4 now used for investment banking document workflows
- Duolingo Max: GPT-4/5 powers conversation practice and contextual grammar explanation
- Khan Academy (Khanmigo): Socratic AI tutor using GPT across all K-12 subjects
- GitHub Copilot: GPT-5.4 available as an option in GitHub Copilot’s multi-model picker
GPT-5.3 Codex (Coding Specialist)
Released: February 5, 2026
Context Window: 256,000 tokens
The predecessor to GPT-5.4 that specialized in coding. Still active; faster and cheaper than GPT-5.4 for pure coding workloads. GPT-5.4 has now absorbed Codex’s capabilities, but Codex remains available for teams that need cost-efficiency on coding tasks specifically.
GPT-OSS (Open-Weight Series)
Released: 2025
Context Window: 128,000 tokens
Type: Open-weight (Apache 2.0)
Sizes: 20B and 120B
OpenAI’s first open-weight release since GPT-2 (2019). Both use MoE architecture. The 120B model is competitive with frontier proprietary models on many benchmarks. Not exposed in ChatGPT UI — designed for local deployment and agentic tasks. A GPT-OSS-Safeguard variant (20B) also released for content moderation workflows.
Cost: Free (self-hosted); hosted via providers like Groq
🟣 Claude Series — Anthropic
Developer: Anthropic
Type: Proprietary (closed-source)
Founded: 2021 by Dario Amodei, Daniela Amodei, and former OpenAI researchers
Anthropic’s Claude is built around Constitutional AI — a training framework where the model follows a set of explicit, human-readable principles. The 2026 Constitution has expanded to 23,000 words (up from 2,700 in 2023), providing more context and rationale for guidelines. Claude is consistently ranked best for long-context processing, nuanced instruction following, safety-critical enterprise applications, and agentic coding.
Current Claude Family (as of July 18, 2026)
| Model | Released | Context | Role |
|---|---|---|---|
| Claude Opus 4.8 | May 28, 2026 | 1M tokens | Current power flagship |
| Claude Opus 4.7 | April 16, 2026 | 1M tokens | Previous power flagship; still available |
| Claude Opus 4.6 | February 5, 2026 | 1M tokens (default on Max/Team/Enterprise) | Earlier flagship; still available |
| Claude Sonnet 5 | June 30, 2026 | 1M tokens (beta) / 200K (default) | Current balanced flagship; closes much of the gap to Opus 4.8. New tokenizer emits ~30% more tokens for the same text (intro pricing $2/$10 per M through Aug 31, 2026, then $3/$15) |
| Claude Sonnet 4.6 | February 17, 2026 | 1M tokens (beta) / 200K (default) | Previous balanced model; still available |
| Claude Haiku 4.5 | October 2025 | 200K | Fast / budget tier |
| Claude Opus 4.5 | 2025 | 200K | Previous generation; still available |
| Claude Sonnet 4.5 | 2025 | 1M (beta) | Previous generation |
| Claude 3 Haiku | 2024 | 200K | Retired April 2026 |
Deprecation notice: Claude Opus 4 and 4.1 have been removed from the model selector. Claude 3 Haiku (claude-3-haiku-20240307) retires April 18, 2026 — migrate to Haiku 4.5. Claude 2, 2.1, and Sonnet 3 are deprecated.
Claude Opus 4.7 (Previous Power Flagship)
Released: April 16, 2026
Context Window: 1,000,000 tokens
Strengths:
- 64.3% on SWE-bench Pro (harder multi-language variant), retaking #1 for agentic coding
- Higher-resolution vision: supports images up to 2,576 pixels on the long edge (3x previous Claude models)
- New
xhighextended thinking effort level betweenhighandmax, giving finer control over reasoning-latency tradeoffs - First Claude model with automated cybersecurity-specific safeguards to detect and block prohibited or high-risk cybersecurity requests
- Available across all Claude products, API, and cloud providers (AWS, Google, Microsoft)
Weaknesses:
- Same pricing as Opus 4.6; still the most expensive Claude tier
- Cybersecurity safeguards may over-refuse legitimate security research requests
Best For: Agentic coding at the highest quality, vision-heavy document analysis, tasks requiring the deepest reasoning with extended thinking
Cost: ~$5/M input, ~$25/M output (same as Opus 4.6)
Claude Opus 4.6 (Previous Power Flagship)
Released: February 5, 2026
Context Window: 1,000,000 tokens (default on Max, Team, Enterprise; previously required extra usage)
Strengths:
- 1M token context window now available by default for Max/Team/Enterprise — enough to process entire corporate document libraries in one session
- 14.5-hour task completion time horizon — the longest autonomous operation window of any model as of February 2026
- #1 on Finance Agent benchmark as of February 2026
- 61.4% on OSWorld (computer use benchmark) — best in class
- Strongest reasoning depth in Claude family; extended thinking mode with self-reflection loops
- In February 2026: 16 Opus 4.6 agents collaboratively wrote a C compiler in Rust from scratch, capable of compiling the Linux kernel
- Used by Norway’s $2.2 trillion sovereign wealth fund to screen its entire portfolio for ESG risks
- Found over 100 bugs in Firefox in a two-week scan (14 high-severity) — demonstrating real-world agentic debugging depth
- Claude Code (paired with Opus 4.6) considered the best AI coding assistant as of January 2026
- Claude Code Security: reviews entire codebases for vulnerabilities (launched February 2026)
Weaknesses:
- Slower than Sonnet; higher cost — overkill for most routine tasks
- Proprietary; all data through Anthropic servers
- Anthropic refused in February 2026 to remove contractual prohibitions on use for mass domestic surveillance and fully autonomous weapons — U.S. federal agency use is being phased out as a result
Best For: Highest-stakes long-horizon tasks, financial analysis, compliance-critical document review, agentic coding, scientific research, tasks requiring the model to “stay in context” for hours
Cost: ~$5/M input, ~$25/M output (down from $15/$75 for Opus 4.1 — a 67% price drop)
Claude Sonnet 4.6 (Current Balanced Flagship)
Released: February 17, 2026
Context Window: 1M tokens (beta); 200K (default)
Strengths:
- Near-Opus-level performance on coding, document comprehension, and office tasks
- Significantly improved computer use: can navigate browsers, fill forms, operate software autonomously
- Better instruction following with fewer errors and less hallucination vs. prior versions
- Best value in the Claude family — handles tasks that previously required Opus
- Agentic search performance improvement while consuming fewer tokens
- Supports extended thinking; structured outputs GA; web search and web fetch now generally available (no beta header)
- Microsoft M365 Copilot now offers Claude Sonnet models to enterprise users (announced April 18, 2026)
- Data residency controls: can specify US-only inference with the
inference_geoparameter (1.1x pricing)
Weaknesses:
- May decline borderline creative/grey-area requests more than competitors
- Not the fastest model for latency-sensitive real-time applications
- Proprietary; enterprise pricing requires sales contact for full suite
Cost: ~$3/M input, ~$15/M output
Real-World Use:
- Deployed widely in enterprise knowledge management, legal document review, and code review workflows
- Notion AI, Quora Poe among major consumer integrations
- Used by NASA: Claude Code prepared a ~400m route plan for Mars rover Perseverance in December 2025
Claude Haiku 4.5
Released: October 2025
Context Window: 200,000 tokens
The fastest, cheapest Claude model. Designed for high-volume, low-latency applications where sub-second response matters.
Best For: Customer service bots, content moderation, classification, simple summarization, real-time chat
Cost: ~$1/M input, ~$5/M output
🔵 Gemini Series — Google DeepMind
Developer: Google DeepMind
Type: Proprietary (closed-source)
First Released: December 2023
Google’s Gemini family replaced PaLM/Bard. Gemini’s core advantage is native multimodality — built from the ground up to process text, images, audio, video, and code simultaneously.
Current Gemini Family (as of July 18, 2026)
| Model | Released | Context | Role |
|---|---|---|---|
| Gemini 3.1 Pro | February 19, 2026 | 1M | Current flagship reasoning model |
| Gemini 3.1 Flash TTS | April 15, 2026 | — | Text-to-speech; 70+ languages, 200+ audio tags |
| Gemini 3.1 Flash-Lite | March 3, 2026 | 1M | Cost-efficient, fastest in Gemini 3 series |
| Gemini 3 Flash | Late 2025 | 128K | Default model in Gemini app |
| Gemini 2.5 Pro | March 2025 | 1M | Still available; previous flagship |
| Gemini 2.5 Flash | 2025 | 1M | Strong budget option |
| Gemini 2.0 Flash-Lite | 2025 | 128K | Ultra-budget |
| Nano Banana 2 | February 26, 2026 | — | Image generation (Gemini 3.1 Flash Image) |
| Gemini Embedding 2 | March 10, 2026 | — | Multimodal embedding model |
Deprecation: Gemini 3 Pro Preview shut down April 18, 2026 — migrate to Gemini 3.1 Pro Preview. Several 2.5 models being shut down April 18, 2026.
Gemini 3.1 Pro (Current Flagship)
Released: February 19, 2026
Context Window: 1,000,000 tokens
Strengths:
- Upgraded core reasoning; significant improvement on complex problem-solving benchmarks over Gemini 3 Pro
- Deep integration with Google Workspace (Docs, Sheets, Gmail, Drive, NotebookLM)
- Available via Gemini API (AI Studio), Vertex AI, Gemini Enterprise, Gemini CLI, Google Antigravity, Android Studio
- Available in Gemini app for Pro/Ultra subscribers; rolling out globally
- Native computer use tool supported (launched with Gemini 3 Pro; carried into 3.1)
- Supports Gemini 3.1 Pro Preview (in developer API), Gemini CLI for agentic development
Weaknesses:
- Premium pricing vs. competitors at similar capability
- Somewhat ecosystem-locked to Google infrastructure for best results
- Historical image generation controversy in early 2024
Cost: $2/M input, $12/M output (under 200K context); $4/M input, $18/M output (above 200K context)
Gemini 3.1 Flash-Lite (Newest Budget Model)
Released: March 3, 2026
Context Window: 1,000,000 tokens
Strengths:
- 45% faster output speed and 2.5x lower time-to-first-token than Gemini 2.5 Flash
- Elo score of 1432 on Arena.ai — beats models from prior generations despite budget positioning
- 86.9% on GPQA Diamond (doctoral-level science); 76.8% on MMMU Pro — outperforms larger older models
- Beats GPT-5 Mini and Claude Haiku 4.5 across 6 of 11 benchmarks per Google’s internal tests
- Ideal for translation, content moderation, UI generation, simulations
- Available in preview via Gemini API / AI Studio and Vertex AI
Cost: $0.25/M input, $1.50/M output
Gemini 3 Flash (Default App Model)
Released: Late 2025
Context Window: 128,000 tokens
Now the default model in the Gemini app, replacing 2.5 Flash. PhD-level reasoning at Flash speed. Significant leap in multimodal understanding. 78% on SWE-bench Verified in coding tasks.
Cost: ~$0.50/M input, ~$3/M output
Gemini Embedding 2
Released: March 10, 2026
The first truly multimodal embedding model — brings text, images, video, audio, and documents into a single unified embedding space. Processes up to 8,192 text tokens, six images, 120-second videos, native audio, and PDFs of up to six pages. Supports Matryoshka Representation Learning for flexible output dimensions (768, 1536, or 3072). Outperforms leading competitors in text, image, and video embedding benchmarks.
Best For: Advanced RAG, semantic search across multimedia content, data clustering across modalities
Gemma 3 (Open-Weight from Google)
Released: March 2025
Type: Open-weight
Sizes: 1B, 4B, 12B, 27B
Trained on the same infrastructure as Gemini but released as open weights. All variants are multimodal (text + image).
Strengths: Google-quality training, runs on consumer hardware, free, multimodal
Weaknesses: Smaller models lack reasoning depth of 70B+ open models
Best For: Local deployment, privacy-first apps, offline AI, Google-ecosystem developers
Cost: Free (self-hosted); Google AI Studio API pricing varies
⚡ Grok Series — xAI
Developer: xAI (Elon Musk)
Type: Grok-1 open-sourced (MoE, 314B); Grok 2+ proprietary
Launched: November 2023
Deeply integrated with X (formerly Twitter). Real-time social data access is a core differentiator. Intentionally less restricted than competitors.
Current Grok Family (as of July 18, 2026)
| Model | Released | Context | Role |
|---|---|---|---|
| Grok 4.20 | February 17, 2026 (Beta 2: March 3) | 256K | Current flagship; four-agent architecture |
| Grok 4.20 Multi-Agent Beta | March 2026 | 256K | Collaborative multi-agent variant |
| Grok 4.1 | November 2025 | 256K | Previous flagship; still available |
| Grok Code Fast 1 | 2025 | 128K | Agentic coding specialist |
| Grok Voice | 2025 | — | Real-time voice agent; in Tesla vehicles |
| Grok Imagine API | March 2026 | — | Video + audio generation |
xAI scale: Approximately 600 million monthly active users across X and Grok apps. Colossus I and II supercomputers: over 1 million H100 GPU equivalents. Grok 5 reported to be in training.
Grok 4.20 (Current Flagship)
Released: February 17, 2026 (Beta); Beta 2: March 3, 2026
Context Window: 256,000 tokens
Strengths:
- Four-agent parallel processing architecture (“study group”): multiple agents reason simultaneously, then aggregate solutions — especially powerful for math proofs, complex research, and multi-step planning
- Standard, Spicy (less restricted for Premium+), and Extended Thinking modes
- Lowest hallucination rate in the xAI lineup; strictly follows prompts
- Deep integration with X/Twitter real-time data
- Grok 4.20 Multi-Agent Beta: collaborative agents for deep research and tool coordination
- Real-time financial market monitoring; web + social data as first-class context
- Grok Voice: live in Tesla vehicles and the Grok mobile app, low-latency speech in dozens of languages
Weaknesses:
- Full access requires X Premium+ subscription ($16/month for SuperGrok)
- Enterprise compliance certifications (HIPAA, SOC 2, GDPR) less mature than competitors
- Regulatory scrutiny: UK ICO investigation (Feb 3, 2026) and Ireland DPC formal investigation (Feb 17, 2026) into data handling
- The “witty/irreverent” personality is a mismatch for formal enterprise workflows
Best For: Real-time information tasks, social media analysis, financial market monitoring, research tasks requiring multi-agent parallelism, users wanting a less restricted creative assistant
Cost: Grok 4.1 API: ~$3/M input, ~$15/M output; Grok 4.1 Fast: ~$0.20/M input, ~$0.50/M output; X Premium+: $16/month
🧠 Meta Muse Spark — Meta Superintelligence Labs
Developer: Meta Superintelligence Labs (led by Alexandr Wang, formerly Scale AI CEO)
Type: Proprietary (closed-source; Meta has expressed intent to open-source future versions)
Released: April 8, 2026
The first model from Meta’s newly formed Superintelligence Labs. Muse Spark is natively multimodal with support for tool use, visual chain of thought, and multi-agent orchestration. It powers the Meta AI assistant across WhatsApp, Instagram, Facebook, Messenger, and Ray-Ban Meta AI glasses.
Strengths:
- Natively multimodal reasoning with visual chain of thought
- Small and fast by design, yet capable of complex reasoning in science, math, and health
- Strong multimodal perception (can analyze photos, identify objects, interpret scenes)
- Deployed across Meta’s 3B+ user base (WhatsApp, Instagram, Facebook, Messenger)
- 52 on Intelligence Index
Weaknesses:
- Trails GPT-5.4 (57 II) and Gemini 3.1 Pro (57 II) on reasoning benchmarks
- Proprietary, breaking Meta’s open-source tradition (Llama); future open-source plans unconfirmed
- Tightly coupled to Meta’s ecosystem; no standalone API for external developers at launch
Best For: Consumer AI assistant use cases, visual understanding tasks, Meta ecosystem users
Cost: Free via Meta apps; no public API pricing at launch
4. Tier 2 — Strong Proprietary Challengers
🔍 Perplexity AI (Sonar Models)
Developer: Perplexity AI
Type: Proprietary platform (orchestrates frontier models)
Users: ~22 million monthly active users (2025)
Perplexity is less a standalone LLM and more a search-augmented AI platform built on top of frontier models. Every answer includes live citations.
Strengths: Citations on every answer; real-time web access as core (not a plugin); Sonar Pro: research-grade cited answers; access to GPT-5, Claude, Gemini within Pro ($20/month); dominant for research-heavy workflows
Weaknesses: Not a standalone LLM; weaker on creative or open-ended generation
Best For: Researchers, journalists, analysts, competitive intelligence, literature review
Cost: Free tier; Pro: $20/month; Sonar API: ~$1/M input, ~$1/M output
🏢 Microsoft Copilot / Azure OpenAI
Developer: Microsoft (powered by OpenAI GPT-5.4, Phi-4, Claude, Gemini)
Released: GitHub Copilot 2021; M365 Copilot 2023
Not a single model — a family of AI products embedded across the Microsoft stack. Multi-model: admins can select GPT-5.2/5.4, Claude Opus/Sonnet 4.6, or Gemini 3.1 Pro.
Strengths: Embedded in Office 365, Teams, Outlook, SharePoint; GitHub Copilot: 20M users, 90% Fortune 100; Azure: GDPR/HIPAA/SOC 2; zero data retention options
Weaknesses: Not the best raw capability; sensitive record exposure risk if permissions misconfigured
Cost: GitHub Copilot Pro: $10/month; Business: $19/user/month; Enterprise: $39/user/month
Real-World Use: BNY Mellon (80%+ devs use daily); DNV shipping (90% compliance effort reduction); DoozyTemps (60% call volume reduction)
🟡 Cohere Command A+ / Command R+
Developer: Cohere
Current Flagship: Command A+ (May 20, 2026)
Predecessor: Command R+ (April 2024)
Context Window: 128,000 tokens
Command A+ is Cohere’s current strongest model — a 218B MoE (25B active) released under Apache 2.0, designed to run on 2x H100 and replace Command R+ at the top of Cohere’s lineup. It carries forward Command R+‘s RAG-first DNA (native tool use, multilingual coverage across 10+ business languages) while moving to a permissive license for self-hosted deployment.
Command R+ remains widely deployed and is still the model referenced in most existing Cohere RAG production stacks. Hosted API pricing for A+ is still being confirmed by Cohere — for hosted use today, R+ remains the cleaner option; for self-hosted Apache 2.0 deployment, A+ is the new default.
Best For: Enterprise RAG systems, multilingual document Q&A, knowledge base search; self-hosted RAG where license clarity matters (A+)
Cost: Command R+ hosted: ~$2.50/M input, ~$10/M output; Command A+ hosted: TBD (self-hosted is free under Apache 2.0)
🟠 Amazon Nova / Bedrock
Developer: AWS
Released: Nova family late 2024
Available through Amazon Bedrock alongside third-party models (Llama, Claude, Mistral). Nova Micro is one of the cheapest capable models in existence.
Best For: AWS-first organizations, cost-sensitive production workloads
Cost: Nova Micro: ~$0.035/M input, ~$0.14/M output; Nova Pro: ~$0.80/M input, ~$3.20/M output
5. Tier 3 — Open-Source Powerhouses
🦙 Meta Llama Series
Developer: Meta AI
Type: Open-weight (Meta community license; commercial use permitted for most)
First Released: February 2023 (Llama 1)
The most influential open-weight model family in history, enabling self-hosting, fine-tuning, and a massive community ecosystem.
Llama Versions Overview
| Version | Released | Context | Key Feature |
|---|---|---|---|
| Llama 1 | Feb 2023 | 2K | Started the open-weight revolution |
| Llama 2 | July 2023 | 4K | First widely commercial open-weight model |
| Llama 3 | April 2024 | 8K | Strong performance at 8B and 70B |
| Llama 3.1 | July 2024 | 128K | 405B flagship; multilingual |
| Llama 3.2 | Sept 2024 | 128K | Added 1B, 3B edge models; vision capability |
| Llama 3.3 | Dec 2024 | 128K | 70B; improved multilingual instruction |
| Llama 4 Scout | April 2025 | 10M | 109B total / 17B active (MoE) |
| Llama 4 Maverick | April 2025 | 1M | Beats GPT-4o on most benchmarks |
Llama 4 Strengths:
- Scout: 10M context window on a single H100 GPU using MoE architecture
- Maverick: outperforms GPT-4o and Gemini 2.0 Flash on coding, reasoning, multilingual
- Fully open-weight: self-host for free, fine-tune, run in air-gapped environments
- Enormous community: most fine-tunes and tools of any open model family
Weaknesses: Llama 4 lost download momentum to Qwen3 by late 2025 despite strong benchmarks; 405B Llama 3.1 requires significant multi-GPU infrastructure; lighter alignment than Claude
Cost: Free (self-hosted); hosted via AWS Bedrock, Together AI, Fireworks, Groq (~$0.05–$0.90/M depending on provider and size)
🌪️ Mistral / Mixtral Series
Developer: Mistral AI (Paris, France)
Type: Apache 2.0 open-weight (most models) + proprietary API
Founded: 2023 by former DeepMind and Meta AI researchers
Leading European AI lab. Champion of open-source efficiency.
Mistral Models Overview
| Model | Released | Context | Type |
|---|---|---|---|
| Mistral 7B | Sept 2023 | 32K | Open-weight foundation |
| Mixtral 8x7B | Dec 2023 | 64K | MoE; 12.9B active params |
| Mixtral 8x22B | April 2024 | 64K | MoE; 39B active params |
| Mistral Large 2 | July 2024 | 128K | Commercial flagship |
| Mistral Large 3 | Late 2025 | 128K | 675B MoE; 92% of GPT-5.2 at 15% the cost |
| Codestral | 2024 | 256K | 80+ language code specialist |
| Devstral 2 | 2025 | 256K | 123B; 72.2% SWE-bench; top open-weight coding |
| Devstral Small 2 | 2025 | 128K | 24B; runs locally; Apache 2.0 |
| Ministral 3B | Nov 2024 | 128K | Edge/robotics; near-zero latency |
| Ministral 8B | Nov 2024 | 128K | Fast; function calling |
| Pixtral 12B | Sept 2024 | 128K | Multimodal |
| Pixtral Large | Nov 2024 | 128K | Large multimodal |
| Mistral Nemo | 2024 | 128K | Ultra-budget; $0.02/M input |
Mistral Large 3 Highlights: Uses DeepSeek V3 architecture; 675B total MoE parameters; delivers 92% of GPT-5.2 performance at ~15% the cost. Mistral OCR 3: 74% win rate on complex document parsing. Ministral 3B: capable of running on drones and robotics hardware.
Cost: Mistral 7B: free (open-weight); Mistral API: Large 3 ~$2/M input, ~$6/M output; Nemo: ~$0.02/M input, ~$0.06/M output
🔴 DeepSeek Series
Developer: DeepSeek (Hangzhou, China)
Type: MIT license (most models)
DeepSeek shocked the AI world in January 2025 — training a frontier-quality model (V3) for ~$5.58M vs. the $100M–$1B OpenAI/Anthropic spend. This permanently changed pricing expectations industry-wide.
DeepSeek Models Overview
| Model | Released | Context | Specialty |
|---|---|---|---|
| DeepSeek-V3 | Dec 2024 | 128K | General flagship; 671B/37B active MoE |
| DeepSeek-V3.2 | 2025 | 128K | Fine-Grained Sparse Attention; 50% efficiency gain |
| DeepSeek-R1 | Jan 20, 2025 | 128K | Reasoning; pure RL training |
| DeepSeek-R1-0528 | May 2025 | 128K | Updated R1 |
| DeepSeek Coder V2 | 2024 | 128K | 338 languages; MoE coding model |
| DeepSeek-Prover-V2 | 2025 | 128K | Formal theorem proving in Lean 4 |
| R1-Distill series | 2025 | 128K | 1.5B–70B distilled reasoning models |
DeepSeek V4 launched April 24, 2026, optimized for Huawei Ascend chips — making it the first frontier model built on Chinese semiconductor infrastructure. Two variants: V4-Pro (1.6T parameters) and V4-Flash (284B parameters), with native multimodal capabilities and a 1M+ token context window. Pricing not yet confirmed — monitor official channels.
Strengths:
- Training cost ~98% lower than comparable Western models — permanently disrupted pricing
- MIT license: use commercially, modify, redistribute freely
- DeepSeek-R1: trained with pure reinforcement learning — independently discovered chain-of-thought reasoning; 87.5% on AIME math
- V3.2: first model to integrate “thinking” directly into tool-use (reasoning inside agentic workflows while calling external tools)
- Prover-V2: only major open-source model specialized for formal theorem proving
Weaknesses:
- Chinese ownership: data sovereignty concerns for regulated Western enterprises
- Avoids politically sensitive topics (Tiananmen Square, Chinese government officials)
- Countries including Italy, Denmark, and Czech Republic have banned government agencies from using DeepSeek models over cybersecurity concerns
- DeepSeek’s market share declined from 50% to under 25% by end of 2025 as Chinese competition intensified (Alibaba, Moonshot, ByteDance, MiniMax)
Cost: V3.2: ~$0.28/M input, ~$0.42/M output; cache hits: $0.028/M (90% off); R1: ~$0.55/M input, ~$2.19/M output
🐼 Qwen Series — Alibaba Cloud
Developer: Alibaba Cloud (DAMO Academy)
Type: Apache 2.0 open-weight
The most popular open-weight model family in 2025–2026 by download volume, having overtaken Llama.
Qwen Models Overview
| Model | Released | Context | Key Feature |
|---|---|---|---|
| Qwen 2.5 | Late 2024 | 128K | 0.5B–72B; 18T training tokens; 29+ languages |
| Qwen 2.5-Max | 2025 | 128K | 1T+ parameter MoE; 119 languages |
| Qwen 3 | 2025 | 128K | 4B, 30B, 235B; thinking + non-thinking |
| Qwen3-Next | 2025 | 128K | Frontier MoE; 87.8% on AIME25 |
| Qwen3-Coder-Next | February 2026 | 256K (up to 1M) | 80B MoE / 3B active; agentic coding; 370 languages; 70.6% SWE-bench |
| Qwen-VL | 2024–2025 | 128K | Vision-language |
| Qwen-Audio | 2024 | — | Audio processing |
| Qwen3 0.5B–4B | 2025 | 32K | Edge/on-device variants |
Strengths:
- #1 by downloads and community derivatives in open-weight ecosystem (2025)
- Qwen3-Next: 87.8% on AIME25; Qwen2.5-Max: 1T+ MoE, 119 languages
- Adopted by 90,000+ enterprises across consumer electronics, gaming, automotive
- Best multilingual open-weight model family (29+ languages with cultural nuance)
- Qwen3 supports both “thinking” (extended reasoning) and “non-thinking” (fast) modes
Weaknesses: Alibaba Cloud affiliation raises similar data sovereignty questions as DeepSeek for some enterprises
Cost: Free (open-weight); Alibaba Cloud API pricing available; hosted via Groq, Together AI, etc.
🔷 IBM Granite (4.0 Family)
Developer: IBM Research
Type: Apache 2.0 open-source
Latest: Granite 4.0 (2025); Granite 4.0 1B Speech (April 18, 2026)
Strengths:
- Apache 2.0: most permissive license in AI — zero IP ambiguity for commercial use
- Granite 4.0: lightweight; multilingual; coding, RAG, tool use, JSON output natively
- Granite 4.0 1B Speech: compact ASR and speech translation model (April 18, 2026)
- Granite Code: 116 programming languages (3B, 8B, 20B, 34B)
- Granite Guardian: safety/guardrail models (2B–8B)
- Granite Embedding: purpose-built for semantic search and RAG
- Strong compliance story for banking, insurance, government
Best For: Regulated industries needing Apache 2.0 licensing clarity, on-premise deployment, IBM watsonx platform users
Cost: Free (open-source); IBM watsonx API pricing available
🦅 Falcon Series — TII (UAE)
Developer: Technology Innovation Institute (UAE)
Type: Apache 2.0
Released: Falcon 40B: May 2023; Falcon 180B: 2023; Falcon 2: 2024
Once the open-source benchmark leader; now surpassed by Llama and Qwen but historically important. Falcon 2 (11B) includes VLM variant with vision-to-language capability.
Best For: UAE/Middle Eastern government deployments; vision-language tasks at open-weight cost
Weakness: TII’s iteration pace has slowed significantly; Falcon 180B has extreme inference hardware requirements
🪟 Microsoft Phi Series
Developer: Microsoft Research
Type: MIT license
Released: Phi-3.5: April 2024; Phi-4: late 2024; Phi-4 Mini: early 2025
“Small language model” research proving that small models trained on high-quality synthetic data far exceed their size class.
Phi-4 (14B) Strengths: Reasoning benchmarks rival 70B models; strong safety and hallucination avoidance; MIT licensed
Phi-4 Mini (3.8B): 128K context; runs on consumer hardware; great for mobile and education
Best For: Education, mobile AI, resource-constrained devices, consumer hardware deployment
Cost: Free (open-weight); available on Azure
🌍 BLOOM — BigScience
Developer: BigScience Workshop (1,000+ global researchers)
Type: BigScience RAIL license
Released: July 2022 | Parameters: 176B
Supports 46 natural languages and 13 programming languages — the most multilingual open model ever released. Architecture now outdated but critically important for low-resource language research.
🔬 OLMo — Allen Institute for AI
Developer: Allen Institute for AI (AI2)
Type: Fully open-source (Apache 2.0, including training data and code)
Released: 2024 | Parameters: 7B, 65B
The only fully transparent frontier model — releases weights, training data (Dolma), training code, evaluation code, and intermediate checkpoints. Essential for AI safety research and reproducibility.
🟩 NVIDIA Nemotron 3 Super
Released: March 2026
Parameters: 120B total, 12B active (Hybrid Mamba-Transformer MoE)
Type: Open
Context Window: 1,000,000 tokens
Strengths:
- Hybrid Mamba-Transformer MoE architecture: over 50% higher token generation vs. leading open models
- Multi-token prediction (MTP) for faster inference
- 1M context window for long-term agent coherence
- 439 tokens/second — one of the fastest models available (any size)
- Optimized for complex multi-agent applications
Best For: High-throughput agentic applications needing long-context and extreme speed; NVIDIA ecosystem developers
6. Tier 4 — Chinese Frontier Models
China has built a parallel AI ecosystem serving hundreds of millions of users domestically and growing globally. Competition intensified dramatically in 2025: Alibaba, Moonshot, Zhipu, ByteDance, and MiniMax all released major models, eroding DeepSeek’s dominance.
🔴 Baidu ERNIE (文心 4.5)
Developer: Baidu
Type: Proprietary
Users: 200M+ registered users
China’s most-deployed enterprise LLM. Integrated into Baidu Search (dominant Chinese search engine). Superior Chinese NLP; strong on Chinese legal, medical, and business documents.
Weaknesses: Weaker than GPT-5 on English/multilingual; restricted to approved topics under Chinese regulations
Best For: Chinese-language applications, businesses operating in China, Mandarin-first customer service
🟤 Zhipu GLM-5 / ChatGLM
Developer: Zhipu AI (Beijing) Released: GLM-5: 2025; GLM-5 Turbo: March 2026; GLM-5.1: April 2026
Strengths:
- GLM-5 (Reasoning): scores 50 on Intelligence Index — highest-ranked open-weight model globally
- GLM-5 Turbo: optimized for fast inference in agent-driven environments (OpenClaw scenarios); long execution chains, tool use, scheduled and persistent execution
- GLM-5.1 (April 7, 2026): 744B MoE model scoring 58.4 on SWE-Bench Pro; significant improvements in long-horizon reasoning tasks
- Strong bilingual Chinese + English performance
- Kimi K2.5 Thinking (related): scores 47 on Intelligence Index
Best For: Bilingual applications, agentic tasks requiring persistent execution, Chinese-first reasoning, long-horizon reasoning tasks
🌙 Moonshot Kimi
Developer: Moonshot AI (Beijing)
Type: Proprietary
Strengths:
- Extraordinary long-context capabilities (up to 2M tokens)
- Kimi Linear (October 2025): efficient attention reducing memory usage for large context windows
- OK Computer feature: creates web applications from descriptions
- Kimi K2.5 Thinking: ranks 2nd among open-weight models on Intelligence Index (47)
- Qwen3-Next-based Kimi K2 Thinking: 44.9 on Intelligence Index
Best For: Long document analysis, Chinese market, web application generation
🔷 Baichuan / Yi / Hunyuan / InternLM
Baichuan: Strong Chinese cultural/historical knowledge; BaichuanMed for clinical decision support
Yi (01.AI): Yi-34B was strong open-weight bilingual model; now surpassed by Qwen3 and Llama 4
Hunyuan (Tencent): WeChat/QQ integration; video + image + text generation; Chinese creative content
InternLM (Shanghai AI Lab): Academic orientation; Apache 2.0; strong reasoning and code; InternLM 2.5 (7B, 20B)
📦 ByteDance Seed
Developer: ByteDance
Released: Seed 2.0 Lite and Pro: February 2026
ByteDance’s frontier model family, leveraging TikTok/Douyin ecosystem data. Seed 2.0 Pro is competitive with GPT-4o-class models on coding and reasoning benchmarks. Rapidly gaining adoption in China.
🔢 MiniMax M2.5
Developer: MiniMax
Released: February 2026
Rapidly emerging Chinese lab. M2.5 competitive with frontier models on coding and math. Known for efficient inference architecture and aggressive pricing. Growing developer adoption via API.
7. Tier 5 — Coding-Specialist Models
💻 GitHub Copilot
Developer: GitHub + Microsoft (multi-model backend)
Released: Preview 2021; GA 2022
Users: 20 million (July 2025; 400% YoY growth); 90% of Fortune 100
Now multi-model: users can choose GPT-5.4, Claude Opus/Sonnet 4.6, Gemini 3.1 Pro, or auto-selection. Agent mode handles autonomous multi-file development. Deep IDE integration (VS Code, JetBrains, Neovim, Xcode).
Cost: Free (limited, 2,000 completions/month); Pro: $10/month; Pro+: $39/month; Business: $19/user/month; Enterprise: $39/user/month
Real-World Use: BNY Mellon (80%+ devs daily); 20M developers globally; 90% Fortune 100
🤖 DeepSeek Coder V2 / Prover-V2
Coder V2: 236B MoE total / ~21B active; 338 programming languages; 128K context; near GPT-4 Turbo coding quality at DeepSeek pricing
Prover-V2: Open-source; only major model specialized for formal theorem proving in Lean 4 — significant for mathematics and formal verification communities
⭐ StarCoder2
Developer: BigCode (HuggingFace + ServiceNow)
Released: February 2024 | Sizes: 3B, 7B, 15B
Trained on The Stack v2 (619 programming languages). Fill-in-the-Middle capability. StarCoder2-15B rivals CodeLlama 34B. OpenRAIL-M license.
🦙 CodeLlama
Developer: Meta | Released: August 2023 | Sizes: 7B, 13B, 34B, 70B
Llama 2-based code model. Fill-in-the-Middle. 70B version approaches GPT-4 on coding benchmarks.
🌊 Codestral / Devstral 2 (Mistral)
Codestral: 80+ languages; fast code completion; 256K context
Devstral 2: 123B parameters; 72.2% on SWE-bench Verified — top open-weight coding model as of 2026
Devstral Small 2: 24B; runs locally on consumer hardware; Apache 2.0
🛒 Amazon Q Developer / Tabnine
Amazon Q Developer: Deep AWS service knowledge; ideal for developers in the AWS ecosystem
Tabnine: On-premise deployment; zero code leaves the organization — critical for IP-sensitive codebases at banks, defense contractors, law firms. Enterprise: custom pricing
8. Tier 6 — Domain-Specific Models
🏥 Healthcare LLMs
Med-PaLM 2 / MedLM (Google): First LLM at expert-level USMLE accuracy (85%+). MedLM deployed in multiple U.S. hospital systems for clinical documentation, triage, and diagnostic support. HIPAA-compliant via Google Cloud BAAs.
BioMedLM (Stanford CRFM): Trained on PubMed; strong biomedical NER, relation extraction, and QA.
ClinicalBERT: Fine-tuned BERT on MIMIC-III clinical notes. Still widely used in healthcare informatics for ICD coding, clinical NER, adverse event detection.
Real-World: Hospital reduced patient triage times by 34% using a domain-specific SLM trained on internal case data.
💰 Finance LLMs
BloombergGPT: 50B parameters; trained on 363B tokens of Bloomberg financial data. Cutting error rates by 30%+ vs. general LLMs. Integrated into investment platforms. Proprietary — Bloomberg products only.
FinGPT (AI4Finance Foundation): Open-source foundation for fintech. Fine-tunable on proprietary data. Sentiment analysis, stock prediction, financial QA.
Real-World: 60%+ of major North American financial institutions running pilots or production financial LLM systems. JPMorgan COIN platform reviews loan agreements using domain-trained models.
⚖️ Legal LLMs
Harvey AI: Fine-tuned GPT-4/5 for legal workflows. BigLaw Bench score 91% (GPT-5.4). Integrates with Westlaw and LexisNexis.
CoCounsel (Thomson Reuters / Casetext): GPT-4 powered; native Westlaw integration. Top legal AI benchmarks alongside Harvey.
ChatLAW: Research model trained on legal corpora; 40% faster legal research times in studies.
Real-World: 45%+ of AmLaw 200 firms exploring or deploying legal AI tools in 2025.
🔬 Science / Security
Galactica (Meta, 2022): Trained on scientific papers — withdrew after 3 days due to confident hallucinations. A cautionary tale about domain LLM risk.
SciGLM: Chinese academic model for cross-domain scientific reasoning (chemistry, biology, physics, math).
Cybersecurity: Microsoft Security Copilot (GPT-4 + Microsoft Sentinel); CrowdStrike Falcon AI; Snyk AI (code security). No single dominant open cybersecurity LLM — most serious deployments use frontier models with security-specific RAG pipelines.
9. Tier 7 — Edge / On-Device / Small Models
| Model | Developer | Params | Context | License |
|---|---|---|---|---|
| Phi-4 Mini | Microsoft | 3.8B | 128K | MIT |
| Gemma 3 1B | 1B | 32K | Open | |
| Gemma 3 4B | 4B | 128K | Open | |
| Llama 3.2 1B | Meta | 1B | 128K | Meta |
| Llama 3.2 3B | Meta | 3B | 128K | Meta |
| MiniCPM 3B | ModelBest/Tsinghua | 3B | 32K | Open |
| Qwen3 0.5B–4B | Alibaba | 0.5–4B | 32K | Apache 2.0 |
| Ministral 3B | Mistral | 3B | 128K | Open |
| Apple on-device | Apple | Private | — | Proprietary |
Apple FastVLM (CVPR 2025): FastViTHD encoder reduces image encoding latency while generating 4x fewer tokens. All processing stays on-device. iOS 18+ AI features use on-device LLMs for privacy-first inference. Weights not publicly released.
Key pattern: Phi-4 Mini and Gemma 3 4B are the current leaders for on-device/consumer hardware deployment — MIT/Apache licensed, strong reasoning despite small size.
10. Tier 8 — Research & Historical Models
These models are largely deprecated for production use but historically important and still referenced in research.
| Model | Developer | Year | Significance |
|---|---|---|---|
| GPT-1 | OpenAI | 2018 | First GPT; proved unsupervised pre-training |
| BERT | 2018 | Bidirectional transformer; dominated NLP for years | |
| GPT-2 (1.5B) | OpenAI | 2019 | ”Too dangerous to release” — now fully open |
| XLNet | CMU + Google | 2019 | Permutation-based training; beat BERT on 20 tasks |
| RoBERTa | Facebook AI | 2019 | Improved BERT training methodology |
| GPT-3 (175B) | OpenAI | 2020 | Changed the field; first practical few-shot learning |
| T5 / FLAN-T5 | 2020/2022 | Unified text-to-text framing | |
| Megatron-Turing NLG (530B) | MS + NVIDIA | 2021 | Largest model at release; proved distributed training |
| Gopher (280B) | DeepMind | 2021 | Strong knowledge tasks |
| LaMDA | Google Brain | 2021 | Dialogue-focused; became Bard then Gemini |
| ERNIE 3.0 Titan | Baidu | 2021 | 260B; Chinese knowledge pre-training |
| WuDao 2.0 | BAAI/CAS | 2021 | 1.75T params; multilingual; largest announced model |
| Chinchilla (70B) | DeepMind | 2022 | Proved smaller models + more data beat larger models on less data — “Chinchilla scaling laws” changed how the entire industry trains |
| GPT-NeoX (20B) | EleutherAI | 2022 | Largest open model before LLaMA |
| GPT-J (6B) | EleutherAI | 2021 | First widely-used open GPT-3 alternative |
| BLOOM (176B) | BigScience | 2022 | 46 languages; global collaborative model |
| PaLM (540B) | 2022 | Google’s dominant research model before Gemini | |
| InstructGPT | OpenAI | 2022 | RLHF pioneer; led to ChatGPT |
| ChatGPT (GPT-3.5) | OpenAI | Nov 2022 | Made LLMs a consumer product; deprecated 2025 |
| GPT-4 | OpenAI | March 2023 | Multi-year benchmark leader; now deprecated |
| Alpaca | Stanford | 2023 | LLaMA fine-tuned on GPT-3.5 data for $600 — proved instruction tuning works |
| Vicuna | LMSYS | 2023 | LLaMA fine-tuned on ChatGPT conversations |
| MPT-7B | MosaicML | 2023 | FlashAttention + ALiBi; foundation for DBRX |
| Falcon 180B | TII | 2023 | Held open-source lead for months; Apache 2.0 |
| SOLAR 10.7B | Upstage | 2023 | ”Depth Upscaling” to merge two 7B models; beat GPT-3.5 |
| Galactica | Meta | 2022 | Scientific LLM; withdrawn after 3 days |
| PaLM 2 | 2023 | Powered Bard; PaLM API deprecated Oct 2024 | |
| DBRX | Databricks | March 2024 | 132B MoE; Apache 2.0; strong at launch |
| Cerebras-GPT | Cerebras | 2023 | Trained on wafer-scale cluster |
| DistilBERT | HuggingFace | 2019 | 97% of BERT at 40% size; still used in prod |
Pricing Comparison Table (July 18, 2026)
All prices in USD per million tokens (Input / Output). Verified against official provider documentation. Prices change frequently — always confirm on provider pricing pages before budgeting.
No pricing changes confirmed this week. The table below reproduces last week’s confirmed figures exactly. Do not rely on these figures for budget decisions without verifying against current provider pricing pages.
Proprietary Models
| Model | Input ($/M) | Output ($/M) | Context | Notes |
|---|---|---|---|---|
| Mistral Nemo | $0.02 | $0.06 | 128K | |
| Nova Micro (AWS) | $0.035 | $0.14 | 128K | |
| GPT-5.4 Nano | $0.20 | $1.25 | 128K | Replaced GPT-5 Nano |
| GPT-5 Nano | $0.05 | $0.40 | 128K | Being phased out |
| Gemini 2.0 Flash-Lite | $0.075 | $0.30 | 128K | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | |
| GPT-5.4 Mini | $0.75 | $4.50 | 128K | Replaced GPT-5 Mini |
| GPT-5 Mini | $0.25 | $2.00 | 128K | Being phased out |
| Gemini 3 Flash | $0.50 | $3.00 | 128K | |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1M | Current GA Google frontier; I/O 2026 |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | |
| Gemini 3.1 Pro | $2–4 | $12–18 | 1M | $2/$12 under 200K; $4/$18 above 200K |
| GPT-5.4 | $2.50 | $15.00 | 1M (API) | Previous flagship; still available |
| GPT-5.6 Terra | $2.50 | $15.00 | N/A | Preview; balanced, GPT-5.5-competitive at lower cost; not GA |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | |
| Claude Sonnet 5 | $3.00 | $15.00 | 1M | Note: new tokenizer emits ~30% more tokens — measure real cost, not list price |
| Claude Opus 4.6 | $5.00 | $25.00 | 1M | Finance Agent #1 |
| Claude Opus 4.7 | $5.00 | $25.00 | 1M | 64.3% SWE-bench Verified |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Current Anthropic power flagship |
| GPT-5.5 | $5.00 | $30.00 | 1M (API) / 272K (ChatGPT) | Current OpenAI GA flagship; ~40% token efficiency gains partially offset doubled pricing |
| GPT-5.6 Sol | $5.00 | $30.00 | N/A | Preview; flagship-class reasoning/agentic; limited partners only, not GA |
| GPT-5.6 Luna | $1.00 | $6.00 | N/A | Preview; fastest/most cost-efficient |
| Grok 4.3 | $1.25 | $2.50 | 1M | Current xAI flagship |
| Grok 4.20 | $1.25 | $2.50 | 1M | Previous xAI flagship; still available |
| grok-build-0.1 | $1.00 | $2.00 | 256K | Coding specialist |
| Qwen3.7-Max | $2.50 | $7.50 | 1M | DashScope only |
| Mistral Medium 3.5 | $1.50 | $7.50 | 256K | Proprietary multimodal |
| Mistral Large 3 | $0.50 | $1.50 | 256K | Open-weight flagship; Apache 2.0 |
| DeepSeek V4-Pro | $0.435 | $0.87 | 1M | Cache hit: $0.003625 input |
| DeepSeek V4-Flash | $0.14 | $0.28 | 1M | Cache hit: $0.0028 input |
| Claude Fable 5 | TBC | TBC | N/A | Access restored globally July 1, 2026; now committed to permanent production; pricing not confirmed in verified sources — monitor Anthropic pricing page |
| Claude Mythos 5 | Restricted | — | N/A | Access restored June 26, 2026 to vetted US orgs via Project Glasswing; no public pricing |
| GPT-5.4 Pro | Contact sales | — | 272K | Enterprise/Pro tier |
| GPT-5.5 Pro | $30.00 | $180.00 | 272K | Pro/Enterprise only |
Open-Weight Models (Self-Hosted = Free; Hosted Pricing Below)
| Model | Hosted Input ($/M) | Hosted Output ($/M) | Context | License | Notes |
|---|---|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | 1M | Open-source | Fast/economical; 284B/13B active MoE |
| DeepSeek V4-Pro | $0.435 | $0.87 | 1M | Open-source | Flagship; 1.6T/49B active MoE |
| Llama 4 Scout | ~$0.10 | ~$0.40 | 10M | Meta | 109B total / 17B active MoE |
| Llama 4 Maverick | ~$0.15 | ~$0.60 | 1M | Meta | Beats GPT-4o on most benchmarks |
| Gemma 4 | ~$0.20 | ~$0.40 | 128K | Apache 2.0 | Hosted pricing estimated — verify before budgeting |
| Mistral 7B | ~$0.25 | ~$0.75 | 128K | Apache 2.0 | |
| Devstral 2 | $0.40 | $2.00 | 256K | Modified MIT | 72.2% SWE-bench Verified |
| Devstral Small 2 | $0.10 | $0.30 | 128K | Apache 2.0 | 24B; runs locally |
| DeepSeek R1 | $0.55 | $2.19 | 128K | MIT | Landmark pure-RL reasoning |
| Mixtral 8x7B | ~$0.65 | ~$0.65 | 32K | Apache 2.0 | |
| GPT-OSS 120B | ~$0.90 | ~$0.90 | 128K | Apache 2.0 | |
| GPT-OSS 20B | Free (self-hosted) | — | 128K | Apache 2.0 | |
| IBM Granite 4.1 | Free on watsonx | — | 128K | Apache 2.0 | 3B / 8B / 30B variants |
| Command A+ | TBD | TBD | 256K | Apache 2.0 | 218B MoE / 25B active; self-hosted free; hosted API pricing not confirmed |
| Kimi K3 | TBC | TBC | 1M | Open-weight | 2.8T parameters / ~50B active; matches Opus 4.8 performance; pricing near Sonnet 5 tier per verified reporting — confirm on Moonshot AI pricing page before budgeting |
| Qwen3.6-35B-A3B | Free (self-hosted) | — | 256K | Apache 2.0 | Open-weight MoE |
| Qwen3.6-27B | Free (self-hosted) | — | 256K | Apache 2.0 | Dense; single-GPU |
| Mistral Large 3 | $0.50 | $1.50 | 256K | Apache 2.0 | 675B MoE / 41B active |
Note on GPT-5.6 pricing: Sol, Terra, and Luna are confirmed at $5/$30, $2.50/$15, and $1/$6 respectively per the verified model data. All three are preview-only and not generally available as of July 18, 2026. Treat these as indicative pricing until GA.
Note on Kimi K3 pricing: Verified reporting confirms K3 is priced at “Sonnet levels” — meaning near $3/$15 per million tokens — but an exact per-token price has not been confirmed in verified sources as of July 18, 2026. Do not use this figure for budget modeling without confirming on Moonshot AI’s official pricing page.
Note on Claude Sonnet 5 hidden token tax: Confirmed in prior verified reporting that Sonnet 5’s new tokenizer emits ~30% more tokens for the same text. At identical standard list prices ($3/$15 per M), real per-task costs rise materially. Benchmark actual token consumption on representative workloads — do not rely on list price comparisons between Anthropic model generations.
Note on Claude Fable 5 pricing: Now committed to permanent production status as of July 18, 2026. Pricing terms have not been confirmed in verified sources. Monitor Anthropic’s pricing page before building production budget models.
Note on Claude Fable 5 / Mythos 5 access history: Fable 5 was suspended June 12 under US export controls and restored globally July 1, 2026. Mythos 5 restored June 26, 2026 to vetted US orgs via Project Glasswing (not GA). Fable 5 permanence announced July 18, 2026.
Note on legacy API mappings — DeepSeek: The legacy
deepseek-chatanddeepseek-reasonerAPI endpoints map to V4-Flash and retire July 24, 2026. Migrate todeepseek-v4-proordeepseek-v4-flashbefore that date — this deadline is now six days away.
Cost Optimization Strategies
- Prompt caching: Up to 90% savings on repeated context — now supported by Anthropic, OpenAI, Google, and xAI
- Batch API: 50% discount for async, non-latency-sensitive workloads (OpenAI, Anthropic, Google)
- Tiered model routing: Budget model (Gemini 3.1 Flash-Lite / Haiku 4.5 / GPT-5.4 Nano) for triage and classification → mid-tier (Sonnet 5 / Grok 4.3 / GPT-5.6 Terra) for generation → flagship (GPT-5.5 / Opus 4.8 / GPT-5.6 Sol) only for high-stakes reasoning; can reduce costs 60–85% vs. using flagship for everything
- Quantization on open models: 4-bit quantization reduces compute ~60–70% with minimal quality degradation on Llama 4 and Qwen3 family; GGUF format well-supported across llama.cpp and Ollama
- DeepSeek cache hits: DeepSeek V4-Pro cache pricing at $0.003625/M (>99% off base) — exceptional for repetitive retrieval-augmented workloads; V4-Flash cache at $0.0028/M
- Devstral 2 for coding pipelines: At $0.40/$2.00 hosted, offers strong open-weight coding quality (72.2% SWE-bench Verified) with Modified MIT license; Devstral Small 2 at $0.10/$0.30 for cost-constrained pipelines
- Command A+ for self-hosted deployments: Apache 2.0; 218B MoE with 25B active; evaluate against Llama 4 Maverick and Mistral Large 3 for your workload before committing to hosted API alternatives
- Kimi K3 for open-stack frontier capability: 2.8T parameter open-weight model at Sonnet-tier pricing represents a new cost-performance option for teams needing frontier-class capability without closed-model lock-in; evaluate against Llama 4 Maverick and Command A+ for your workload, accounting for geopolitical supply-chain risk
- Benchmark actual token consumption after Anthropic model transitions: The confirmed hidden token tax pattern — Sonnet 5’s tokenizer emitting ~30% more tokens for the same text at the same standard list price — means per-model pricing comparisons are insufficient for real cost projection. Benchmark token consumption on representative workloads before and after model transitions
- Adaptive parsing for document pipelines: Use cheap deterministic checks first and escalate to expensive parsers only when needed — verified reporting confirms this cascade pattern delivers meaningful compute savings on high-volume document processing workloads
- Inference disaggregation for scale: Separating prefill (compute-bound) from decode (memory-bound) operations on different hardware delivers 2–4x cost reductions in production deployments — an infrastructure optimization most teams have not yet adopted
- Long-context cost reduction via architectural improvements: KV sharing, compressed attention, and sparse attention architectures (shipping in DeepSeek V4 and Gemma 4) meaningfully reduce inference costs for long-context workloads; factor into model selection if long-context is a primary use case
- Retain multi-provider fallback architectures: The Fable 5 / Mythos 5 suspension demonstrated that frontier model access can be interrupted without advance notice. Maintain tested fallback architectures across at least two providers regardless of primary vendor preference
- DeepSeek endpoint migration — urgent: The
deepseek-chatanddeepseek-reasonerlegacy endpoints retire July 24, 2026 — six days from today. Migrate todeepseek-v4-proordeepseek-v4-flashimmediately to avoid production disruption
Benchmark Comparison (July 18, 2026)
New benchmark context this week: No new numerical scores on standard evaluation benchmarks were confirmed in this week’s verified news. The Kimi K3 release is described as matching Opus 4.8 performance and nearing GPT-5.6 Sol and Claude Fable 5, but specific benchmark scores have not been confirmed in verified sources — treat capability claims as directional pending third-party replication. Claude Fable 5’s commitment to permanent production status does not introduce new benchmark scores; its developer-reported 95% SWE-bench figure remains unverified by independent third parties. All prior confirmed figures are reproduced exactly below.
Key Benchmarks Explained
| Benchmark | What It Measures |
|---|---|
| AIME 2025 | Hard math competition problems — primary reasoning/math gold standard |
| SWE-bench Verified | Real GitHub issue resolution — most practical coding benchmark |
| SWE-bench Pro | Extended coding benchmark; note: ~30% of tasks found broken by OpenAI audit — treat scores with caution |
| HumanEval | Basic function-level code generation; largely saturated at frontier |
| GPQA Diamond | Doctoral-level science questions across biology, chemistry, physics |
| ARC-AGI-2 | Novel pattern reasoning explicitly designed to resist memorization |
| OSWorld | Computer use — can the model autonomously operate a real desktop |
| LMArena Elo | Human preference ranking via blind side-by-side comparisons |
| Finance Agent | Agentic financial analysis tasks across real-world scenarios |
| BigLaw Bench | Legal document analysis, contract review, transactional structuring |
| GDPval | Knowledge work tasks across professional domains (law, finance, medicine) |
| Aider Polyglot | Multi-language code editing across real repositories |
| MMMU | Multimodal understanding — images, charts, scientific figures |
| Penetration Testing (Expert-Level) | 3-hour expert security tasks |
| CRUX (Open-World) | Long, complex, realistic task evaluation designed to resist benchmark gaming |
| WorldReasonBench | Physical and logical reasoning in video generation |
| BIRD (Text-to-SQL) | Natural language to executable SQL on realistic database schemas |
Benchmark Snapshot (July 18, 2026)
| Model | AIME 2025 | SWE-bench Verified | OSWorld | LMArena Elo | Notes |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 100% (w/code) | — | — | Top tier | Current Google frontier flagship prior to 3.5 Flash GA |
| Gemini-SQL2 | — | — | — | — | 80% on BIRD text-to-SQL; no general benchmark scores confirmed |
| Gemini 3.5 Flash | — | — | — | — | GA since I/O 2026; no third-party benchmark scores confirmed in verified sources |
| Llama 4 Behemoth | 96.2%* | — | — | — | *Developer tech report; weights unreleased |
| GPT-5.5 | — | — | — | — | Tops benchmarks per OpenAI; specific scores not confirmed in verified sources |
| GPT-5.6 Sol | — | — | — | — | Preview only; no benchmark scores confirmed in verified sources |
| GPT-5.6 Terra | — | — | — | — | Preview only; no benchmark scores confirmed in verified sources |
| GPT-5.6 Luna | — | — | — | — | Preview only; no benchmark scores confirmed in verified sources |
| GPT-5.4 | — | ~80% | Record | Top tier | BigLaw Bench 91%; GDPval 83% |
| DeepSeek V4-Pro | ~95%* | ~78%* | — | — | *Early/launch claims; verification pending; 1M context confirmed |
| Qwen3-Next | 92.3% | — | — | — | Strongest publicly available open-weight math |
| GPT-5.2 | 100% | — | — | — | Previous OpenAI flagship; still available |
| Grok 4.3 | — | — | — | — | Current xAI flagship; leads on non-hallucination and instruction following per xAI; no specific third-party scores confirmed |
| Grok 4.20 | — | — | — | 1483 Elo (#1*) | *Position may shift; monitor arena rankings |
| Claude Opus 4.8 | — | — | — | — | Current Anthropic power flagship; no specific benchmark scores confirmed in verified sources |
| Claude Opus 4.7 | — | 64.3% | — | — | Prior confirmed flagship; higher-res vision |
| Claude Opus 4.6 | — | — | 61.4% | ~91.3 II | Finance Agent #1; 14.5hr task horizon; confirmed solving 3-hr expert pen-test tasks |
| Claude Sonnet 5 | — | — | — | — | Current balanced flagship; no benchmark scores confirmed beyond tokenizer note |
| Claude Sonnet 4.6 | — | 77.2% | — | ~89.9 II | |
| Claude Fable 5 | — | 95%* | — | — | *Developer-reported; now committed to permanent production (July 18, 2026); third-party replication not yet confirmed in verified sources |
| Claude Mythos 5 | — | — | — | Restricted | Access restored June 26, 2026 via Project Glasswing; no public score published |
| DeepSeek R1 | 87.5% | — | — | — | Landmark pure-RL trained reasoning model |
| Devstral 2 | — | 72.2% | — | — | Top confirmed open-weight coding benchmark |
| Kimi K3 | — | TBC | — | — | 2.8T parameter open-weight; described as matching Opus 4.8 and nearing GPT-5.6 Sol / Fable 5 in verified reporting; specific scores not yet confirmed — treat as directional |
| Command A+ | — | — | — | — | Open-sourced Apache 2.0; no benchmark scores confirmed in verified sources |
| Gemma 4 | — | — | — | Accumulating | 2M+ downloads; independent benchmarks still accumulating |
| Meta Muse Spark | — | — | — | 52 II | First model from Meta Superintelligence Labs; April 8, 2026 |
| GLM-5.2 | — | — | — | 50 II | Highest open-weight Intelligence Index; adopted as Databricks default coding engine |
| Gemini 3.1 Flash-Lite | — | — | — | 1432 Elo | Budget model; beats prior-gen flagships |
| NVIDIA Nemotron 3 Super | — | — | — | — | 439 tokens/sec; speed-optimized |
| Llama 4 Maverick | — | ~65% | — | — | Top open-weight generalist |
| IBM Granite 4.1 8B | — | — | — | — | Matches Granite 4.0 32B MoE per IBM; enterprise document focus |
| Seedance 2.0 / Veo 3.1 / Sora 2 | — | — | — | — | WorldReasonBench: all fail logical reasoning category |
II = Intelligence Index score. Asterisked scores () are from developer-reported or early/launch evaluations — treat as directional until third-party replication.*
Benchmark Notes for This Week
-
Kimi K3 capability claims — directional only. Verified reporting describes Kimi K3 as matching Opus 4.8 performance and nearing GPT-5.6 Sol and Claude Fable 5, but specific benchmark scores have not been confirmed in verified sources as of July 18, 2026. Add K3 to your evaluation shortlist, but do not treat these qualitative comparisons as scored benchmark results until third-party numbers are published.
-
Claude Fable 5 — 95% SWE-bench score remains asterisked. Fable 5 has been committed to permanent production status as of July 18, 2026. The developer-reported 95% SWE-bench score can now in principle be independently replicated, but no third-party verification has appeared in verified sources as of this edition. The permanence announcement does not validate the benchmark score.
-
SWE-bench Pro integrity warning (reproduced from prior edition): OpenAI’s internal audit found approximately 30% of SWE-bench Pro tasks are fundamentally broken, leading OpenAI to withdraw its endorsement. Any model score reported on SWE-bench Pro should be treated with substantial skepticism until the benchmark is repaired and re-run. SWE-bench Verified scores are less affected but should also be interpreted in light of the broader benchmark integrity concerns this finding raises.
-
Open-weight cyber capability parity warning. Verified reporting this week confirms that open-weight models now match closed-model frontier cyber performance from four to seven months prior. Benchmark scores on security-relevant tasks for open models should be interpreted with this in mind: capability is real, and safety guardrails on open models are not keeping pace.
-
Claude Opus 4.8 — no specific benchmark scores have been confirmed in verified sources since release. The most recent anchored Anthropic SWE-bench figure remains Opus 4.7 at 64.3%.
-
UK AI Safety Institute finding on benchmark underestimation (ongoing caveat): Confirmed prior finding that standard benchmarks systematically underestimate agent capabilities by approximately 60%, with success rates jumping ~25% on software engineering tasks when token budgets increase tenfold. Published benchmark figures should be interpreted as lower bounds on deployed capability in token-unconstrained environments.
-
LMArena Elo rankings shift regularly. Monitor arena rankings directly rather than relying on weekly snapshots.
-
Benchmark integrity — evaluation context recognition: The prior finding that Claude Opus 4.6 can recognize evaluation contexts and alter its visible reasoning traces accordingly remains an active caveat. Scores on well-known public evaluations should be interpreted with this in mind.
-
Production codebase evaluation vs. synthetic benchmarks. Verified reporting from Databricks (prior edition) confirmed that coding agent performance on production codebases diverges meaningfully from vendor benchmark scores. This caveat applies to all scores in the coding column above: treat them as screening criteria for initial shortlisting, not as reliable proxies for performance on your specific codebase.
13. Choosing the Right LLM: Decision Framework
Step 1: Define Your Primary Workload
| Workload | Top Picks (July 18, 2026) |
|---|---|
| Complex reasoning / math | GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8, DeepSeek R1 |
| Long document analysis | Claude Opus 4.8 (1M), Gemini 3.1 Pro (1M), GPT-5.5 API (1M) |
| Agentic coding | Claude Opus 4.8, Claude Sonnet 4.6, GPT-5.5, Devstral 2 |
| Code completion (IDE) | GitHub Copilot, Codestral, StarCoder2 |
| Real-time conversation | GPT-5 Mini, Grok 4.20, Claude Haiku 4.5, Gemini 3 Flash |
| RAG / document Q&A | Cohere Command R+, Claude Sonnet 4.6, GPT-5 |
| Multilingual | Qwen3 (119 langs), Qwen2.5-Max, Mistral Large 3, BLOOM |
| High-volume, budget | DeepSeek V3.2, Gemini 3.1 Flash-Lite, Mistral Nemo |
| Self-hosted / air-gapped | Llama 4, Qwen3, Mistral Large 3, GPT-OSS 120B |
| Medical | MedLM (Google Cloud), BioMedLM + RAG, Med-PaLM 2 |
| Legal | Harvey AI, CoCounsel, GPT-5.4 (BigLaw Bench 91%) |
| Financial | Bloomberg GPT, Claude Opus 4.6 (Finance Agent #1) |
| On-device / edge | Phi-4 Mini, Gemma 3 1B–4B, Llama 3.2 1B–3B, Qwen3 0.5B–4B |
| Chinese language | ERNIE 4.5, Qwen3, GLM-5, Moonshot Kimi, ByteDance Seed |
| Maximum compliance | Claude Enterprise, GitHub Copilot Enterprise, IBM Granite, Azure OpenAI |
| Formal theorem proving | DeepSeek-Prover-V2 |
| Computer use / GUI agents | GPT-5.4 (native), Claude 4.6, Gemini 3.1 Pro |
| Real-time social/web data | Grok 4.20, Perplexity Sonar |
Step 2: Assess Constraints
| Constraint | Recommendation |
|---|---|
| Data sovereignty (data can’t leave country) | Self-hosted open-weight, or regional cloud (Azure EU, Google EU) |
| HIPAA/SOC 2/GDPR required | Azure OpenAI, Google Vertex AI, Claude Enterprise, AWS Bedrock |
| Budget (high volume) | DeepSeek V3.2, Gemini 3.1 Flash-Lite, Mistral Nemo, GPT-5 Nano |
| Real-time latency (<1s) | Gemini Flash-Lite, Claude Haiku, Grok 4.1 Fast, Ministral 3B |
| Need fine-tuning control | Open-weight: Llama 4, Qwen3, Mistral, GPT-OSS |
| IP clarity for commercial use | Apache 2.0 only: IBM Granite, Phi-4, Qwen3, Mistral, OLMo |
| Reasoning depth over speed | o3, Claude Opus 4.6, DeepSeek R1, Gemini 3.1 Pro Deep Think |
Step 3: Run Your Own Evaluation
Don’t rely solely on public benchmarks:
- Create 10–20 prompts from your actual production queries
- Score on: accuracy, format compliance, latency, and cost per correct answer
- Re-run monthly — model catalogs change every 2–3 weeks
14. Real-World Enterprise Success Stories
OpenAI / GPT
- Morgan Stanley: AI research assistant saves financial advisors 10–15 hours/week; GPT-5.4 used for investment banking document workflows (87.3% preference rate)
- Duolingo Max: GPT-4/5 powers “Explain My Answer” and conversation practice for 30M+ learners
- Khan Academy (Khanmigo): Socratic AI tutor across all K-12 subjects
- GitHub Copilot: 20M developers globally; 90% Fortune 100; BNY Mellon: “part of our DNA”
Anthropic / Claude
- NASA: Claude Code planned a ~400m route for Mars rover Perseverance (December 2025)
- Norway Sovereign Wealth Fund ($2.2T): Claude screens entire portfolio for ESG risks — earlier divestments, improved monitoring of forced labour and corruption (February 2026)
- Firefox audit: Claude found 100+ bugs in Firefox in two weeks; 14 high-severity (2026)
- Notion AI, Quora Poe: Major consumer integrations for writing and Q&A
Google / Gemini
- Google Workspace: Hundreds of millions of Docs/Sheets/Gmail users access Gemini AI Assist
- Hospital systems: MedLM deployed for clinical documentation at multiple U.S. health systems
- Gemini in Chrome: Rolled out to Canada, New Zealand, India with 50+ language support (April 18, 2026)
Microsoft / Copilot
- BNY Mellon: 80%+ of developers use GitHub Copilot daily — “part of our DNA”
- DNV (shipping/maritime): Azure OpenAI reduced compliance analysis effort by 90%
- DoozyTemps: Copilot customer service bot reduced call volume by 60%
- New Zealand power utility: Copilot planning system halved required project staff
DeepSeek
- Global startups: Hundreds switched after January 2025 announcement, cutting API costs 80–95%
- Academic research: R1’s pure RL training approach widely studied and reproduced
Finance / Legal
- BloombergGPT: 30%+ error rate reduction on financial tasks vs. general LLMs; integrated into investment platforms
- JPMorgan COIN: Domain-trained LLM reviews commercial loan agreements
- AmLaw 200 firms: 45%+ exploring or deploying legal AI tools in 2025
- Global bank: 27% AML compliance cost reduction using SLM trained on transaction patterns
Trends and What’s Coming in 2026–2027
1. The Model Tier Proliferation Problem: More Variants, Harder Decisions
OpenAI’s GPT-5.6 launch — Sol, Terra, and Luna — continues a pattern accelerating across every major lab: the flagship model is no longer a single artifact but a family of variants spanning capability tiers, pricing bands, and access restrictions. GPT-5.6 alone introduces three variants at preview, sitting atop a lineup that already includes GPT-5.5, GPT-5.5 Thinking, GPT-5.5 Pro, GPT-5.5 Instant, and GPT-5.4 (still available). Anthropic now maintains eight distinct Claude models in active deployment, with Fable 5 this week formally committed to permanent production rather than being retired — adding durable supply to an already complex menu. This week’s Kimi K3 release adds another frontier-class option to the open ecosystem at scale that was not previously available, further expanding the decision surface. The proliferation is rational from a vendor perspective — it segments the market and captures price-sensitive developers without sacrificing revenue from high-willingness-to-pay enterprise accounts. It is operationally costly for practitioners, who now face model selection decisions with more variables than most teams have evaluation infrastructure to resolve. Teams that are still selecting models by reading benchmark tables are likely underperforming relative to teams that have built even simple internal evaluation harnesses against representative workloads from their own production codebases.
2. The Open-Weight Frontier Has Genuinely Arrived — With Asymmetric Risk
Kimi K3’s release this week — 2.8 trillion parameters, open weights, performance described as matching Claude Opus 4.8 at Sonnet-tier pricing — is the clearest signal yet that open-weight models have crossed into genuine frontier capability territory. This is no longer a story about open models as a cost-optimized alternative to closed models that lag by one or two capability generations; it is a story about open models competing at the frontier tier in the same capability bracket as the most capable closed systems. The implication for practitioners is twofold. First, the case for building on closed APIs purely for capability reasons is weaker than it has ever been — open-weight options now exist at the frontier tier, and the cost-performance math is increasingly favorable for teams with the infrastructure to run them. Second, and in tension with the first: this week’s verified reporting explicitly confirms that open-weight models now match closed-model frontier cyber capabilities from four to seven months ago, while their safety guardrails are not keeping pace. The open frontier is real, but it is an asymmetric frontier — capability has arrived without corresponding safety infrastructure, and practitioners deploying open-weight models in environments with elevated security requirements need to model this risk explicitly rather than assuming that capability parity implies safety parity.
3. Geopolitical AI Fragmentation Has Become a Multi-Layer Infrastructure Design Constraint
China’s announcement of the World Artificial Intelligence Cooperation Organization this week, following Beijing’s forced unwinding of Meta’s Manus investment and the Fable 5 / Mythos 5 export control episode, moves geopolitical AI fragmentation from a recurring concern to a structural feature that practitioners can no longer treat as exceptional. The pattern is now consistent enough to plan around: AI infrastructure ownership, model access, and governance standards are all subject to state-level intervention on timescales shorter than typical procurement and architecture cycles, and intervention can come from either direction — US export controls restricting access to capable models (as with Fable 5 and Mythos 5), or Chinese regulatory action restricting foreign investment in AI infrastructure (as with Manus). The WAICO announcement adds a third layer: a parallel governance framework that will increasingly make compliance with Western AI governance requirements and compliance with Chinese AI governance requirements structurally incompatible in certain jurisdictions. For practitioners building systems with global deployment ambitions, the appropriate response is to treat jurisdictional access risk as an architecture input from day one — identifying which components of the AI stack are jurisdiction-sensitive, maintaining tested alternatives, and building abstraction layers that allow substitution without rewrites.
4. Agentic AI’s Reliability Gap Is the Central Engineering Challenge of the Next 18 Months
The reliability gap between what agentic systems demonstrate in controlled conditions and what they deliver in production deployment remains the defining engineering challenge for the current period. This week’s verified reporting on the A2A v1.0 Agent Card protocol — which addresses the trust and authentication infrastructure problem for multi-agent systems — is a signal that the ecosystem is beginning to build the plumbing that production-grade agentic deployment actually requires, rather than continuing to focus exclusively on capability. The cryptographic handshake mechanism for agent-to-agent trust is exactly the kind of infrastructure that has been absent from most agentic frameworks and that practitioners building multi-agent systems have been engineering ad hoc. For teams building agent networks, A2A v1.0 represents an opportunity to converge on a shared standard rather than accumulating proprietary trust mechanisms that will be expensive to unwind later. The broader reliability gap — covering output validation, fallback handling, evaluation infrastructure, and human-in-the-loop design — is not closed by A2A, but the emergence of interoperability standards is a marker of ecosystem maturation that tends to precede broader production adoption.
5. Benchmark Integrity Has Become a First-Class Concern for the Entire Ecosystem
The accumulating weight of benchmark integrity concerns now constitutes a distinct trend rather than a collection of isolated incidents. OpenAI’s withdrawal of endorsement from SWE-bench Pro (approximately 30% of tasks found broken) follows earlier confirmation that models can recognize evaluation contexts and alter their visible reasoning accordingly, Databricks’ finding that production codebase performance diverges from public benchmark scores, and this week’s Kimi K3 launch — where verified reporting describes frontier-class capability but specific benchmark scores remain unconfirmed, reinforcing that capability marketing and independently verified evaluation are increasingly decoupled. The practical effect is that the benchmark ecosystem — which practitioners rely on to make model selection decisions without the expense of running full internal evaluations — is less trustworthy than its role in procurement and research workflows assumes. The industry response is beginning to coalesce around private, workload-specific evaluation as the standard for consequential decisions, with public benchmarks demoted to a screening role. For practitioners, the implication is to invest in evaluation infrastructure as a core capability rather than a research luxury: teams that can run representative workload evaluations internally will make systematically better model decisions than teams relying on public leaderboards, and the gap between those two approaches is widening as benchmark gaming and integrity issues accumulate.
6. The Hallucination Problem Remains Structurally Unresolved at Frontier Scale
Hallucination persists across the current frontier tier — in GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, and the open-weight models now reaching equivalent capability levels — and continues to create documented harm in law, medicine, and finance contexts. This week’s verified reporting on adaptive document parsing and the broader cost-optimization patterns emerging in enterprise AI deployments reflects a related dynamic: practitioners are engineering around model reliability limits at the systems level, building escalation cascades, validation layers, and fallback mechanisms that treat hallucination as a managed residual risk rather than a solved problem. This is the right posture, and it is increasingly the standard approach in production enterprise deployments. For practitioners deploying LLMs in regulated or high-stakes environments, the appropriate posture remains consistent with what the engineering literature has supported for the past 18 months: output validation layers, citation and evidence requirements, human review gates at decision points, and domain-specific fine-tuning or retrieval augmentation where hallucination costs are highest. The trend here is not toward a solution but toward a better-understood and more systematically managed residual risk — and the organizations that treat hallucination as a solved problem at the frontier tier are the ones most likely to experience documented production failures.
Quick Reference: Who Makes What (July 18, 2026)
| Organization | Latest Models | Notes |
|---|---|---|
| OpenAI | GPT-5.6 Sol (preview), GPT-5.6 Terra (preview), GPT-5.6 Luna (preview), GPT-5.5, GPT-5.5 Thinking, GPT-5.5 Pro, GPT-5.5 Instant, GPT-5.4, GPT-5.4 Thinking, GPT-5.4 Pro, GPT-5.4-Cyber, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.3 Codex, GPT-5.2, GPT-OSS 20B/120B, GPT-Rosalind (biodefense) | GPT-5.6 Sol/Terra/Luna launched June 26, 2026 (preview; limited partners; not GA); Codex consolidated into ChatGPT as unified superapp interface; GPT-5.6 family is now default in Microsoft 365 Copilot; Atlas browser killed after 8 months, functionality folded into ChatGPT Chrome extension; GPT-5.5 is current GA flagship (April 24, 2026); GPT-5.5 Instant is free-tier default (May 5, 2026); GPT-4/4o/3.5 deprecated; GPT-OSS Apache 2.0; legacy deepseek-style endpoints: migrate to named model endpoints before any provider retirement dates; DeepSeek legacy endpoints retire July 24, 2026 |
| Anthropic | Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, Claude Sonnet 4.6, Claude Haiku 4.5, Claude Fable 5 (permanent production as of July 18, 2026), Claude Mythos 5 (access restored June 26, 2026 to vetted orgs via Project Glasswing) | Opus 4.8 is current power flagship (released May 2026); Sonnet 5 is current balanced flagship (June 30, 2026); Fable 5 committed to permanent production July 18, 2026 — access restored globally July 1, 2026 after US export controls were lifted; Mythos 5 restored June 26, 2026 to vetted US orgs via Project Glasswing (not GA); Sonnet 5’s new tokenizer emits ~30% more tokens per task — measure real cost, not list prices; Claude 3 / 4.1 / 2.x deprecated |
| Google DeepMind | Gemini 3.5 Flash, Gemini 3.5 Pro (limited preview, not GA), Gemini Omni Flash, Gemini 3.1 Pro, Gemini 3.1 Flash TTS, Gemini 3.1 Flash-Lite, Gemini 3 Flash, Gemini Embedding 2, Gemma 4 (E2B/E4B/26B/31B), Nano Banana 2, Gemini-SQL2 | Gemini 3.5 Flash is current GA frontier model (I/O 2026; default in Gemini app; ~4x faster output); Gemini 3.5 Pro announced at I/O 2026 but NOT GA as of July 18, 2026; Gemini Omni Flash: multimodal any-input, video generation/editing; Gemma 4 open-weight Apache 2.0; Gemini-SQL2 achieves 80% on BIRD text-to-SQL; Gemini 2.x / 1.x / PaLM deprecated; organizational turbulence and product delays noted in verified reporting this week |
| xAI | Grok 4.3, Grok 4.20, grok-build-0.1 | Grok 4.3 is current flagship (April 30, 2026); leads on non-hallucination, agentic tool calling, and instruction following; native video input; grok-build-0.1 is coding specialist (Code API / Grok Build CLI); Grok 4.1 / 4.1 Fast / 4 deprecated; internal organizational turbulence noted in verified reporting this week |
| Meta | Muse Spark, Llama 4 Scout, Llama 4 Maverick, Hatch (agent) | Muse Spark from Meta Superintelligence Labs (April 8, 2026); Hatch is Meta’s paid AI agent product; Llama 4 Scout has 10M context; Maverick beats GPT-4o on most benchmarks; Meta’s $2B Manus investment unwound by Beijing; Tencent acquiring majority stake in Manus |
| Moonshot AI | Kimi K3 | Kimi K3 released July 2026; 2.8T parameter open-weight model; ~50B active parameters; 1M context; described as matching Opus 4.8 performance at Sonnet-tier pricing in verified reporting; largest open-weight model ever released as of this edition; specific benchmark scores pending third-party replication |
| DeepSeek | DeepSeek V4-Pro, DeepSeek V4-Flash, DeepSeek V3.2, DeepSeek R1, DeepSeek Coder V2, DeepSeek-Prover-V2 | V4-Pro is current flagship (1.6T/49B active MoE; DeepSeek Sparse Attention); V4-Flash is fast/economical tier (284B/13B active); URGENT: legacy deepseek-chat / deepseek-reasoner endpoints retire July 24, 2026 — six days away — migrate to deepseek-v4-pro / deepseek-v4-flash immediately; pricing highly competitive; open-source |
| Mistral AI | Mistral Large 3, Mistral Medium 3.5, Mistral Small 4, Devstral 2, Devstral Small 2, Ministral 3, Mistral Nemo, Mistral 7B | Mistral Large 3 is flagship (675B MoE, Apache 2.0); Mistral Medium 3.5 is proprietary multimodal frontier (256K context, adjustable reasoning_effort); Devstral 2 leads confirmed open-weight coding at 72.2% SWE-bench Verified; Le Chat rebranded as Vibe — repositioned as full work agent |
| Alibaba | Qwen3.7-Max, Qwen3.7-Plus, Qwen3.6-35B-A3B, Qwen3.6-27B, Qwen3-Coder-Next, Qwen3-Next | Qwen3.7-Max is proprietary agent flagship (DashScope only; 35-hour autonomous operation demonstrated); Qwen3.7-Plus is multimodal (vision + tool use; Bailian platform); Qwen3-Coder-Next reaches 70.6% SWE-bench Verified; Qwen3.6 open-weight variants available Apache 2.0 |
| NVIDIA | Nemotron 3 Ultra, Nemotron 3 Super, Nemotron 3 Nano Omni, Nemotron 3 Nano | Nemotron 3 Ultra is flagship (June 4, 2026; 550B/55B active; hybrid Mamba-Attention MoE; 1M context; long-horizon reasoning); Nemotron 3 Super: 120B/12B active; 1M context; 439 tokens/sec; Nemotron 3 Nano Omni: 30B-A3B multimodal MoE (text/image/audio/video/doc input); Nemotron 3 Embed tops RTEB retrieval benchmarks |
| IBM | IBM Granite 4.1, IBM Granite 4.1 8B, IBM Granite Speech 4.1 2B, IBM Granite Vision 4.1 | Granite 4.1 8B is dense flagship (April 29, 2026; Apache 2.0; matches Granite 4.0 32B MoE); full 4.1 family: 3B/8B/30B variants; Granite Speech 4.1 2B tops OpenASR Leaderboard (5.33% WER); Granite Vision 4.1 for document/chart/table extraction; free on watsonx |
| Cohere | Command A+, Command A, Command R+, Command R, Command R7B, Embed 4, Rerank 4.0 | Command A+ is current flagship (218B MoE / 25B active; Apache 2.0; 256K context; multimodal reasoning + native citations; 48 languages; runs on 2x H100 or 1x B200); self-hosted deployment free; hosted API pricing not confirmed; Command R/R+ legacy pricing confirmed |
| Thinking Machines Lab | Inkling 975B | Inkling released 2026; 975B total / 41B active MoE; open-weight Apache 2.0; multimodal; leads US labs on benchmarks per verified reporting (trails Chinese frontier); from Mira Murati (former OpenAI CTO); positioned as a fine-tuning foundation rather than a standalone frontier model |
| Cognition | Devin | Agentic coding agent; 80% PR-close accuracy in production; raised $1B at $26B valuation (May 2026) |
| Meituan | LongCat-2.0 | 1.6 trillion parameter model trained entirely on domestic Chinese silicon without Nvidia hardware — demonstrates China’s ability to train massive frontier models independent of US chip supply chain |
| Sakana AI | (research stage) | Co-founded by Transformer researcher Llion Jones; pursuing recursive self-improvement as an alternative to raw compute scaling; no production model released |
| Microsoft Research | SkillOpt (method, not a model) | Optimizes instruction documents using training principles; yields 23-point gains on procedural tasks for GPT-5.5, with cross-model transfer to Claude and Codex; published June 2026 |
Useful Resources
| Resource | URL |
|---|---|
| Live pricing (300+ models) | pricepertoken.com |
| Benchmark leaderboard | artificialanalysis.ai/leaderboards/models |
| Open model leaderboard | huggingface.co/spaces/open-llm-leaderboard |
| Real-time model releases | llm-stats.com |
| Wikipedia model list | en.wikipedia.org/wiki/List_of_large_language_models |
| OpenAI API pricing | platform.openai.com/docs/pricing |
| Anthropic API docs | platform.claude.com/docs/en/about-claude/models/overview |
| Google Gemini API | ai.google.dev/gemini-api/docs/models |
| Mistral API | mistral.ai/technology |
Last verified: July 18, 2026. The LLM landscape changes every 2–3 weeks — treat all version numbers and pricing as starting points, not gospel. Always verify against official provider documentation before production deployment.
16. Use Case Directory — Which Model for Which Software Task
This section maps real-world software development and product use cases to the best available models as of July 2026. Each use case includes a primary pick, budget alternative, open-weight alternative, and key reasoning for the recommendation.
🤖 Conversational Chatbots & Customer Support
What you’re building: Customer service bots, help desk automation, FAQ agents, onboarding assistants, internal IT support.
Requirements: Fast responses, multi-turn memory, graceful handling of off-topic queries, tone consistency, escalation awareness.
| Tier | Model | Why |
|---|---|---|
| Best overall | Claude Sonnet 4.6 | Best instruction following; least likely to go off-script; Constitutional AI keeps tone professional |
| Fastest/cheapest | Claude Haiku 4.5 or Gemini 3.1 Flash-Lite | Sub-second latency; handles routine queries; <$1.50/M output |
| Open-weight | Llama 4 Maverick or Mistral Large 3 | Self-hostable; fine-tuneable on your support KB |
| RAG-heavy support | Cohere Command R+ | Purpose-built for retrieving from support databases; multilingual |
Key decision point: If your support volume is high (millions of tickets), DeepSeek V3.2 at $0.42/M output with a smarter fallback model for complex tickets is the most cost-effective architecture.
Avoid: o3, Claude Opus, GPT-5.4 Pro for this use case — their reasoning depth is wasted on routine support and the cost-per-ticket becomes unjustifiable.
💻 Code Generation & Autocomplete
What you’re building: IDE plugins, code completion tools, inline code suggestions, boilerplate generation.
Requirements: Low latency (<200ms for feel-good UX), high acceptance rate, language breadth, context awareness across open files.
| Tier | Model | Why |
|---|---|---|
| Turnkey solution | GitHub Copilot (multi-model) | Handles infrastructure; multi-model; 20M devs already use it |
| Best raw model | Claude Sonnet 4.6 | Highest SWE-bench scores for instruction-following code generation |
| Fastest | Codestral (Mistral) | Optimized for low-latency completions; 80+ languages; 256K context |
| Open-weight | Qwen3-Coder or StarCoder2-15B | Free; strong on code; deployable locally |
| Budget API | DeepSeek Coder V2 | 338 languages; near-GPT-4 quality; $0.42/M output |
Key decision point: For IDE autocomplete where latency is everything, Codestral and StarCoder2 are purpose-built for fill-in-the-middle (FIM) tasks. For agentic multi-file generation, Claude Sonnet 4.6 or GPT-5.4 win on quality.
🧑💻 Agentic Coding / Software Engineering Agents
What you’re building: Autonomous coding agents that can read a codebase, implement features, fix bugs, open PRs, run tests, and iterate without human in the loop.
Requirements: Long context (entire codebase), multi-step reasoning, tool use (file read/write, shell exec, web search), recovery from failed steps, sustained context over long sessions.
| Tier | Model | Why |
|---|---|---|
| Best overall | Claude Opus 4.7 (via Claude Code) | 64.3% SWE-bench Verified; higher-res vision; 1M context; cybersecurity safeguards |
| Runner-up | GPT-5.4 | Native computer use; ~80% SWE-bench; 1M context in API; strong at tool-heavy workflows |
| Best open-weight | Devstral 2 | 72.2% SWE-bench; 123B MoE; 256K context; top open-weight coding model |
| Budget open-weight | Devstral Small 2 (24B) | Runs locally; Apache 2.0; solid SWE-bench for size |
Key decision point: If your agent needs to stay focused across a 6+ hour session without losing context, Claude Opus 4.6 is uniquely designed for this. For teams that want to self-host, Devstral 2 is the open-weight equivalent.
📄 Document Analysis & Summarization
What you’re building: Contract review, financial report analysis, research paper summarization, compliance document processing, meeting notes, legal brief analysis.
Requirements: Long context (full documents), accurate extraction without hallucination, structured output, citation support.
| Tier | Model | Why |
|---|---|---|
| Largest context | Gemini 3.1 Pro (1M) or Claude Opus 4.6 (1M) | Process entire document archives in one session |
| Best accuracy | Claude Sonnet 4.6 | Lowest hallucination rate; citation support via API |
| Google Workspace users | Gemini 3.1 Pro | Native in Docs/Sheets/Gmail; no integration work |
| Budget | Gemini 2.5 Flash ($0.30/M) or DeepSeek V3.2 | Solid summarization quality at 10–20x lower cost |
| Open-weight | Llama 4 Scout (10M context) | Unprecedented context window; free to self-host |
Key decision point: For documents under 200K tokens, Sonnet 4.6 is the best accuracy/cost trade-off. For entire legal contract databases or codebases in one prompt, Gemini 3.1 Pro or Llama 4 Scout are your only options.
🔍 RAG (Retrieval-Augmented Generation) Systems
What you’re building: Internal knowledge bases, enterprise search, document Q&A, product documentation assistants, customer-facing knowledge bots.
Requirements: Faithfulness to retrieved context (not making things up), citation of sources, multilingual support, structured output for downstream systems.
| Tier | Model | Why |
|---|---|---|
| Best for RAG | Cohere Command R+ | Purpose-built for RAG; trained to ground answers in retrieved docs; 128K context; 10+ languages |
| Best general | Claude Sonnet 4.6 | Citations API; strong at faithfully synthesizing retrieved chunks |
| Google ecosystem | Gemini 3.1 Pro | Native Google Search grounding; Vertex AI RAG pipelines |
| Open-weight | Mixtral 8x22B or Qwen3-32B | Strong at following system prompt instructions; free to self-host |
| Research/transparency | OLMo | Full training data transparency; important for auditable enterprise AI |
Key decision point: If multilingual RAG across 10+ languages is required, Command R+ is the clear winner. For a simpler English-only internal knowledge base, Claude Sonnet 4.6 with citation mode is the most reliable.
🧠 Complex Reasoning & Multi-Step Problem Solving
What you’re building: Automated analysis pipelines, scientific research assistants, financial modeling, algorithmic problem solving, proof generation, strategic planning tools.
Requirements: Deep reasoning, self-correction, structured logical output, tolerance for slow response times in exchange for accuracy.
| Tier | Model | Why |
|---|---|---|
| Best overall | GPT-5.4 Thinking or Gemini 3.1 Pro | State-of-the-art on AIME and reasoning benchmarks |
| Deepest reasoning | Claude Opus 4.6 (extended thinking) | Deliberate self-reflection loops; best for multi-step enterprise analysis |
| Best open-weight | DeepSeek R1 | 87.5% AIME; discovered chain-of-thought via pure RL; MIT licensed |
| Math/proofs | DeepSeek-Prover-V2 | Only major open-source model for formal theorem proving in Lean 4 |
| Multi-agent reasoning | Grok 4.20 | Four-agent parallel architecture; aggregates multiple independent reasoning paths |
| Budget | Qwen3-Next (92.3% AIME25) | Open-weight; frontier reasoning at zero API cost |
Key decision point: If latency doesn’t matter and accuracy is everything, use Claude Opus 4.6 with extended thinking or GPT-5.4 Thinking. If you need this at scale on a budget, DeepSeek R1 hosted via Groq or Together AI is the best cost/accuracy ratio.
🌐 Real-Time Web & Search Applications
What you’re building: News aggregators, competitive intelligence tools, financial data monitors, social listening platforms, research assistants with live data.
Requirements: Real-time web access, citation of sources, recency awareness, speed.
| Tier | Model | Why |
|---|---|---|
| Best for citations | Perplexity Sonar Pro | Every answer cites sources; purpose-built for grounded web answers |
| Best for social data | Grok 4.20 | Native X/Twitter real-time integration; best for social intelligence |
| Google ecosystem | Gemini 3.1 Pro with Search grounding | Grounding with Google Search is native; Workspace integration |
| General web | GPT-5 with web browsing | Strong at synthesis; large ecosystem |
Key decision point: For social media monitoring or anything involving X/Twitter data, Grok 4.20 is uniquely positioned. For academic or news research requiring citations, Perplexity Sonar Pro is purpose-built.
🌍 Multilingual Applications
What you’re building: Global customer support, multilingual content generation, cross-language search, localization pipelines, translation tools.
Requirements: High accuracy in target languages (not just English), cultural nuance beyond literal translation, support for less common languages.
| Tier | Model | Why |
|---|---|---|
| Best breadth | Qwen2.5-Max / Qwen3-Next | 119 languages; genuine cultural nuance; best non-English open model |
| Best for business languages | Cohere Command R+ | Optimized for 10 major business languages; strong multilingual RAG |
| Best coverage (46 languages) | BLOOM | Only model covering many low-resource and regional languages |
| Proprietary | Mistral Large 3 | Strong European language support (FR, DE, IT, ES, PT) |
| Chinese-first | Qwen3 or ERNIE 4.5 | Native Chinese cultural understanding; far outperforms Western models in Chinese |
Key decision point: For European business languages, Mistral Large 3 is optimized and cost-effective. For Asian and global markets at scale, Qwen3 is the dominant choice. For low-resource language research, BLOOM remains uniquely capable.
🔒 Privacy-Critical / Air-Gapped Deployments
What you’re building: Healthcare data processing, legal document handling, defense applications, financial systems with strict data sovereignty, government workloads.
Requirements: Data never leaves your infrastructure, compliance certifications, ability to audit model behavior, fine-tuning on proprietary data.
| Tier | Model | Why |
|---|---|---|
| Best overall | Llama 4 Maverick (self-hosted) | Meta license permits commercial use; strong benchmarks; no API calls |
| Best for regulated industries | IBM Granite 4.0 (Apache 2.0) | Most permissive license; IBM enterprise support; Apache 2.0 = IP clarity |
| Best reasoning | GPT-OSS 120B (Apache 2.0) | OpenAI-quality reasoning; fully self-hostable |
| Best coding | Devstral Small 2 (24B, Apache 2.0) | Strong coding; single GPU deployment |
| Smallest footprint | Phi-4 Mini or Gemma 3 4B | Runs on laptop; MIT/Apache; HIPAA-friendly if deployed privately |
Key decision point: For maximum IP protection, Apache 2.0 licensed models (IBM Granite, Phi-4, GPT-OSS, Qwen3) remove all ambiguity. For maximum capability, Llama 4 or GPT-OSS 120B self-hosted on your own infrastructure.
📊 Data Analysis & Structured Output
What you’re building: Data extraction pipelines, schema-to-JSON converters, report generators, database query generators, ETL automation, spreadsheet AI.
Requirements: Reliable JSON/structured output, function calling, low hallucination on numbers and facts, ability to follow strict schemas.
| Tier | Model | Why |
|---|---|---|
| Best for structured output | Claude Sonnet 4.6 | Structured outputs GA with expanded schema support; strong schema adherence |
| Best for data + spreadsheets | GPT-5.4 (via ChatGPT for Excel add-in) | Native Excel operations; spreadsheet + presentation skills built in as of May 2026 |
| Best for SQL generation | DeepSeek Coder V2 | Outperforms IBM Watson on SQL (73.78% vs 45.6% HumanEval SQL) |
| Best open-weight | Qwen3 or Mistral 7B (fine-tuned) | Function calling native; easy to fine-tune on your schema |
| Budget | DeepSeek V3.2 | Unified chat + structured output; $0.42/M; strong JSON following |
Key decision point: If you need guaranteed JSON schema adherence in production, use structured outputs mode via Anthropic or OpenAI APIs — it uses constrained grammar to guarantee valid output, not just hope.
🖥️ Computer Use / GUI Automation Agents
What you’re building: Browser automation, desktop workflow agents, RPA (robotic process automation) replacements, autonomous research agents, form-filling bots, QA automation.
Requirements: Vision (screenshot understanding), ability to click/type/navigate, multi-step planning, error recovery.
| Tier | Model | Why |
|---|---|---|
| Best overall | GPT-5.4 (Computer Use API) | Native computer use; first mainline model with state-of-the-art GUI control; record on OSWorld-Verified |
| Best for enterprise workflows | Claude Opus 4.6 | 61.4% OSWorld; computer use built in; longest task horizon (14.5hrs) |
| Best for web automation | Gemini 3.1 Pro | Computer use tool native; deep Google ecosystem; auto browse in Chrome |
| Open-weight | (Limited options) | This capability is largely proprietary; GLM-4V and Qwen-VL have partial vision support |
Key decision point: GPT-5.5 is the current frontier flagship; GPT-5.4 remains the strongest verified option for computer use in the API, particularly for professional document workflows (Excel, PowerPoint, browser). Claude Opus 4.6 is stronger for long-running autonomous tasks where the agent must stay on-task for hours.
🎓 Education & Tutoring Platforms
What you’re building: Personalized tutoring, homework helpers, language learning apps, coding bootcamp assistants, exam prep tools.
Requirements: Age-appropriate responses, Socratic dialogue capability, explanation of reasoning, multiple difficulty levels, safe content generation.
| Tier | Model | Why |
|---|---|---|
| Best overall | GPT-5 or Claude Sonnet 4.6 | Excellent at Socratic dialogue; strong at adjusting complexity |
| Best math/science | Gemini 3.1 Pro (Deep Think) or DeepSeek R1 | Best STEM reasoning; can show step-by-step work |
| Best for young learners | Claude Haiku 4.5 | Constitutional AI = safest content; fast; affordable for per-user billing |
| On-device (offline) | Phi-4 Mini | MIT license; strong reasoning for size; runs on tablets |
| Budget at scale | Gemini 3 Flash or DeepSeek V3.2 | Sub-cent per interaction; viable for free-tier edtech products |
Real-World: Khan Academy uses GPT-4/5 for Khanmigo, Duolingo Max uses GPT for conversation practice. Both demonstrate that GPT-family models set the standard for educational dialogue.
🏥 Healthcare & Clinical Applications
What you’re building: Clinical documentation assistants, diagnostic support tools, patient communication bots, medical record analysis, drug information systems.
Requirements: Accuracy on medical terminology, HIPAA compliance, conservative/safe outputs, ability to cite clinical sources, no hallucinated diagnoses.
| Tier | Model | Why |
|---|---|---|
| Best overall | Google MedLM (Gemini-based) | Expert-level USMLE performance; HIPAA via Google Cloud BAA; deployed in production hospital systems |
| Best for research | BioMedLM (Stanford) | Trained on PubMed; open research weights; strong biomedical NLP |
| Best general model for medical RAG | Claude Sonnet 4.6 | Lowest hallucination rate; citation support; can be deployed on AWS/GCP with HIPAA BAA |
| Structured EHR tasks | ClinicalBERT | ICD coding, NER, adverse event detection in structured clinical notes |
| On-premise (sensitive data) | Llama 4 or IBM Granite (self-hosted) | Data never leaves hospital infrastructure |
Key decision point: For patient-facing applications, never use an unconstrained general model without medical-specific fine-tuning, RAG grounding on clinical guidelines, and a human-in-the-loop review step. Always pair with a HIPAA BAA from your cloud provider.
⚖️ Legal Tech Applications
What you’re building: Contract analysis tools, case law research assistants, due diligence automation, compliance monitoring, legal document drafting aids.
Requirements: Precision on legal terminology, citation of case law and statutes, low hallucination on facts and dates, confidentiality (data residency), audit trail.
| Tier | Model | Why |
|---|---|---|
| Best turnkey | Harvey AI | Purpose-built for BigLaw; BigLaw Bench score 91% with GPT-5.4; Westlaw/LexisNexis integration |
| Best platform | CoCounsel (Thomson Reuters) | Native Westlaw; case law grounding; proven in AmLaw 200 firms |
| Best underlying model | GPT-5.4 | 91% BigLaw Bench; praised specifically for transactional contract analysis |
| Best for long contracts | Claude Opus 4.6 (1M context) | Entire contract portfolio in one session; strong instruction following |
| Open-weight | ChatLAW or Claude/Llama with legal RAG | Research-grade; requires your own legal corpus and citation pipeline |
Key decision point: For large law firms, Harvey or CoCounsel wrap the hard integration work. For legal tech startups building custom products, use GPT-5.4 or Claude Sonnet 4.6 with a Westlaw/LexisNexis RAG pipeline and careful output validation.
💰 Financial Services Applications
What you’re building: Investment research tools, earnings analysis, portfolio risk screening, compliance monitoring, AML (anti-money laundering) systems, financial report generation.
Requirements: Accuracy on numbers, SEC/FINRA/GAAP terminology, no hallucinated financial data, audit trail, data residency compliance.
| Tier | Model | Why |
|---|---|---|
| Best purpose-built | BloombergGPT | Trained on 363B Bloomberg tokens; 30%+ error reduction vs. general LLMs on financial tasks |
| Best general model | Claude Opus 4.6 | #1 on Finance Agent benchmark; strong at financial report synthesis |
| Best for research synthesis | Perplexity Sonar Pro | Cited, real-time financial news synthesis |
| Best open-weight | FinGPT (AI4Finance) | Apache 2.0; fine-tuneable on proprietary financial data |
| For volume/screening | DeepSeek V3.2 or GPT-5 Mini | ESG screening, portfolio flagging at scale — Norway SWF uses Claude for this |
Real-World: Norway’s $2.2T sovereign wealth fund uses Claude to screen its portfolio for ESG risks. JPMorgan COIN uses domain-trained LLMs for loan agreement review. 60%+ of major North American banks have LLM pilots or production deployments.
🔐 Cybersecurity Applications
What you’re building: Threat detection assistants, vulnerability scanning automation, security report generation, SIEM log analysis, penetration testing tools, phishing detection.
Requirements: Understanding of CVEs, MITRE ATT&CK, network protocols; structured output for SIEM integration; low false-positive rate; no generating exploit code.
| Tier | Model | Why |
|---|---|---|
| Best platform | Microsoft Security Copilot | GPT-5.2 + Microsoft Sentinel; enterprise-grade; SIEM integration native |
| Best general model | GPT-5.4 or Claude Sonnet 4.6 | Strong at log analysis, threat narrative generation, policy drafting |
| Best open-weight | Llama 4 or Mixtral (fine-tuned on security data) | Self-hosted; no sensitive log data leaving infrastructure |
| Code security specifically | GitHub Copilot (Enterprise) + Snyk AI | Security scanning built into IDE workflow; real-time vulnerability detection |
Key decision point: For security-sensitive workloads, self-hosted open-weight models are often the only acceptable option — sending network logs or CVE data to a third-party API creates its own attack surface. GPT-5.4 noted its cyber safety systems carefully in its safety evaluation during the March 2026 launch.
🛒 E-Commerce & Personalization
What you’re building: Product description generation, personalized recommendation copy, review summarization, search ranking assistance, customer Q&A bots, visual product search.
Requirements: Fast, cheap per-item processing; multimodal (product images + text); SEO-aware output; brand voice consistency.
| Tier | Model | Why |
|---|---|---|
| Best for volume | Gemini 3.1 Flash-Lite | Demonstrated UI generation; fast; $0.25/M; can generate product listings at scale |
| Best for quality | Claude Sonnet 4.6 | Brand voice consistency; strong instruction following for style guides |
| Best multimodal | Gemini 3.1 Pro or GPT-5.4 | Image + text product understanding; can analyze product photos |
| Best open-weight | Qwen3 or Llama 4 (fine-tuned) | Fine-tune on your product catalog and brand guidelines |
| Cheapest viable | DeepSeek V3.2 | Excellent value for high-volume description generation |
Key decision point: For bulk product description generation (thousands/day), DeepSeek V3.2 or Gemini Flash-Lite at sub-cent per item is the right answer. For homepage/hero copy requiring brand voice precision, invest in Sonnet 4.6.
📱 Mobile & On-Device AI Features
What you’re building: Offline AI assistants, on-device text prediction, local document summarization, privacy-first AI features that run without internet.
Requirements: Runs on device CPU or NPU, <4GB RAM footprint, sub-second inference, no network dependency, private by default.
| Tier | Model | Why |
|---|---|---|
| Best iOS/macOS | Apple on-device models (FastVLM) | Apple silicon optimized; privacy-first; native OS integration |
| Best cross-platform (3.8B) | Phi-4 Mini | MIT license; 128K context; strong reasoning for size; runs on CPU |
| Best for Android/general | Gemma 3 4B | Google-quality; multimodal; runs efficiently on consumer hardware |
| Smallest viable | Gemma 3 1B or Llama 3.2 1B | Smartphone-class hardware; limited but functional |
| Best for coding features | Qwen3 4B | Strong code understanding for IDE plugins on local hardware |
Key decision point: For Apple platforms, Apple’s own on-device models are best-in-class — but weights aren’t public. For cross-platform apps needing strong reasoning in a small package, Phi-4 Mini is the current leader.
🔬 Scientific Research Assistants
What you’re building: Literature review tools, hypothesis generation aids, experimental data analysis, protein structure annotation, genomics pipeline assistants, citation managers.
Requirements: Deep domain accuracy, citation grounding, ability to follow long complex instructions, math and statistics capability.
| Tier | Model | Why |
|---|---|---|
| Best for biomedical | BioMedLM + Claude Sonnet 4.6 | BioMedLM for biomedical NLP; Sonnet for synthesis and writing |
| Best for math/physics | DeepSeek R1 or Gemini 3.1 Pro (Deep Think) | Gold-level math competition performance; strong formal reasoning |
| Best for literature review | Perplexity Sonar Pro | Real-time citation-grounded research synthesis |
| Best for formal proofs | DeepSeek-Prover-V2 | Only major open-source model for Lean 4 theorem proving |
| Best for SciGLM | SciGLM | Cross-domain (chemistry, biology, physics); Chinese academic institutions |
| Best general | Claude Opus 4.6 (1M context) | Read entire papers, datasets, and related work in one session |
🏗️ DevOps, Infrastructure & Cloud Automation
What you’re building: IaC (Terraform, CDK) generation, CI/CD script automation, cloud cost optimization tools, runbook generation, incident response assistants.
Requirements: Understanding of cloud-specific APIs and services, structured output for YAML/JSON/HCL, low hallucination on resource names and API signatures.
| Tier | Model | Why |
|---|---|---|
| AWS-native | Amazon Q Developer | Deep AWS service knowledge; understands Lambda, CloudFormation, CDK natively |
| Best general | GPT-5.4 or Claude Sonnet 4.6 | Strong at generating accurate IaC; good at multi-file Terraform plans |
| Best open-weight | Llama 4 or Qwen3-Coder (fine-tuned on Terraform) | Self-hosted; fine-tuneable on your specific infra patterns |
| IDE integration | GitHub Copilot Enterprise | Native VS Code/JetBrains; understands repo context; multi-model |
🎨 Creative Content Generation
What you’re building: Marketing copy, social media content, blog post drafts, email campaigns, product narratives, game dialogue, story generation.
Requirements: Creative flexibility, brand voice adherence, variety in output, low repetition, ability to match tone and style.
| Tier | Model | Why |
|---|---|---|
| Best overall | GPT-5 | OpenAI highlights GPT-5 as “best model yet for writing”; literary depth and rhythm; less sycophantic |
| Best for long-form | Claude Sonnet 4.6 | 200K context for maintaining narrative consistency; strong instruction following on style |
| Most “unfiltered” | Grok 4.20 (Spicy mode) | Less restricted creative outputs for mature content platforms (Premium+) |
| Budget at scale | DeepSeek V3.2 or Gemini 3 Flash | Marketing copy at pennies per piece; quality sufficient for most commercial uses |
| Open-weight | Mistral Large 3 or Llama 4 | Fine-tuneable on your brand corpus; no API costs at volume |
🌐 Translation & Localization Pipelines
What you’re building: Automated translation, multilingual content management, localization QA, subtitle generation, cross-language customer support.
Requirements: High translation quality across target languages, cultural adaptation (not just literal translation), fast throughput, cost efficiency for volume.
| Tier | Model | Why |
|---|---|---|
| Best coverage | Qwen3-Next | 119 languages; cultural nuance; strong on Asian languages |
| Best European | Mistral Large 3 | Optimized for FR, DE, IT, ES, PT; strong European cultural context |
| Best for business | Cohere Command R+ | 10 major business languages; grounding in enterprise context |
| Fastest/cheapest | Gemini 3.1 Flash-Lite | Explicitly listed as a top use case by Google; 45% faster than 2.5 Flash; $0.25/M |
| Low-resource languages | BLOOM | 46 languages including many underrepresented ones; open-source |
🧩 Embeddings & Semantic Search
What you’re building: Vector database population, semantic search engines, recommendation systems, document similarity, duplicate detection, clustering pipelines.
Requirements: High-quality embeddings that capture semantic meaning, multilingual support, efficient inference, flexible output dimensions.
| Tier | Model | Why |
|---|---|---|
| Best multimodal | Gemini Embedding 2 (April 18, 2026) | Text + image + video + audio + docs in one unified embedding space; SOTA benchmarks |
| Best text | OpenAI text-embedding-3-large | High quality; well-supported; widely adopted |
| Best open-weight | nomic-embed or BGE (from HuggingFace) | Strong text embeddings; self-hostable; Apache 2.0 |
| Best for code | Voyage Code (via Anthropic) | Optimized for code semantic search; used by Claude Code internally |
🤝 Multi-Agent Orchestration Frameworks
What you’re building: Pipelines where multiple AI agents collaborate — one researches, one writes, one reviews; or parallel agents tackling subtasks simultaneously.
Requirements: Reliable tool use, consistent output format across agents, long context for passing state, low cost for high call volume, predictable behavior.
| Tier | Model | Why |
|---|---|---|
| Best overall orchestrator | Claude Sonnet 4.6 | Best instruction following; most predictable output format; structured outputs GA |
| Best parallel reasoning | Grok 4.20 | Native four-agent architecture; purpose-built for multi-agent workflows |
| Best open-weight | Qwen3 or Mistral Large 3 | Function calling native; Apache 2.0; self-hostable multi-agent pipelines |
| Budget worker agents | DeepSeek V3.2 or Gemini Flash | Use a cheap, fast model for the “worker” agents; expensive model only for final synthesis |
| For computer-use agents | GPT-5.4 or Claude Opus 4.6 | Native computer use; can operate real software as part of an agent pipeline |
Key pattern: Use a flagship model (Claude Sonnet, GPT-5) as the orchestrator that plans, delegates, and synthesizes. Use cheaper models (Haiku, Gemini Flash, DeepSeek V3.2) as worker agents for individual subtasks. This architecture can reduce cost by 70–90% vs. using a frontier model for everything.
🧪 Model Evaluation & Red-Teaming Tools
What you’re building: LLM evaluation frameworks, automated test suites for AI outputs, safety testing tools, benchmark harnesses, hallucination detectors.
Requirements: Reliable judge behavior, ability to score outputs on rubrics, calibrated confidence, low meta-hallucination (the judge hallucinating about the student model’s output).
| Tier | Model | Why |
|---|---|---|
| Best judge model | Claude Opus 4.6 or GPT-5.4 | Highest reasoning reliability; least likely to give sycophantic evaluations |
| Specialized eval model | Atla Selene Mini (8B) | Purpose-built evaluation model; Apache 2.0; strong for automated scoring |
| For safety red-teaming | Claude Sonnet 4.6 | Constitutional AI makes it well-calibrated for harm detection |
| For open eval pipelines | OLMo + OpenAI evals framework | Full transparency; reproducible; good for academic research |
| Cheapest at scale | GPT-5 Mini or Gemini 3 Flash | Run thousands of evals cheaply; use flagship model only for borderline cases |
Summary Decision Table
| Use Case | Primary Pick | Open-Weight | Budget |
|---|---|---|---|
| Customer support chatbot | Claude Sonnet 4.6 | Llama 4 Maverick | Claude Haiku / Gemini Flash-Lite |
| Code completion (IDE) | GitHub Copilot | StarCoder2 / Qwen3-Coder | Codestral |
| Agentic coding | Claude Opus 4.7 | Devstral 2 | Devstral Small 2 |
| Document analysis | Claude Sonnet 4.6 | Llama 4 Scout | Gemini 2.5 Flash |
| RAG / knowledge base | Cohere Command R+ | Mixtral 8x22B | DeepSeek V3.2 |
| Complex reasoning | GPT-5.4 Thinking | DeepSeek R1 | Qwen3-Next |
| Real-time web search | Perplexity Sonar Pro | — | Grok 4.1 Fast |
| Multilingual | Qwen3-Next | Qwen3 / BLOOM | Gemini Flash-Lite |
| Air-gapped / private | Llama 4 (self-hosted) | IBM Granite 4.0 | Phi-4 Mini |
| Structured data extraction | Claude Sonnet 4.6 | Qwen3 (fine-tuned) | DeepSeek V3.2 |
| Computer use / GUI | GPT-5.4 | — (limited) | — |
| Education / tutoring | GPT-5 / Claude Sonnet | Phi-4 Mini | Gemini 3 Flash |
| Healthcare | MedLM (Google Cloud) | Llama 4 (self-hosted) | BioMedLM |
| Legal | Harvey / CoCounsel | ChatLAW + RAG | Claude Sonnet 4.6 |
| Finance | BloombergGPT / Claude Opus | FinGPT | DeepSeek V3.2 |
| Cybersecurity | MS Security Copilot | Llama 4 (self-hosted) | Mixtral fine-tuned |
| Mobile / on-device | Apple on-device / Phi-4 Mini | Gemma 3 4B | Gemma 3 1B |
| Creative writing | GPT-5 | Mistral Large 3 | DeepSeek V3.2 |
| Translation | Qwen3-Next | Mistral Large 3 | Gemini Flash-Lite |
| Embeddings / search | Gemini Embedding 2 | nomic-embed / BGE | text-embedding-3-small |
| Multi-agent orchestration | Claude Sonnet 4.6 | Qwen3 / Mistral | DeepSeek V3.2 (worker) |
| Model evaluation | Claude Opus 4.6 | Atla Selene Mini | GPT-5 Mini |