Office Hours — Have you noticed Claude Opus degrading in quality recently, and if so what are the workarounds? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-09-28T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — Have you noticed Claude Opus degrading in quality recently, and if so what are the workarounds?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

Have you noticed Claude Opus degrading in quality recently, and if so what are the workarounds?

I haven’t seen credible evidence of widespread Claude Opus 5 degradation in the signal sources I track or from practitioners I talk to regularly. What I have seen is a pattern of confusion around which Opus you’re using, what “quality” means in your specific task, and how context and prompt structure interact with model capability in ways that look like degradation but often aren’t.

The Opus Timeline Matters

Claude Opus 5 (July 2026) is the current flagship. Claude Opus 4.8 (May 2026) is the previous version, still available. If you’re running Opus 4.8 and comparing it to Opus 5, you should see improvement, not degradation. If you’re comparing Opus 5 to an older checkpoint you remember from your head, the perception gap is real but measurement is hard.

The other pattern I see: people running older code that still calls Opus 4.6 (or even earlier versions) without realizing it, then wondering why results feel off. Check your API calls. Verify you’re actually hitting Opus 5.

What Actually Breaks Claude’s Consistency

Anthropic’s own benchmarks show Opus 5 is stronger on coding, reasoning, and novel problem-solving (ARC-AGI-3 especially) than Opus 4.8. But consistency within a single task depends heavily on factors that have nothing to do with model degradation:

Context window pollution. If you’re reusing cached prompts across different tasks, or if your context is mixing old system instructions with new data, Claude will pick up on the inconsistency. The model isn’t degrading, your prompt architecture is leaking signal. Rebuild your context from scratch for each task rather than appending to a shared cache.

Temperature creep. If you lock temperature to 0, Claude should be deterministic. If it’s not, check whether you’re actually hitting the cached input or if a recent code change is regenerating the prompt slightly differently each time. A single space change in your prompt can break cached input entirely and force a full recompute at higher effective temperature.

Instruction drift. Claude responds strongly to framing. If your system prompt changed subtly (swapped a phrase, reordered instructions, added “be concise”), that’s a real change in behavior that looks like degradation. Version your system prompts explicitly. Compare side-by-side what you’re actually sending.

The Token Explosion Issue

Claude Sonnet 5’s new tokenizer (June 30, 2026) emits roughly 30% more tokens for the same text. If you switched from Opus 4.8 to Sonnet 5 to save money, you’re now paying ~40% more per task in real terms, not less. That’s not degradation, that’s a misaligned cost model. If you’re using Opus 5 (which didn’t change tokenizers), this doesn’t apply.

Concrete Workaround: Rebuild Your Test Suite

Rather than vibe-checking degradation, construct a small eval set of tasks where you care about consistency. Run the same prompt three times at temperature 0 with fresh API calls (no caching) against both Opus 5 and Opus 4.8. If Opus 5 is more consistent, your perception is backwards. If Opus 4.8 wins, you’ve found a real regression, which is worth reporting to Anthropic.

import anthropic

def eval_opus_consistency(prompt, task_name, runs=3):
    client = anthropic.Anthropic()
    results_opus5 = []
    results_opus48 = []
    
    for model in ["claude-opus-5-20250627", "claude-opus-4-8-20250509"]:
        for i in range(runs):
            response = client.messages.create(
                model=model,
                max_tokens=500,
                messages=[{"role": "user", "content": prompt}]
            )
            results = results_opus5 if "opus-5" in model else results_opus48
            results.append(response.content[0].text)
    
    # Compare consistency within each model, not across
    opus5_unique = len(set(results_opus5))
    opus48_unique = len(set(results_opus48))
    
    print(f"{task_name}: Opus5 uniqueness={opus5_unique}/{runs}, Opus4.8 uniqueness={opus48_unique}/{runs}")

Run this on five to ten representative tasks from your actual workload. If Opus 5 is consistently less variable, you’ve got your answer. If they’re the same, the issue is upstream of the model.

If You Find Real Degradation

Anthropic takes regression reports seriously. Document the exact prompt, model ID, and reproducible example. Post in their forum or contact support with a clear reproduction. My read of the platform is that if there’s a real quality drop they’ll fix it quickly—it’s bad for business and bad for trust.

The more likely scenario: your prompt needs tuning for Opus 5’s different reasoning style, or your context is noisier than you realize. Opus 5 is stronger on complex reasoning but also more sensitive to instruction clarity. Make your prompts tighter.

Bottom line: Before assuming degradation, verify you’re running Opus 5 (not an older version), rebuild your prompts with fresh context, and run a small reproducible eval. Most perceived degradation is actually prompt drift or caching confusion.

Question via Hacker News