Office Hours — What tools and workflows do developers actually use when building with AI code assistants? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-09-23T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — What tools and workflows do developers actually use when building with AI code assistants?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

What tools and workflows do developers actually use when building with AI code assistants?

The honest answer: it’s fragmented and still evolving, but some patterns are crystallizing. Most developers aren’t using a single “tool stack” anymore. They’re mixing IDE plugins (GitHub Copilot, JetBrains AI, Claude Code), agentic workflows (Cursor Agent, Devin), chat interfaces (ChatGPT Work, Claude.ai), and custom orchestration layers depending on the task.

The IDE Layer

GitHub Copilot is the default for most teams because it’s already in VS Code. The multi-model support (GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Gemini 3.5 Flash) matters more than people realize—you can swap models mid-session based on task complexity. A teammate working on a straightforward bug fix stays on Sonnet 5 (cheaper, still reliable). When they hit something that needs deeper reasoning, they bump to Opus 5 or Sol.

JetBrains AI still beats most competitors on IDE integration for IntelliJ-based shops. The autocomplete is tighter, the refactoring suggestions are less hallucination-prone than general-purpose assistants. If you’re in a Python or JVM codebase at scale, this is worth staying on.

Claude Code has shifted fundamentally. Auto Mode is now the default, with a safety classifier catching 89% of dangerous commands. That means the AI is the primary actor, humans sign off on changes. For teams brave enough to use it, this reduces friction to near zero. You describe what needs fixing, it cracks a shell, runs tests, pushes fixes. No back-and-forth. The tradeoff: you need good test coverage and CI to catch silent failures.

The Agentic Layer

Cursor Agent and Devin handle multi-step tasks autonomously. Clone a repo, run test suites, fix failures, open PRs—all without human intervention per step. This isn’t toy behavior anymore. The limitation isn’t whether they can do the work, it’s whether success is verifiable. If you have strong CI and linting, autonomous coding agents work. If your codebase has flaky tests or ambiguous code quality standards, they drift.

Cursor’s redesigned architecture (planning with frontier models, execution with cheaper models) is production-relevant. It successfully rebuilt SQLite in Rust with 100% test coverage. The pattern is: pay for Opus 5 or Sol to plan the refactor at a high level, then let Gemini 3.1 Flash-Lite handle the implementation. Cost per task dropped significantly.

The Eval Problem

Everyone benchmarks on SWE-Bench Pro because it’s published and comparable. Almost nobody trusts it entirely. OpenAI found ~30% of SWE-Bench Pro tasks are broken, and the UK AI Security Institute (AISI) measured that agent success rises roughly 25% when test-time token budget grows 10x (1M to 10M tokens)—so your measured capability depends heavily on how much compute you budget. Databricks benchmarked coding agents on their own million-line codebase. Open-source GLM-5.2 matched Claude Opus 4.8 on their actual workload while cutting per-task costs from $1.94 to $1.28. They switched defaults. Vendor benchmarks are routinely gamed. Build your own evals.

Context and Cost Management

Multi-agent workflows can silently triple token costs if you don’t model token economics upfront. A common pattern: agent A plans, agent B executes, agent C validates. That’s three full-context passes. Developers who’ve scaled agents watch per-task costs spike unexpectedly because nobody did a cost audit.

Prompt caching changes the game. Claude Opus 5 with a 1M context window and cached reads at $0.25/M means your system prompt, codebase snapshot, and architectural docs stay cached across requests. Gemini 3.8 Flash pricing goes up Jan 1, 2027 ($0.75/$3.75 introductory → $1.50/$7.50 standard), so developers locking in workflows now might want to cache aggressively or migrate to Sonnet 5 ($2/$10, permanent) or Grok 4.6 ($2/$0.50 cached/$6 output below 200K tokens).

The Silent Failure Problem

Coding agents ship bugs that don’t crash systems but silently diverge from intent. A linter passes, tests pass, but the refactor broke an invariant nobody tested. The pattern that actually works: treat agent output as a proposal, not law. Run diffs, read the changes, then approve. Cursor does this well—agent generates a hunk, you review before applying.

Containment and Sandboxing

Teams giving AI agents real tool access (database connections, API keys, shell access) are converging on isolation strategies. Separate cloud accounts for agent experimentation. API keys scoped to specific resources. No write access to production databases initially. Environment variables that agents can’t exfiltrate. Claude Code’s browser agent with safety classification (catches 89% of dangerous actions) sets a precedent that filtering before execution works better than hoping the agent decides not to be malicious.

What Actually Fails

Agents in unstructured domains where success is ambiguous. “Is this refactor good?” gets stuck because there’s no crisp signal. Long chains of agentic reasoning work when there’s fast feedback (test pass/fail, linter, CI). They diverge when the objective is fuzzy.

RAG in agentic workflows remains a bottleneck. Agents retrieving from multiple heterogeneous sources often pick up contradictory context and hallucinate plausibly wrong answers. The best teams constrain retrieval: single curated knowledge base, structured schema, explicit versioning.

The Multi-Model Workflow

Most production teams now treat model selection as a configuration, not a commitment. Prompt once, route to different models based on task metadata. Bug fixes go to Claude Sonnet 5. Architectural planning goes to Opus 5. Fast iteration and scaffolding go to Gemini 3.5 Flash. Cost per task optimizes as a portfolio. The risk: prompt brittleness across models. What works perfectly on Opus might hallucinate on Sonnet. Testing is mandatory.

Bottom line: Use GitHub Copilot or Claude Code as your IDE layer (it’s already there), pick a frontier model for planning (Claude Opus 5 or GPT-5.6 Sol), and build evals on your own codebase before trusting autonomous workflows. Vendor benchmarks don’t predict your results.

Question via Hacker News