Office Hours — How should you interview and evaluate developers in a world where AI handles routine coding tasks? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-10-08T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — How should you interview and evaluate developers in a world where AI handles routine coding tasks?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

How should you interview and evaluate developers in a world where AI handles routine coding tasks?

The premise needs inverting. You’re not hiring for routine coding anymore—you’re hiring for judgment calls that AI gets confidently wrong. That changes almost everything about how you screen candidates.

The Confidence-Correctness Gap Is the Real Problem

Autonomous coding agents can now pass SWE-Bench Pro at 72% and handle multi-step tasks with zero human intervention. But they’re also confidently hallucinating tool calls, inventing APIs that don’t exist, and shipping plausible garbage that passes linters. Claude Code, Cursor Agent, and Devin can clone repos, run tests, and open PRs unsupervised. What they can’t do reliably is know when they’re wrong.

This means your interview should probe for the exact opposite of what you’d test five years ago. You don’t need to watch someone write a binary search from scratch. You need to know if they can tell the difference between a working solution and a broken one that looks convincing.

What to Actually Test

Reality-checking ability. Give them a small codebase with a subtle bug introduced by an LLM agent. Not obviously broken—something that compiles, passes most tests, but has a real logic error or performance cliff. Ask them to find it and explain why the agent missed it. This tests whether they read code critically or just trust the machine.

Example: An agent refactors a database query to use async/await, changes look good, tests pass locally, but it silently introduces a race condition under load. Can they spot it? Do they think to load-test before shipping?

Judgment under ambiguity. Describe a real production incident where an AI agent took an action that was technically correct but operationally wrong. Like: “The agent found that renaming a variable makes the code 12% faster. It renamed 247 variables across the codebase. Tests pass. Should we land it?” What questions do they ask? Do they immediately think about git blame, reviewability, team friction? Or do they just measure the performance win?

Architectural awareness. Ask them to design a system where an AI agent can safely modify code autonomously. What safeguards do they propose? Hard budget caps on spending? Automatic rollback triggers? Rate-limited deployments? Output validation before merge? Most candidates will miss at least three of the six critical controls. That gap tells you something.

What to Stop Testing

Don’t whiteboard code. Seriously. An hour of watching someone pseudo-code their way through a sorting algorithm tells you almost nothing about whether they can manage a system where sorting is written by a model.

Don’t test raw algorithm knowledge unless it’s genuinely core to the role. If the job is “build backend services,” ask about caching strategies, API design, deployment patterns. If it’s “build AI agent systems,” ask about cost limits, token economics, hallucination recovery. Optimize for signal on the things that matter now.

Don’t ask them to code a feature from scratch in real-time. Instead, give them a feature implemented by an agent with bugs or edge cases missing, and ask them to debug and improve it. That’s the actual job.

The Pairing Session That Actually Works

Instead of a coding interview, do a “agent oversight” session. Pair them with Claude Code or Cursor Agent on a small, real problem from your codebase. Give them 60 minutes. Then ask:

  • Did they catch what the agent did wrong before shipping?
  • Did they validate the changes or just trust the tests?
  • When the agent got stuck, did they pivot to a different approach or keep pushing harder?
  • Did they think about costs, latency, or operational concerns?

Most candidates will let the agent ship plausible garbage. The ones who ask “wait, does this actually scale?” are the ones who can work in this world.

Red Flags and Green Flags

Red flags: They don’t question the agent’s output. They assume passing tests means correctness. They’re focused on speed over verification. They’ve never thought about token budgets or cost limits. They don’t know what hallucination looks like in their domain.

Green flags: They read the agent’s code like they’d read a peer’s PR—skeptical but fair. They ask about data dependencies, edge cases, and whether the solution works for scale-1 and scale-1000. They’ve deployed agents before and know what breaks. They think about operational concerns (observability, rollback, cost controls) before implementation.

What This Means for Onboarding

Your first week with a new engineer shouldn’t be “set up your IDE and start coding.” It should be “here’s how we use agents, here’s where they work well, here’s where they’ve failed us, and here’s your power to shut them down.” Pair them with someone who’s already managed AI-generated code in production.

The Uncomfortable Truth

Some strong engineers from five years ago won’t adapt to this. They’ll either over-trust agents or spend six hours manually rewriting AI output because they don’t trust it. Neither extreme works. You need people who can operate in the middle: leverage the speed, catch the confidence hallucinations, validate the critical path.

You’re also going to discover that people who thrived at whiteboard interviews often struggle here. And vice versa. Rebalance your hiring rubric accordingly.

Bottom line: Stop testing raw coding ability and start testing judgment under uncertainty. Your job isn’t hiring people who code faster than AI—it’s hiring people who know when to trust it and when to slow down.

Question via Hacker News