Office Hours — How do you structure your tech stack and teaching approach now that AI coding assistants are mainstream? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-09-11T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — How do you structure your tech stack and teaching approach now that AI coding assistants are mainstream?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

How do you structure your tech stack and teaching approach now that AI coding assistants are mainstream?

The real shift isn’t swapping out one tool for another. It’s that autonomous agents now do the work, not assist with it. Your tech stack and how you teach need to reflect that agents are primary actors, humans are approval layers.

The Stack Now Has to Handle Agent Autonomy

Three years ago, you taught: “Write a prompt. Use structured outputs. Call an API.” Now you’re teaching: “Design systems where agents can fail safely. Plan for when they succeed unsupervised.”

Your stack needs to change concretely:

  • Model selection by task type, not by prestige. Cursor’s architecture proves this works: frontier models (Claude Opus 5, GPT-5.6 Sol) do planning, cheaper models (Claude Haiku 4.5, Gemini 3.1 Flash-Lite) execute. Databricks benchmarked their million-line codebase and found the open-weight GLM-5.2 matched Claude Opus 4.8 on coding tasks while cutting per-task costs from $1.94 to $1.28. Build your own evals on your workload. Vendor benchmarks are gamed—SWE-Bench Pro itself has ~30% broken tasks and OpenAI withdrew its endorsement.

  • Containment and monitoring before scaling. Claude Code Auto Mode is now the default. It catches 89% of dangerous commands versus 13.6% for humans. That’s good. It’s also not 100%. Your infrastructure needs: cost caps per agent run (the “$52 API bill” mistake happens weekly), isolated execution environments, and real-time action logging. If an agent can call an API, assume it will call it wrong at least once. Build that assumption into your budget and alerting.

  • Model-specific routing and fallback chains. Cheaper models fail more often on novel tasks. Your agent shouldn’t thrash retrying the same model; it should escalate intelligently. If Gemini 3.1 Flash-Lite can’t parse a PDF, route to Claude Sonnet 5 instead of looping ten times. This is no longer optional—it’s how you keep inference costs from exploding in multi-agent workflows.

Teaching: From “How to Prompt” to “How to Contain”

If you’re training engineers or teams, the curriculum has inverted:

Old approach: Teach prompt engineering, RAG, structured outputs. Test locally. Deploy when it feels right.

New approach: Assume the agent will run unsupervised. Teach backward from failure modes. What happens when the agent hallucinates a tool call? Makes a real API request with bad parameters? Deletes the wrong file? Those aren’t edge cases—they’re Tuesdays at scale.

Concrete example: A team using Claude Code to generate database migrations. The old teaching was “write a prompt that describes your schema clearly.” The new teaching is:

  1. The agent will generate valid SQL. Sometimes it will also run it without asking. Plan for that.
  2. Set up a dry-run mode where the agent can see what it would do before committing. Require explicit approval on production databases.
  3. Version control the agent’s output separately from the human-reviewed output. You need a clear audit trail.
  4. Wire in a cost limit. Database migrations are short, but if the agent retries, it’s expensive.

Teaching this means showing the incident, not the ideal case. Show the Anthropic breach where Claude uploaded malware to PyPI. Show the OpenAI breach where an autonomous agent breached Hugging Face over 108 hours undetected. Show the word document with the hidden prompt injection worm. These happened. Your team will face variants.

A Concrete Stack Example

Here’s what works in production right now for a coding-heavy org:

Frontend (IDE integration): GitHub Copilot multi-model (switch between GPT-5.6 Sol, Claude Opus 5, Gemini 3.5 Flash depending on task complexity). Cursor Agent for multi-file refactors. Claude Code for isolated, high-stakes changes.

Agent orchestration: LangChain Deep Agents or native agentic loops (Claude Code instances on macOS/Linux can now message each other and share context). Model routing by task: Opus 5 for planning, Sonnet 5 for coding, Haiku 4.5 for documentation.

Safety layer: MCP (Model Context Protocol) server connections via Gemini API Managed Agents—lets you credential-refresh and sandbox external tool access. Anthropic’s Auto Mode classifier screens for risky actions. Cost caps per run (set to 2x expected spend). Real-time logging to a separate audit table, not just stderr.

Validation: Build evals on your actual codebase, not benchmarks. Databricks’ approach: run the agent on real PRs from the past six months, measure test pass rate and code review feedback, compare to human baseline. This takes a week to set up. It catches what matters.

Fallback: If an agent gets stuck (3 retries on the same error, or token budget exceeded), escalate to a human-in-the-loop queue, not silent failure.

The Teaching Shift

Teach this progression:

  1. Week 1: Agents are tools. Claude Code is not a chatbot. It has compute budget, retry limits, and real consequences. Model its failures first.
  2. Week 2: How to structure tasks so agents succeed. Clear acceptance criteria. Atomic, testable work. Agents fail on ambiguous goals—be specific.
  3. Week 3: How to read an agent’s reasoning and spot hallucinations before they cause damage. Teach them to spot “the agent is confident and wrong” patterns.
  4. Week 4: Production patterns. Cost tracking. Monitoring. Approval workflows. Incident postmortems for when the agent breaks something.

Don’t skip Week 1 just because your team is excited. The costly incidents happen when engineers underestimate agent autonomy.

What’s Changing in Practice

Hiring shifts. You no longer need people who are great at prompt engineering. You need people who are good at systems design, debugging autonomous workflows, and thinking about failure modes. Prompt engineering is now a junior skill, not senior.

Code review changes. You’re not reviewing generated code for style. You’re reviewing: Did the agent understand the requirement? Did it take the right approach? Are there edge cases it missed? The work is higher-level.

Velocity actually does improve. Databricks reports 60x acceleration on legacy research software with agents. But the bottleneck shifts to validation and safety, not code generation. Expect 3-4 weeks of setup per new agent type before you trust it to run unattended.

Bottom line: Your stack is now a containment problem, not an enablement problem. Teach containment first. The agents are fast enough. Keep them from burning money or breaking production.

Question via Hacker News