Office Hours — What workflow changes have you made to embrace rather than fight AI over-reliance in your development process?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
What workflow changes have you made to embrace rather than fight AI over-reliance in your development process?
The instinct to fight over-reliance on AI is understandable. It feels safer to keep yourself as the primary actor and treat AI as a suggestion engine. In practice, that approach burns tokens, delays decisions, and wastes the one scarce resource you actually have: your attention.
I’ve stopped trying to be the bottleneck. Instead, I’ve restructured workflows to make AI the default actor and myself the checkpoint.
Stop Treating AI as a Suggestion Layer
For three years I treated Claude and GPT like advanced autocomplete: I’d write most of a solution, ask the model to fill gaps, then review everything. This created friction at every step. The model was operating under artificial constraints (incomplete context, mid-task handoff), and I was making micro-decisions on output that the model could have contextualized better from the start.
What changed: I now hand off whole problem statements, let the model propose a complete approach (including tradeoffs and alternatives), and my job is to evaluate the proposal, not fill in the blanks. For a typical feature, this looks like:
- Model generates a full implementation, including test cases and deployment considerations.
- I read the generated code once, checking for architectural fit and whether it matches our actual constraints (auth patterns, database schema, error handling for our specific failure modes).
- If it’s 80% there, I commit it and fix edge cases in the next PR cycle. If it’s fundamentally wrong, I re-prompt with what I learned.
This sounds risky until you realize: my “review-first” approach was also producing bugs. The difference is that now bugs are caught in CI or staging, not in half-baked code I wrote myself.
Route Based on Verify-ability, Not Capability
Not every task should go to the most expensive model. I’ve started routing work based on whether success is verifiable, not whether the task looks hard.
Tasks where tests pass, linters run, and CI is green? Send to a cheaper model (Claude Sonnet 5 or Gemini 3.5 Flash). The model generates code, tests validate it. Done.
Tasks requiring judgment—should we add a caching layer here, is this database design future-proof—those still need a frontier model (Claude Opus 5 or GPT-6 Astra). Not because they’re harder, but because I need reasoning I can audit, not just code that compiles.
A concrete example: I asked Claude Sonnet 5 to implement a database migration for a schema change affecting 200K records. The cost was $0.47 total. It generated the migration, I ran it in staging, migrations succeeded, tests passed. Done in 15 minutes. If I’d used Opus, the cost would have been $3.50 and the time would’ve been the same because I still had to verify the result.
Embrace Asynchronous Delegation with Clear Acceptance Criteria
I define what “done” looks like before handing off work to an agent. No fuzzy success metrics.
Instead of “write a function that handles user authentication,” I now write:
Write a function that:
- Accepts username/password strings
- Returns (user_id: str, session_token: str) on success
- Raises AuthenticationError with message "Invalid credentials" on failure
- Should pass the test suite in tests/auth_test.py
- Must use bcrypt for password hashing (not plain SHA256)
Then I kick off the agent and check back when it reports completion. If tests pass, I move on. If they don’t, I see exactly which assertions failed and can ask for a targeted fix.
The key insight: vague prompts create ambiguous outputs. Clear acceptance criteria let the AI self-correct against an objective signal instead of guessing what I want.
Shift from “Did the AI Get It Right?” to “Does This Pass CI?”
I’ve stopped doing vibe checks on AI output. Instead, I’ve instrumented CI to be my first reviewer.
My GitHub Actions pipeline now:
- Runs type checkers (mypy, eslint) on AI-generated code.
- Runs the full test suite.
- Runs security scanners (bandit for Python, semgrep).
- Only then does a human open the PR.
If any check fails, the bot requests changes inline. This removes the need for me to mentally simulate whether the code works—I just let the machine tell me.
Result: I approve 3-4x more pull requests from agents because I’m not second-guessing logic that passes automated validation.
Accept That Over-Reliance Requires Monitoring, Not Prevention
The real risk isn’t that you rely on AI too much. It’s that you rely on AI without visibility.
I now treat AI-generated code like code from any other team member I don’t sit next to: I require clear commit messages, I check for suspicious patterns (hardcoded IPs, missing error handling), and I monitor what it produces in staging before it hits production.
But I don’t refuse to use it or artificially limit it. That’s just risk denial disguised as caution.
A concrete guard: I’ve set up alerts that fire if any single PR generated by an agent exceeds 500 lines of code. Not because agents can’t write 500+ line PRs, but because that’s a sign the task decomposition was poor and I should have broken it into smaller chunks. The alert is a check on my prompting, not the model.
Measure Output Quality, Not Input Effort
I used to feel guilty when an AI wrote something that took me 10 seconds to prompt and 30 minutes for me to write. That guilt was backwards. The output is what matters.
Now I track:
- Did the tests pass on first try?
- How many issues were found in code review (by humans or automation)?
- How many production bugs came from this code in the first 30 days?
For most AI-generated code: 85%+ passes tests on first try, code review surfaces 0-2 issues per 100 lines (typical for human code is 1-3), and production defect rate is indistinguishable from code I wrote myself.
That last point took me six months to believe. But the data is clear: I’m not worse at shipping reliable code when I delegate to AI. I’m just faster.
Bottom line: Stop trying to keep AI as an assistant and yourself as the actor. Flip it: make AI the default actor, put yourself in the approval workflow, and measure success by whether tests pass and users aren’t complaining, not by how much mental effort you invested.
Question via Hacker News