Office Hours — How are developers using spec-driven development with LLMs to improve code generation reliability?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
How are developers using spec-driven development with LLMs to improve code generation reliability?
Spec-driven development with LLMs flips the traditional prompt-and-hope workflow: instead of asking the model to generate code and then validating it afterward, you give the LLM an executable specification first, and use it as a continuous feedback loop to steer generation toward working code.
The core idea is simple but powerful. You write a test, a type signature, a schema, or a formal constraint upfront. You feed that spec to the LLM as part of the prompt, along with examples of how the spec is validated. Then you ask the model to generate code that satisfies it. When it doesn’t pass, you run the test again, show the failure output to the model, and let it iterate. The spec acts as an objective signal instead of vague English expectations.
Why specs matter more than prompts
LLMs are pattern-matching machines, not theorem provers. When you say “write a function that parses CSV files,” the model generates plausible-looking code that might handle the happy path but silently fail on quoted fields, embedded newlines, or CRLF line endings. The code looks confident. It’s wrong.
When you say “write a function that parses CSV files and passes these 12 test cases,” the model has a concrete target. If it generates code that fails three tests, you can show it the failures. The model sees exactly what it got wrong and adapts. This is orders of magnitude more reliable than iterating on prose feedback.
The difference between “write a parser” and “write code that passes this test suite” is the difference between guessing and measuring.
Concrete patterns in production
Test-driven generation: Write minimal unit tests first. Feed the test file into the context window along with the test framework docs and a few passing examples. Ask the model to generate the function. Run the tests. If they fail, paste the error output back into the next prompt. Claude Opus 5 and GPT-6 Astra both handle this cycle well. Most teams see the failure rate drop from ~40% on first-pass generation to ~5% after 2-3 test-guided iterations.
Schema validation: For data transformation tasks, define your input and output schemas upfront (JSON Schema, Pydantic, Protocol Buffers). Show the model concrete examples of valid inputs and expected outputs. When the model generates code, validate the output against the schema before accepting it. If validation fails, return the schema error and a few failing examples back to the model in the next prompt. This works particularly well for ETL pipelines where silent corruption is worse than loudly failing.
Type-driven generation: Give the model a TypeScript or Python interface. The type signature itself is a spec. Show the model similar functions with the same types. Ask it to fill in the implementation. The type checker becomes your validation gate. If the code doesn’t type-check, show the type errors to the model and ask it to fix them. Frontier models now have good enough understanding of type systems that this loop converges quickly, usually in one or two iterations.
Property-based testing: Instead of writing specific test cases, write properties that your code must satisfy for all inputs within a range. Feed those properties to the model along with a property-based testing framework (QuickCheck, Hypothesis, etc.). The model generates code. The property tester runs thousands of randomized inputs. When it finds a counterexample, show the model the specific input that broke its code. This catches edge cases that hand-written tests miss.
Cost and latency tradeoffs
Spec-driven development increases the number of LLM calls. You’re not just calling the model once; you’re calling it multiple times per task to iterate on failures. On Claude Opus 5 or GPT-6 Astra, this can add up. A simple function that cost $0.005 on the first try might cost $0.015 across three iterations.
But the tradeoff is real: code that actually works. One developer at a mid-size fintech firm benchmarked this internally. First-pass code generation (no specs, no iteration) succeeded 30% of the time. Same developer, same prompts, but with unit tests fed into the prompt and test-driven iteration, succeeded 87% of the time. The cost per successful implementation was actually lower, because the failed attempts in the first approach required human debugging and rework.
If cost is critical, use cheaper models for iteration. Generate with Claude Opus 5 or GPT-6 Astra once to get the skeleton, then iterate failures using Claude Sonnet 5 or Gemini 3.8 Flash. Frontier models are better at understanding why they failed and fixing it, but cheaper models can handle the grinding work of passing tests once the structure is sound.
What still fails
Spec-driven development works when your spec is actually executable and fast to validate. It breaks when your success criteria are subjective. “Write a function that handles edge cases” is not a spec. “Write a function that passes these 47 test cases covering null inputs, empty arrays, and pathological Unicode” is.
If your spec is slow to run, iteration becomes expensive. If your problem is fundamentally about design judgment (“should we cache this result?”), specs won’t save you. Specs enforce correctness, not wisdom.
Also, specs don’t solve the hallucination problem at the architectural level. An LLM can generate code that passes your tests but uses a library that doesn’t exist, or makes an API call to a function with the wrong signature. Show the model the actual API docs, not your assumptions about the API. Include concrete examples of real, working function calls in your spec.
Practical setup
The easiest entry point is to use existing test infrastructure you already have. If you’ve got a test suite in pytest or Jest, you’re most of the way there. Write a simple bash loop: generate code with the LLM, run the tests, capture failures, feed them back to the LLM with the same context. Cursor Agent and Claude Code both have built-in iteration loops that do this automatically. For one-off tasks, just keep your terminal and LLM context window open side-by-side.
For larger systems, think about which tests matter most. You don’t need to spec every function. Spec the critical path: parsing logic, validation, data transformations, authorization checks. Leave the plumbing to faster, cheaper prompts.
Bottom line: Feed executable tests or type signatures into your LLM prompts as specs, run them after generation, and iterate on failures. This converts reliability from a hope into a measurable outcome. Most teams see error rates drop 70%+ with three iterations, and the cost per working function is usually lower than the cost of debugging broken first-pass code.
Question via Hacker News