Office Hours — When is fine-tuning a small LLM worth it versus just using a larger model API?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
When is fine-tuning a small LLM worth it versus just using a larger model API?
This is the question I get asked most often, and the honest answer is: fine-tuning rarely wins on pure capability or total cost, but it wins decisively when you care about latency, data privacy, or controlling your own inference stack.
The math rarely favors fine-tuning
Let’s be concrete. Suppose you want a model that extracts structured data from support tickets. You could fine-tune Qwen3.6-35B (an open-weight model) on 5,000 labeled examples. That costs real money upfront: GPU rental, engineering time to prepare data, validation, iteration. You’re probably looking at $2,000-$5,000 for a competent fine-tuning run.
Alternatively, you send every support ticket to Claude Opus 5 or GPT-6 Astra at inference time. Opus 5 runs about $3/$50 per million tokens input/output. If each ticket is 500 tokens and your extraction returns 200 tokens, you’re at roughly $0.0015 per ticket. That’s $7.50 per 5,000 tickets. Over a year processing 100,000 tickets, that’s $150.
The fine-tuned model looks cheaper per inference after you amortize the upfront cost across enough predictions. But here’s what kills that math: fine-tuning didn’t actually solve your problem better than the base model would have. Databricks benchmarked this directly on their production codebase and found GLM-5.2 (an open-weight model) matched Claude Opus 4.8 on coding tasks at $1.28 per task versus $1.94. But they still run it as a smaller frontier model by default because the performance difference only mattered for edge cases. Most of the time, Sonnet 5 or a good open model does the job.
Fine-tuning buys you marginal improvements—maybe 5-15% accuracy gains on specific distributions—not categorical capability jumps. You’re paying thousands to get 5% better. That only pencils out if you’re processing millions of examples, or if the cost savings from fewer API calls directly flows to revenue.
Latency and privacy flip the equation
This is where fine-tuning actually wins. If you need sub-200ms response time for a customer-facing feature, calling Claude or GPT-6 Astra remotely is a non-starter. Round-trip latency to OpenAI or Anthropic is 500ms-2s even when things are fast. You need local inference. That’s when you fine-tune Qwen3.6-35B or Mistral Large 3, run it on your own GPU, and own the latency.
Same with privacy. If you’re processing medical records, financial statements, or anything covered by HIPAA or PCI compliance, sending data to a third-party API is legally and operationally risky. You need it on-prem. That forces you to fine-tune. Fine-tuning in that context isn’t “versus an API,” it’s the only option.
The right question is: are you actually solving a distribution problem?
Fine-tuning works when your data distribution is radically different from what the base model was trained on. If you’re extracting entities from highly specialized technical domains, fine-tuning might help. But if you’re doing anything generic, the frontier models have already seen enough variety that prompt engineering or RAG beats fine-tuning every time.
Real example: A team was fine-tuning Claude Sonnet 4.6 for customer support classification. They hit a wall at ~87% accuracy. They switched to Claude Opus 5 with a better prompt and context window strategy. Accuracy jumped to 94% without fine-tuning. The frontier model was just better at reasoning through ambiguous cases. The problem wasn’t task-specific; it was that the base model underperformed on hard examples.
When fine-tuning actually makes sense
Fine-tuning wins in narrow cases. Your use case has to hit at least two of these:
- You need local inference for latency or privacy reasons.
- You’re running hundreds of thousands or millions of predictions annually and the cost per inference matters at scale.
- Your data distribution is genuinely unusual and you have thousands of labeled examples to prove it works.
If you only hit one, stick with the API. If you hit two, run a pilot. If you hit all three, fine-tune.
The current landscape makes API-first the default
Claude Opus 5 and GPT-6 Astra are so capable that fine-tuning a 7B or 35B model to match them is increasingly difficult. Databricks’ real-world validation on their codebase showed the open-weight GLM-5.2 matched Opus only because they carefully benchmarked on their specific workload. For generic tasks, the gap widens fast.
And frontier model pricing keeps falling. Gemini 3.8 Flash is introductory at $0.75/$3.75 per million tokens through year-end. Sonnet 5 is $2/$10 permanently. Claude Fable 5.1 brings agentic reasoning to $10/$50. These price points make fine-tuning economically harder to justify unless you’re hitting millions of predictions.
The only counterargument is control. If you want determinism, auditability, and to not depend on OpenAI or Anthropic staying in business, fine-tuning an open-weight model on your infrastructure is the answer. That’s a real argument. But it’s not a cost or capability argument. It’s a strategic one.
Bottom line: Default to using a frontier model API. Fine-tune only if you need local inference for latency/privacy, or if you’ve actually validated on your data that fine-tuning closes a capability gap that the base API model won’t.
Question via Hacker News