Office Hours — If embeddings are so powerful, why are they mostly used only for retrieval instead of other tasks?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
If embeddings are so powerful, why are they mostly used only for retrieval instead of other tasks?
Embeddings work phenomenally well for retrieval because the problem is well-defined: you have a query, you need to find similar items from a fixed set, and similarity in embedding space correlates strongly with human judgment. You don’t need the embedding to understand anything about the world, just to cluster semantically related text close together. That’s a narrow, solvable problem.
Beyond retrieval, embeddings hit harder constraints. They compress information into fixed-dimensional vectors, which destroys detail. A 1536-dimensional embedding of a movie review loses the plot summary, character names, and whether the critic actually liked it, keeping only “vibes.” You can recover some of that by retrieving the original text, but then the embedding was just a search key, not a representation doing cognitive work.
Why Embeddings Fail at Other Tasks
Classification and tagging: if you want to classify whether a customer support ticket is about billing or shipping, an embedding alone can’t distinguish between “my bill is shipping late” (billing) and “my shipment billing is wrong” (shipping). You need to feed the embedding to a classifier head, and at that point you’re doing traditional ML, not leveraging any special property of embeddings. A frontier model like Claude Opus 5 or GPT-6 Astra just reads the text and decides directly, and it’s more accurate because it has semantic reasoning built in.
Structured prediction: embeddings are designed for ranking, not for generating outputs with constraints. If you want to extract a person’s name, age, and phone number from a messy resume, embeddings won’t help you generate those three specific fields in the right format. You either use an LLM with structured output (which parses text end-to-end), or you build a feature extraction pipeline (which is feature engineering, not embedding tricks).
Why LLMs Replaced Embeddings for Most of the Pipeline
The real shift is that frontier models got cheap enough that you can skip the embedding layer entirely. A year ago, using Claude Haiku or GPT-4.1 Nano for classification felt wasteful. Today, running GPT-4.1 Nano or Claude Sonnet 5 end-to-end on a task costs nearly the same as embedding + classifier + retrieval stacks, and it’s simpler, more accurate, and requires no training. You get interpretability for free because the model explains itself.
Look at a real comparison: a team using embeddings for document tagging would embed documents, cluster them, train a tiny supervised classifier on labeled examples, and maintain that pipeline. Now they just prompt an LLM: “Tag this document with one of [category1, category2, category3].” Done. No embedding infrastructure, no classifier training, no drift when document styles change. For a 1000-document batch, Claude Sonnet 5 at $2/$10 per M tokens costs roughly the same as a hand-rolled embedding pipeline, and it works better because the model understands context.
The Niche Where Embeddings Still Own the Problem
Retrieval persists because it’s the one task where you need to:
- Pre-compute representations once, before you know what queries will come.
- Search billions of items in milliseconds without calling an LLM (embeddings + vector index do this; LLM inference does not).
- Keep search costs sub-cent per query in high-volume settings.
A customer support system handling 100K searches a day needs vector search. Running GPT-6 Astra on every search would cost thousands. Embedding + retrieval costs cents. For any problem where you can afford to call an LLM, though, you stop using embeddings as a feature extractor and use the LLM directly.
When People Still Try (and Why It Fails)
Teams sometimes attempt to use embeddings for:
Semantic clustering without an LLM, to save costs. This works until the clusters need to encode judgment (is this user complaint about poor quality or poor shipping?). Then you’re stuck explaining why your embedding space put them together when they need different handling.
Anomaly detection in embedding space, looking for outliers. The problem: embeddings are smooth approximations of language, so outliers in embedding space often just mean “rare word combination,” not “fraud” or “security issue.” You end up with false positives and no interpretability.
Building “embedding search + LLM reranking” pipelines as a middle ground. This is common and sometimes necessary at massive scale (>100M documents), but the embedding part is now just a fast prefilter. The LLM does the real decision-making.
A Concrete Example: Why a Startup Ditched Embeddings for Direct Classification
A fintech startup was using embeddings to classify transaction descriptions (restaurant vs. grocery vs. fuel) and reranking with a fine-tuned classifier. Pipeline: embed → vector search on labeled examples → classify. Cost per transaction: $0.00003. Accuracy: 87%.
They switched to passing transaction descriptions directly to GPT-4.1 Nano with a system prompt: “Classify this transaction into one of [restaurant, grocery, fuel]. Respond with only the category name.” Cost per transaction: $0.00004. Accuracy: 94%. Latency dropped because they eliminated the embedding server. They killed the whole pipeline.
The embedding approach made sense in 2022 when LLM inference was expensive and slow. Now? LLMs are the default tool, and embeddings are the specialized trick you reach for when you have a billion items and millisecond latency requirements.
Bottom line: Use embeddings for retrieval at scale, where pre-computed vectors and vector indices are the only way to search billions of items cheaply and fast. For everything else—classification, extraction, tagging, reranking—use an LLM directly; it’s cheaper, simpler, and more accurate than embedding + ML classifier stacks.
Question via Hacker News