The Open-Source AI Inference Stack: Ollama, vLLM, llama.cpp, LiteLLM, and the Tooling That Makes Self-Hosted Inference Practical
How the open-source AI inference stack fits together: Ollama, vLLM, llama.cpp, LiteLLM, and the tooling that makes self-hosted models practical in 2026.
How the open-source AI inference stack fits together: Ollama, vLLM, llama.cpp, LiteLLM, and the tooling that makes self-hosted models practical in 2026.
The Open-Source AI Inference Stack: Ollama, vLLM, llama.cpp, LiteLLM, and the Tooling That Makes Self-Hosted Inference Practical
Self-hosted inference has moved from hobbyist curiosity to production necessity. API costs at scale compound fast — a mid-volume application making 10M requests/month against a frontier API can spend $50K–$200K monthly on tokens alone. Open-weight models like Mistral Large 3, Qwen3.6, DeepSeek-V4, and Llama 4 now match or approach the performance of proprietary models on many tasks, which means the “quality gap” argument for API-only architectures has narrowed. The open-source inference stack has matured to fill that gap with real production tooling.
This post covers the full stack: the inference engines that run models on hardware, the serving layers that expose them as APIs, the compatibility shims that let applications treat self-hosted models like OpenAI endpoints, and the operational tooling that makes all of it manageable. The focus is on what each tool actually does, where they overlap, where they don’t, and how to choose between them.
Table of Contents
- The Stack at a Glance
- llama.cpp: The Foundation Layer
- Ollama: Local Inference Made Trivial
- vLLM: Production-Grade GPU Serving
- Other Inference Engines Worth Knowing
- LiteLLM: The Unified API Layer
- Quantization: The Practical Reality
- Hardware Mapping: What Runs Where
- Putting It Together: Architecture Patterns
- Operational Concerns
- When to Self-Host vs When to Use an API
- Summary
- Further Reading
The Stack at a Glance
The open-source inference stack has four functional layers. Most deployments use one tool from each layer, though some tools span multiple layers.
The four layers of self-hosted inference: hardware, engine, serving API, and routing proxy.
Each layer has multiple options with different tradeoffs. The key insight is that these tools are composable — llama.cpp handles inference on CPU/Apple Silicon, vLLM handles GPU serving at scale, and LiteLLM sits in front of either (or both, plus external APIs) to give applications a single endpoint.
llama.cpp: The Foundation Layer
llama.cpp started as Georgi Gerganov’s C/C++ port of Meta’s LLaMA inference code in March 2023. It has since become the de facto standard for running quantized models on consumer hardware and the foundation that several other tools (including Ollama) build on.
What It Actually Does
llama.cpp implements transformer inference in pure C/C++ with optional hardware acceleration. No Python runtime, no PyTorch dependency, no CUDA requirement. It runs on:
- CPUs (x86, ARM) using AVX2/AVX-512/AMX instructions
- Apple Silicon using Metal for GPU acceleration
- NVIDIA GPUs using CUDA
- AMD GPUs using ROCm/HIP
- Intel GPUs using SYCL
- Vulkan for cross-platform GPU support
The primary model format is GGUF (GGML Universal Format), a single-file binary format that packages model weights, tokenizer, and metadata together. GGUF files are self-contained — no separate config files, no tokenizer JSONs, no Python scripts needed to load them.
llama-server
llama.cpp ships with llama-server, a built-in HTTP server that exposes an OpenAI-compatible chat completions API. This is the simplest path from “I have a GGUF file” to “I have an API endpoint.”
# Serve a model with GPU offloading
llama-server \
-m mistral-small-4-Q4_K_M.gguf \
--host 0.0.0.0 \
--port 8080 \
-ngl 99 \ # offload all layers to GPU
-c 8192 \ # context length
--n-predict 4096 # max output tokens
The server supports streaming (SSE), multiple concurrent requests (with --parallel N), and basic features like grammar-based constrained decoding. It does not support continuous batching at the level vLLM does — concurrent requests share a KV cache pool but don’t get the same level of batch scheduling optimization.
Key Quantization Formats
llama.cpp pioneered practical quantization for inference. The naming convention encodes the method:
| Format | Bits/Weight | Typical Quality | Use Case |
|---|---|---|---|
| Q8_0 | 8.0 | Near-lossless | When VRAM allows |
| Q6_K | 6.6 | Excellent | Sweet spot for large models |
| Q5_K_M | 5.5 | Very good | Balanced quality/size |
| Q4_K_M | 4.8 | Good | Most common default |
| Q4_0 | 4.0 | Acceptable | Tight VRAM budgets |
| Q3_K_M | 3.9 | Noticeable degradation | Only when necessary |
| Q2_K | 2.6 | Poor | Last resort |
| IQ4_XS | ~4.3 | Good (importance-matrix) | Better than Q4_0 at similar size |
The K suffix indicates k-quant methods (different precision per layer based on sensitivity), and _M indicates a medium allocation strategy. Importance-matrix quants (IQ*) use calibration data to assign more bits to important weights — they achieve better quality at the same size but take longer to produce.
The llama.cpp pipeline: convert HuggingFace weights to GGUF, quantize, serve.
When to Use llama.cpp Directly
- Running models on Apple Silicon (llama.cpp’s Metal backend is the most mature option)
- CPU-only inference (no GPU available)
- Single-model deployments where simplicity matters more than throughput
- Edge or embedded deployments where minimizing dependencies is critical
- When the model is already available as a GGUF file (most popular models are, via HuggingFace)
Ollama: Local Inference Made Trivial
Ollama wraps llama.cpp (and increasingly other backends) in a Docker-like UX for model management. It reduces the “download a model and run it” workflow to two commands.
What It Adds Over Raw llama.cpp
- Model registry and pull semantics:
ollama pull mistral-smalldownloads a quantized model from Ollama’s library, handles format details, and stores it locally - Modelfile abstraction: declarative configuration for system prompts, parameters, quantization level, and template format
- Automatic hardware detection: detects available GPUs and Apple Silicon, configures offloading without manual flags
- Background daemon: runs as a service, manages model loading/unloading, handles concurrent requests
- OpenAI-compatible API: serves at
http://localhost:11434/v1/chat/completions
# Install and run — that's it
curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen3.6:27b-q4_K_M
ollama run qwen3.6:27b-q4_K_M
Modelfile Example
FROM qwen3.6:27b-q4_K_M
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
PARAMETER top_p 0.9
SYSTEM """You are a helpful coding assistant. Respond with concise, working code."""
ollama create my-coding-assistant -f Modelfile
ollama run my-coding-assistant
Ollama’s API
Ollama exposes two API styles: its native API and an OpenAI-compatible endpoint. The OpenAI-compatible endpoint means most tools that work with OpenAI’s API work with Ollama out of the box — LangChain, LlamaIndex, Continue (VS Code), and any application using the OpenAI Python/Node SDK.
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama" # required by SDK, ignored by Ollama
)
response = client.chat.completions.create(
model="qwen3.6:27b-q4_K_M",
messages=[{"role": "user", "content": "Explain B-trees in 3 sentences."}],
temperature=0.7
)
Limitations
Ollama optimizes for developer experience, not production throughput. Specific gaps:
- No continuous batching: concurrent requests queue rather than batch efficiently — throughput under load is substantially lower than vLLM
- Limited multi-GPU: supports GPU offloading but not the tensor parallelism needed for models that don’t fit on a single GPU
- No speculative decoding: llama.cpp supports it, but Ollama doesn’t expose it
- Model library lag: new models appear on HuggingFace days or weeks before they’re available in Ollama’s registry (though GGUF files can be imported directly)
Ollama wraps llama.cpp, adding model management and a daemon, then exposes an OpenAI-compatible API.
When to Use Ollama
- Local development and prototyping against open models
- Single-developer or small-team setups where managing llama.cpp flags is unwanted overhead
- IDE integration (Continue, Cursor-compatible setups, Cody)
- Running on laptops or workstations with 16–64GB RAM/VRAM
Ollama is not the right choice for production serving at scale. That’s vLLM’s territory.
vLLM: Production-Grade GPU Serving
vLLM is an inference engine built for throughput. Its core contribution is PagedAttention — a memory management technique that treats the KV cache like virtual memory pages, eliminating the fragmentation that wastes 60–80% of GPU memory in naive implementations.
Why PagedAttention Matters
In standard transformer inference, each request needs a contiguous block of GPU memory for its KV cache, sized to the maximum possible sequence length. A model serving 2048-token sequences allocates KV cache for 2048 tokens per request even if most responses are 200 tokens. Unused memory sits wasted.
PagedAttention breaks the KV cache into fixed-size pages (typically 16 tokens each) and allocates them on demand, like a virtual memory system. Pages can be non-contiguous in physical GPU memory. This achieves near-optimal memory utilization and allows 2–4x more concurrent requests on the same hardware.
PagedAttention replaces contiguous pre-allocated KV cache with paged on-demand allocation, cutting waste from 60–80% to under 5%.
Core Features
- Continuous batching: new requests join an in-flight batch without waiting for the current batch to finish — keeps GPUs saturated
- Tensor parallelism: splits a single model across multiple GPUs (essential for 70B+ models)
- Pipeline parallelism: chains multiple GPU groups for very large models
- Speculative decoding: uses a small draft model to propose tokens, verified in parallel by the main model — can increase throughput 1.5–2x
- Quantization: supports GPTQ, AWQ, SqueezeLLM, and FP8 quantized models
- OpenAI-compatible API: drop-in replacement for OpenAI endpoints
- LoRA serving: hot-swap LoRA adapters at request time without reloading the base model
- Prefix caching: automatically reuses KV cache for shared prompt prefixes across requests
Running vLLM
pip install vllm
# Serve a model — single command
vllm serve Qwen/Qwen3.6-27B \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--gpu-memory-utilization 0.9 \
--port 8000
The vllm serve command downloads the model from HuggingFace (if not cached), loads it onto GPUs, and starts an OpenAI-compatible server. For models that need quantization:
vllm serve Qwen/Qwen3.6-27B-AWQ \
--quantization awq \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--port 8000
vLLM vs llama.cpp: The Tradeoff
| Dimension | vLLM | llama.cpp |
|---|---|---|
| Primary target | NVIDIA GPUs (datacenter) | CPU, Apple Silicon, consumer GPU |
| Batching | Continuous batching | Basic parallel requests |
| Multi-GPU | Tensor + pipeline parallelism | Layer offloading only |
| Model format | HuggingFace (safetensors) | GGUF |
| Quantization | GPTQ, AWQ, FP8 | GGML quants (Q4_K_M, etc.) |
| Memory efficiency | PagedAttention (~95% utilization) | Good but lower under load |
| Throughput at scale | 10–50x higher under concurrent load | Adequate for single/few users |
| Dependencies | Python, CUDA, PyTorch | C/C++ only |
| Setup complexity | Moderate | Low (single binary) |
| CPU inference | Not designed for it | Excellent |
| Apple Silicon | Not supported | Best-in-class |
The decision is usually clear: if the deployment target is NVIDIA datacenter GPUs serving many concurrent users, use vLLM. If it’s a Mac, a CPU server, or a single-user workstation, use llama.cpp (directly or via Ollama).
Other Inference Engines Worth Knowing
Text Generation Inference (TGI)
HuggingFace’s Rust-based inference server. Supports continuous batching, tensor parallelism, quantization (GPTQ, AWQ, EETQ, bitsandbytes), and speculative decoding. Used as the backend for HuggingFace’s Inference Endpoints service.
TGI’s main advantage is tight HuggingFace ecosystem integration — it pulls models directly from the Hub and supports the widest range of HuggingFace model architectures. Its throughput is competitive with vLLM but benchmarks vary by model and workload. TGI is licensed under Apache 2.0 as of 2024.
SGLang
A serving framework that emphasizes structured generation and complex prompting patterns. SGLang (Structured Generation Language) treats prompt construction as a program, enabling efficient KV cache reuse across branching prompt structures. Strong on workloads that involve constrained decoding, multiple generation calls per request, or tree-of-thought patterns. Throughput is competitive with vLLM on standard workloads and can exceed it on structured/branching workloads.
ExLlamaV2
Optimized specifically for GPTQ and EXL2 quantized models on consumer NVIDIA GPUs. Achieves the highest tokens/second for a single user on a single GPU, particularly at 4-bit and 3-bit quantization levels. Less suited for multi-user serving but excellent for single-user setups on RTX 3090/4090 hardware.
The inference engine landscape: vLLM and TGI compete for datacenter GPU serving, SGLang optimizes for structured generation, ExLlamaV2 and llama.cpp target consumer/edge hardware.
LiteLLM: The Unified API Layer
LiteLLM solves the “100 providers, 100 APIs” problem. It provides a single OpenAI-compatible interface that routes to any backend — self-hosted vLLM, local Ollama, OpenAI, Anthropic, Google, Bedrock, Vertex, and 100+ others.
Why This Matters
Applications that use self-hosted models almost always also need access to external APIs for certain tasks. A coding assistant might use a local Qwen3.6-27B for routine completions and route complex architectural questions to Claude Opus 5. Without a unified layer, the application needs separate client code, error handling, and retry logic for each provider.
LiteLLM eliminates this. The application calls one endpoint with one SDK, and LiteLLM handles routing, retries, and format translation.
Proxy Mode
LiteLLM’s proxy server is the primary production deployment pattern. It runs as a standalone service with a YAML configuration:
# litellm_config.yaml
model_list:
- model_name: fast-local
litellm_params:
model: ollama/qwen3.6:27b-q4_K_M
api_base: http://localhost:11434
- model_name: fast-local
litellm_params:
model: openai/qwen3.6-27b
api_base: http://vllm-server:8000/v1
api_key: none
- model_name: frontier
litellm_params:
model: anthropic/claude-opus-5
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: frontier
litellm_params:
model: openai/gpt-6-astra
api_key: os.environ/OPENAI_API_KEY
router_settings:
routing_strategy: least-busy
num_retries: 3
fallbacks:
- fast-local: [frontier]
litellm --config litellm_config.yaml --port 4000
This configuration does several things:
- Two instances of
fast-local(Ollama and vLLM) — LiteLLM load-balances between them - Two instances of
frontier(Claude Opus 5 and GPT-6 Astra) — load-balanced - If
fast-localfails, requests fall back tofrontier - Retry logic handles transient errors
LiteLLM as a unified proxy: applications hit one endpoint, LiteLLM routes to local or cloud backends with fallback.
Features That Matter in Production
- Cost tracking: logs token counts and costs per request, per model, per user (even for self-hosted models, with configurable per-token pricing)
- Rate limiting: per-user, per-model token and request limits
- Spend budgets: cap spending per API key or team
- Caching: optional Redis-backed response caching
- Guardrails integration: hooks for content filtering before/after generation
- Virtual keys: issue API keys to teams/users, each with their own model access, rate limits, and budgets
SDK Mode
For simpler setups or direct library use:
from litellm import completion
# Calls local vLLM
response = completion(
model="openai/qwen3.6-27b",
messages=[{"role": "user", "content": "Explain RAFT consensus."}],
api_base="http://localhost:8000/v1",
api_key="none"
)
# Same interface, calls Anthropic
response = completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Review this architecture."}]
)
Quantization: The Practical Reality
Quantization is what makes large models runnable on affordable hardware. The core idea: store model weights in lower precision (4-bit, 8-bit) instead of the original 16-bit or 32-bit, reducing memory requirements proportionally.
Memory Math
A model’s memory footprint in FP16 is roughly parameters × 2 bytes. A 27B parameter model needs ~54GB in FP16. At Q4_K_M (approximately 4.8 bits/weight), that drops to ~16GB — fitting on a single 24GB GPU with room for KV cache.
| Model Size | FP16 | Q8_0 | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 7–8B | 16GB | 8GB | 5GB | 4GB |
| 14B | 28GB | 14GB | 9GB | 7GB |
| 27B | 54GB | 27GB | 16GB | 13GB |
| 70B | 140GB | 70GB | 40GB | 33GB |
These are weight-only numbers. Actual runtime memory includes the KV cache (proportional to context length × batch size), activations, and CUDA/framework overhead. For a single-user setup with 8K context, add 1–3GB. For concurrent serving, KV cache can dominate memory usage.
GGUF vs GPTQ vs AWQ
The quantization ecosystem has settled into two camps:
GGUF (llama.cpp ecosystem): Quantization happens at inference time or via the llama-quantize tool. Supports mixed-precision (different layers get different bit widths). Runs on CPU, Apple Silicon, and GPU. The Q4_K_M and Q5_K_M variants are the most popular defaults. Importance-matrix quants (IQ4_XS, IQ3_XXS) push quality further at a given size.
GPTQ / AWQ (GPU ecosystem): Calibration-based methods that require a calibration dataset during quantization. Generally produce slightly better quality than GGUF at the same bit width, but only run on GPUs with CUDA. Used with vLLM, TGI, ExLlamaV2, and the HuggingFace Transformers library.
Three quantization families: GGUF for universal hardware, GPTQ/AWQ for GPU serving, FP8 for datacenter GPUs with native support.
FP8: NVIDIA H100 and newer GPUs have native FP8 support. FP8 quantization cuts memory in half vs FP16 with minimal quality loss — it’s effectively “free” compression on supported hardware. vLLM supports FP8 natively.
Quality Impact
The quality vs quantization tradeoff is nonlinear and model-dependent. General patterns observed across many models:
- Q8 → Q6: Quality loss is imperceptible on almost all benchmarks
- Q6 → Q5_K_M: Slight degradation, usually <1% on standard benchmarks
- Q5_K_M → Q4_K_M: Noticeable on reasoning-heavy tasks, usually 1–3% benchmark drop
- Q4_K_M → Q3_K_M: Measurable quality loss, 3–5% on reasoning, significant on math
- Below Q3: Quality degrades rapidly; not recommended for tasks requiring accuracy
For coding tasks, Q4_K_M is generally the lowest practical quantization. For creative writing and conversation, Q3_K_M is often acceptable. For math and complex reasoning, Q5_K_M or higher is advisable.
Hardware Mapping: What Runs Where
Apple Silicon
Apple’s unified memory architecture makes it uniquely suited for running large models — the GPU can access all system memory, so a 96GB M2 Max or M4 Max can run a 70B Q4_K_M model that wouldn’t fit on any consumer GPU’s VRAM.
| Chip | Memory | Practical Max Model | Speed (tok/s, Q4_K_M) |
|---|---|---|---|
| M2/M3/M4 (8GB) | 8GB | 7–8B | 15–25 |
| M2/M3/M4 Pro (18GB) | 18GB | 14B | 20–30 |
| M2/M3/M4 Max (32–64GB) | 32–64GB | 27–35B | 15–30 |
| M2/M3/M4 Max (96–128GB) | 96–128GB | 70B | 10–20 |
| M2/M3/M4 Ultra (192GB) | 192GB | 100B+ | 10–20 |
Token generation speed on Apple Silicon is limited by memory bandwidth, not compute. The M4 Max with its ~546 GB/s bandwidth generates tokens approximately 1.5–2x faster than the M2 Max at ~400 GB/s.
Tool choice on Apple Silicon: llama.cpp (directly or via Ollama) using the Metal backend. vLLM does not support Apple Silicon.
Consumer NVIDIA GPUs
| GPU | VRAM | Practical Max Model (Q4_K_M) | Notes |
|---|---|---|---|
| RTX 3060 12GB | 12GB | 7–8B | Budget entry point |
| RTX 3090 | 24GB | 14B (comfortable), 27B (tight) | Still the price/performance king for used market |
| RTX 4090 | 24GB | 14B (comfortable), 27B (tight) | Fastest single consumer GPU |
| RTX 5090 | 32GB | 27B (comfortable) | Current generation |
| 2× RTX 3090 | 48GB | 35B (comfortable), 70B (tight) | Requires multi-GPU support |
For consumer GPUs: llama.cpp, ExLlamaV2, or Ollama for single-user. vLLM works on RTX 3090/4090/5090 but its overhead means it’s only worthwhile if serving multiple concurrent users.
Datacenter GPUs
| GPU | VRAM | FP16 Models | Quantized Models | Best For |
|---|---|---|---|---|
| A100 40GB | 40GB | 27B | 70B (Q4) | Legacy deployments |
| A100 80GB | 80GB | 35B | 70B (Q5/Q6) | Workhorse |
| H100 80GB | 80GB | 35B | 70B (FP8) | Current standard |
| L40S | 48GB | 27B | 70B (Q3, tight) | Inference-optimized |
| 2× H100 | 160GB | 70B (FP16) | 100B+ | Tensor parallel |
| 8× H100 | 640GB | 400B+ | Large MoE models | Multi-node |
For datacenter GPUs: vLLM is the default choice. TGI and SGLang are alternatives worth benchmarking for specific workloads.
Hardware tiers map to different tools: Apple Silicon uses llama.cpp, consumer GPUs use llama.cpp or ExLlamaV2, datacenter GPUs use vLLM.
Putting It Together: Architecture Patterns
Pattern 1: Solo Developer / Small Team
The simplest viable setup. One machine, one model, direct access.
Single-machine setup: Ollama serves a model, IDE and application code connect directly.
- Hardware: Mac with 32GB+ RAM, or workstation with RTX 3090/4090
- Software: Ollama running one or two models
- Model: Qwen3.6-27B-Q4_K_M or Mistral Small 4 for general tasks; Devstral 2 for coding
- Cost: Hardware only (one-time)
- Throughput: Single user, 20–40 tok/s
Pattern 2: Team Serving (5–20 developers)
Shared inference server behind a proxy.
Team serving: vLLM on dedicated GPUs, LiteLLM in front for routing and key management, cloud APIs as fallback.
- Hardware: 2× A100 80GB or 2× H100 (cloud instances or on-prem)
- Software: vLLM serving 1–2 models, LiteLLM proxy for routing
- Models: 70B model at FP8 for primary workloads, frontier API for complex tasks
- Cost: ~$5–10/hour for GPU instances + API costs for fallback
- Throughput: 10–20 concurrent users at acceptable latency
Pattern 3: Production Application
Full stack with redundancy, monitoring, and cost optimization.
Production deployment: load balancer → LiteLLM cluster → vLLM pool (primary) and cloud APIs (burst/fallback), with observability.
- Hardware: 4–8 GPUs in a pool, potentially across multiple nodes
- Software: vLLM with tensor parallelism, LiteLLM proxy cluster (multiple instances behind a load balancer), Langfuse or similar for observability
- Models: Primary self-hosted model for 80–90% of traffic, cloud APIs for burst capacity and tasks requiring frontier capabilities
- Cost: GPU infrastructure + minimal API spend for overflow
- Architecture decisions:
- Route simple tasks (summarization, extraction, classification) to self-hosted models
- Route complex tasks (multi-step reasoning, code review) to frontier APIs
- Use LiteLLM’s router to implement this logic based on model name or request metadata
Pattern 4: Multi-Model Router
For applications that benefit from using different models for different task types.
# LiteLLM config for multi-model routing
model_list:
# Fast, cheap: classification, extraction, simple Q&A
- model_name: fast
litellm_params:
model: openai/mistral-small-4
api_base: http://vllm-small:8000/v1
# Balanced: general coding, analysis
- model_name: balanced
litellm_params:
model: openai/qwen3.6-27b
api_base: http://vllm-medium:8000/v1
# Frontier: complex reasoning, architecture review
- model_name: frontier
litellm_params:
model: anthropic/claude-sonnet-5
# Max capability: when nothing else works
- model_name: max
litellm_params:
model: anthropic/claude-opus-5
The application selects the model name based on task type. A classification request goes to fast, a code review goes to balanced or frontier, and a complex multi-file refactor goes to max. This pattern typically reduces costs by 60–80% compared to routing everything to a frontier model.
Operational Concerns
Model Updates and Versioning
Open-weight models update frequently. Qwen3.6 succeeds Qwen3.5, Mistral Small 4 succeeds Small 3, and each version has different quantizations from different providers on HuggingFace. Production deployments need a strategy.
Practical approach:
- Pin specific model versions (e.g.,
Qwen/Qwen3.6-27Bat a specific commit hash, or a specific GGUF filename) - Test new versions against an eval suite before deployment
- Use LiteLLM’s model aliasing —
model_name: balancedcan point to different backends without application changes - Keep the previous model loaded for quick rollback
Health Checks and Monitoring
vLLM exposes a /health endpoint. Ollama has /api/tags (returns loaded models) and responds to chat requests with model-not-found errors if the model isn’t loaded. LiteLLM has /health that checks all configured backends.
Key metrics to monitor:
- Time to first token (TTFT): measures model loading + prompt processing latency
- Tokens per second (TPS): generation throughput
- Queue depth: number of waiting requests (vLLM reports this)
- GPU utilization and VRAM usage: approaching VRAM limits causes OOM crashes
- KV cache utilization: vLLM reports this; high KV cache usage means new requests may queue
# Prometheus metrics from vLLM (exposed at /metrics)
# Key metrics:
# vllm:num_requests_running — current batch size
# vllm:num_requests_waiting — queue depth
# vllm:gpu_cache_usage_perc — KV cache utilization
# vllm:avg_generation_throughput_toks_per_s — throughput
Context Length vs Memory
Context length is a direct tradeoff with concurrent users. The KV cache for a single 32K-context request on a 27B model at FP16 attention consumes roughly 2–4GB of VRAM. Serving 10 concurrent 32K-context requests requires 20–40GB just for KV cache, on top of the model weights.
Practical limits:
- Set
--max-model-lenin vLLM to the actual maximum context needed, not the model’s theoretical maximum - If the application only uses 4K context, setting the limit to 4K instead of 128K dramatically increases concurrent capacity
- vLLM’s PagedAttention helps but doesn’t eliminate this fundamental tradeoff
Container Deployment
Both vLLM and Ollama have official Docker images. vLLM’s image includes CUDA and all dependencies:
# docker-compose.yml for vLLM + LiteLLM
services:
vllm:
image: vllm/vllm-openai:latest
command: >
--model Qwen/Qwen3.6-27B
--tensor-parallel-size 2
--max-model-len 16384
--gpu-memory-utilization 0.9
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 2
capabilities: [gpu]
volumes:
- model-cache:/root/.cache/huggingface
litellm:
image: ghcr.io/berriai/litellm:main-latest
command: --config /app/config.yaml --port 4000
volumes:
- ./litellm_config.yaml:/app/config.yaml
ports:
- "4000:4000"
depends_on:
- vllm
volumes:
model-cache:
Security Considerations
Self-hosted models eliminate the concern of sending proprietary data to external APIs. But they introduce new attack surfaces:
- Model endpoint exposure: vLLM and Ollama endpoints have no authentication by default — put them behind a reverse proxy or use LiteLLM’s key management
- Model supply chain: GGUF and safetensors files from HuggingFace are generally safe (safetensors was designed to prevent arbitrary code execution), but pickle-based model formats (older PyTorch
.binfiles) can contain malicious code - Prompt injection: self-hosted models are as vulnerable to prompt injection as API models — the same guardrails are needed
- Resource exhaustion: a long-context request can consume all VRAM and crash the server — set
--max-model-lenand consider request timeouts
When to Self-Host vs When to Use an API
The decision framework comes down to four variables: volume, latency requirements, data sensitivity, and capability needs.
Decision factors: volume drives cost math, data sensitivity forces self-hosting, frontier capability needs favor APIs.
Self-host when:
- Token volume exceeds ~$2K–5K/month in API costs (the breakeven depends on GPU rental costs, but at this scale self-hosting typically saves 50–80%)
- Data cannot leave the organization (regulated industries, proprietary code, PII)
- Latency requirements demand co-location (inference latency is 5–20ms for local, 100–500ms for API round-trips)
- The task is well-served by a 7B–70B model (classification, extraction, summarization, code completion, simple Q&A)
Use APIs when:
- Tasks require frontier reasoning capabilities (complex multi-step reasoning, novel problem-solving, advanced code architecture)
- Volume is low enough that API costs are under the operational overhead of running infrastructure
- Rapid model updates matter (API providers deploy new models instantly; self-hosted requires re-download and revalidation)
- No in-house GPU infrastructure or ML ops expertise
Hybrid (the most common production pattern):
- Self-host a 14B–27B model for 80% of requests (fast, cheap, private)
- Route the remaining 20% to frontier APIs (Claude Opus 5, GPT-6 Astra) for tasks that require maximum capability
- Use LiteLLM to make this transparent to the application
The cost math example: serving Qwen3.6-27B-Q4_K_M on a single A100 80GB cloud instance costs roughly $1.50–2.50/hour ($1,100–1,800/month). That same instance can serve ~500K tokens/minute. At Claude Sonnet 5 pricing ($2/$10 per M tokens), 500K tokens/minute would cost approximately $2,160–7,200/month in API fees at moderate input/output ratios. Self-hosting breaks even at roughly 30–50% utilization.
Summary
The open-source inference stack in 2026 is production-ready for most workloads that don’t require frontier-model capabilities.
llama.cpp is the universal inference engine — runs everywhere, uses GGUF format, best choice for CPU, Apple Silicon, and single-user GPU deployments. Use it directly via llama-server or through Ollama.
Ollama wraps llama.cpp in a developer-friendly package. Ideal for local development, IDE integration, and small-team use. Not designed for high-concurrency production serving.
vLLM is the production GPU serving engine. PagedAttention, continuous batching, tensor parallelism, and speculative decoding make it the clear choice for serving open-weight models at scale on NVIDIA GPUs.
LiteLLM unifies everything behind a single OpenAI-compatible API. Its proxy mode adds routing, fallback, cost tracking, and key management — the glue layer that makes hybrid self-hosted + API architectures practical.
Quantization is not optional for most deployments. Q4_K_M is the practical default for most tasks. Q5_K_M or Q6_K for quality-sensitive workloads. FP8 on H100/B200 hardware gives near-lossless compression for free.
The dominant architecture pattern is hybrid: self-host a capable open model (Qwen3.6-27B, Mistral Small 4, or similar) for the bulk of traffic, route complex tasks to frontier APIs, and use LiteLLM to make the routing transparent to application code. This typically reduces costs 60–80% compared to API-only architectures while maintaining access to frontier capabilities when needed.
Further Reading
- llama.cpp GitHub repository — The C/C++ inference engine, GGUF format spec, quantization tools, and llama-server documentation
- vLLM documentation — Official docs covering PagedAttention, serving configuration, quantization, and tensor parallelism
- Ollama GitHub repository — Model library, Modelfile reference, and API documentation
- LiteLLM documentation — Proxy configuration, supported providers, router settings, and virtual key management
- HuggingFace Text Generation Inference — TGI source code and deployment guides
- SGLang GitHub repository — Structured generation framework with serving backend
- ExLlamaV2 GitHub repository — Optimized GPTQ/EXL2 inference for consumer GPUs
- GGUF format specification — Technical specification for the GGUF model format
- vLLM PagedAttention paper — “Efficient Memory Management for Large Language Model Serving with PagedAttention” (Kwon et al., 2023)
- TheBloke on HuggingFace — Largest collection of pre-quantized GGUF and GPTQ models (many community quantizers have continued this work)