Library of the Week — llm
A weekly teardown of one open-source AI/ML library: what it does, why it stands out, and when to use it.
llm — a CLI tool and Python library for running prompts against LLMs, locally and via API
GitHub · Language: Python · License: Apache 2.0
What it does
Simon Willison’s llm gives you a single, unified interface for interacting with dozens of LLM providers and local models — from the command line or from Python. It solves the “I just want to run a prompt” problem without wiring up provider SDKs, managing auth boilerplate, or spinning up a server. It’s aimed at developers who work across multiple models and want logging, templating, and scripting built in from day one.
Why it stands out
- Plugin architecture that actually works: providers like Anthropic, Gemini, Ollama, and Mistral are first-class plugins installable with one
pip installcommand, and the plugin API is stable enough that the ecosystem has grown substantially - Automatic SQLite logging: every prompt and response is stored locally by default, queryable later — invaluable for debugging and cost auditing without any extra instrumentation
- Prompt templates baked in: named, reusable templates with variable substitution live on disk, so you can standardize prompts across projects and invoke them by name from the CLI
- Conversation continuity:
llm -ccontinues the previous conversation, threading context automatically — genuinely useful for interactive debugging sessions at the terminal
Quick start
pip install llm
pip install llm-anthropic # or llm-gemini, etc.
llm keys set anthropic # store API key once
# one-shot prompt
llm "Summarize this in one sentence: $(cat notes.txt)"
# continue a conversation
llm "What were the main themes?" -c
# use a saved template
llm --system "You are a terse code reviewer." \
"Review this function" < myfile.py
import llm
model = llm.get_model("claude-sonnet-5")
response = model.prompt("Explain tail call optimization")
print(response.text())
When to use it
- You frequently switch between providers (Claude Fable 5.1, GPT-6 Astra, local Qwen3.6) and want a single interface rather than maintaining separate SDK integrations
- You want persistent, searchable logs of every LLM call across projects without building your own observability layer
- You’re building shell scripts or lightweight automation around LLM calls and need something composable with standard Unix pipes
When to skip it
- You need streaming-first, production-grade async throughput —
llmis optimized for developer ergonomics, not high-concurrency inference pipelines; reach for a proper inference client or gateway at scale - If your team’s primary need is agent orchestration with tool-calling graphs, you’ll quickly outgrow
llm’s scope and want something like LangGraph instead
The verdict
llm is one of those tools that quietly becomes load-bearing in a developer’s workflow — you install it to try one prompt and six months later you’re piping everything through it. It’s exceptionally well-designed for the “daily driver” use case: fast experimentation, cross-provider comparison, and lightweight scripting. If you work with multiple LLM providers and aren’t already using it, install it today.