Office Hours — How do you design persistent memory systems that work effectively across multiple AI agents?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
How do you design persistent memory systems that work effectively across multiple AI agents?
This is the question everyone’s pretending they’ve solved. The reality is messier.
The core problem is simple: you need agents to accumulate knowledge across tasks, share context efficiently, and avoid hallucinating or forgetting critical state. What makes this hard is that every architectural choice introduces tradeoffs that compound at scale.
The Straightforward Approach: Shared Vector Store
Most teams start here. You throw a vector database (Pinecone, Weaviate, Milvus) behind your agents and call it “persistent memory.” Each agent writes summaries or raw data, other agents retrieve via semantic search. It feels clean.
The first failure mode hits within weeks. Agent A retrieved a document about “Q3 budget allocation,” Agent B independently retrieved the same document, but they disagree on the exact number. One got the cached embedding, one got a fresh retrieval. Neither can reconcile. You’ve traded consistency for convenience.
The second failure is worse. Your agents aren’t actually learning over time—they’re pattern-matching against a static embedding space. If Agent A solved a problem by calling a specific API with specific parameters, Agent B doesn’t learn “use this API with these params.” It learns “documents similar to this one.” The difference is the gap between actual knowledge transfer and statistical coincidence.
Graph-Based Memory: Structure Without Brittleness
Better teams move to a hybrid approach: a graph database backing agent memory, with nodes representing entities, relationships, and decisions. Claude Cowork uses this pattern natively—agents can share typed context across terminals on the same machine.
The win here is specificity. You encode “Agent A called weather_api with lat=37.7749, lon=-122.4194” as a relationship. Agent B can look up “what APIs worked for location queries?” and get exact parameters, not semantic matches.
The cost is operational. Graph databases require schema design. Every agent needs to understand your ontology. When you get it wrong, you end up with inconsistent node types or relationships that agents can’t interpret. You’re trading flexibility for correctness, and if your domain isn’t stable, you’ll refactor frequently.
A practical middle ground: use Claude Opus 5 or GPT-6 Astra as a context compiler. The agent keeps a flat log of decisions. Before each task, a frontier model reads the log, extracts relevant structured facts, and builds a clean context window. This defers consistency requirements to the model’s reasoning, which is often more robust than brittle schema enforcement.
The Token Economy Problem
Here’s what nobody mentions in talks about multi-agent memory: shared context gets expensive fast.
Say you have three agents. Agent A builds a knowledge base of 50 KB of relevant facts. Agent B needs some of those facts, so it retrieves them. That’s an extra 50 KB of input tokens. Agent C retrieves an overlapping 40 KB. Now you’re paying for redundant context across every task.
Nvidia’s SoL-Pi system (mentioned in recent Daily Signal coverage) cuts agent token consumption nearly in half by optimizing the control layer—what context gets passed to the model versus what stays in the environment. The principle applies here: explicit memory structures should hold facts, but only relevant subsets should enter the LLM’s context window.
A working pattern: implement a two-level memory hierarchy. Level 1 is fast, cheap, and lossy: a Redis cache of recent decisions and summaries, indexed by agent ID and task type. Level 2 is durable and rich: a graph or vector store, queried only when the cache misses or when doing long-horizon planning. This keeps token costs proportional to actual information need, not total memory size.
Cross-Agent Coordination Without Hallucination
The hardest problem is this: when Agent A writes “we used API endpoint /v2/allocate,” how do you ensure Agent B interprets it the same way three weeks later, after an API migration to /v3?
Don’t rely on agents to maintain version consistency. Version your facts explicitly. Every memory entry should include a timestamp and a schema version. Before retrieving facts for a task, check whether they’re still valid. This sounds obvious, but most teams skip it and end up with agents confidently using deprecated information.
For higher-stakes coordination (like multi-agent coding workflows on the same codebase), give agents read-only access to a canonical source of truth: Git history, a design document, or a running test suite. They can reference it, not rewrite it from memory. Claude Code instances can now message each other and share context across terminals, which works because they’re all referencing the same filesystem—single source of truth.
A Concrete Example
Here’s a minimal persistent memory system that actually works:
class AgentMemory:
def __init__(self, agent_id: str, graph_db, cache_ttl=3600):
self.agent_id = agent_id
self.graph = graph_db
self.cache = {}
self.cache_ttl = cache_ttl
def record_decision(self, task_id: str, decision: dict, params: dict):
# Write to both cache and durable storage
key = f"{self.agent_id}:{task_id}"
self.cache[key] = (decision, time.time())
# Graph: (Agent) -[made_decision]-> (Decision) -[used_params]-> (Params)
self.graph.add_node(f"decision_{task_id}",
type="decision",
value=decision,
schema_version=1)
self.graph.add_edge(f"agent_{self.agent_id}",
f"decision_{task_id}")
def retrieve_relevant(self, query: str, k=5):
# Check cache first
recent = [v for k, (v, t) in self.cache.items()
if time.time() - t < self.cache_ttl]
if recent:
return recent[:k] # Fast path
# Graph search: find similar decisions by type and param overlap
return self.graph.query(query, limit=k)
This isn’t production-grade, but it illustrates the pattern: fast cache for recent decisions, durable graph for discovery, and explicit schema versioning.
The Real Constraint
The limiting factor isn’t architecture—it’s whether agents actually learn. If Agent B can retrieve what Agent A did but doesn’t update its own behavior based on success or failure, you have a storage system, not a learning system.
This is why researchers like Robert O’Callahan (who recently left Google DeepMind citing unsustainable AI development velocity) worry about agentic automation: the system accumulates memory but not wisdom. It can reference past decisions but can’t reliably distinguish good ones from lucky ones.
Bottom line: Start with a vector store plus an explicit log of decisions (not just embeddings). Add a graph layer only if your agents need to share precise, actionable facts. Version everything and query a canonical source of truth for high-stakes coordination, don’t rely on agents to maintain consistency in memory alone.
Question via Hacker News