Office Hours — What techniques work best for making legacy codebases legible and navigable to AI coding agents? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-09-25T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — What techniques work best for making legacy codebases legible and navigable to AI coding agents?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

What techniques work best for making legacy codebases legible and navigable to AI coding agents?

Legacy codebases are the worst-case scenario for AI agents. They’re often undocumented, architecturally baroque, built across multiple languages and frameworks, and sprawling enough that even a 1M context window feels small. The real problem isn’t that agents can’t read code—it’s that they read everything equally, wasting tokens on noise and losing sight of what matters.

The Codebase Walkthrough Problem

An agent hitting a 500K-line monolith cold will spend 60% of its context budget just establishing what the hell it’s looking at. It’ll read helper functions that are irrelevant, trace imports three levels deep, and then still miss the architectural intent because that knowledge lives in someone’s head, not in a docstring.

The best fix is explicit, lightweight triage before the agent even starts. This isn’t about rewriting your codebase—it’s about making your codebase transparent to machines in the way it already is to humans who’ve spent six months there.

Concrete Technique: Semantic File Mapping

Build a single “navigation document” at the repo root that a frontier model generates once and then updates sparingly. This should be:

  • A flat list of 30-50 key files and their purposes (one sentence each)
  • Dependency flow (which modules call which)
  • Any architectural constraints or patterns (monolith vs. microservices, API conventions, data flow)
  • Known hazard zones (files where logic is particularly tangled, deprecated patterns still in use, places where agents have broken things before)

Example format:

## Core Architecture
- `auth/jwt_handler.py`: Validates tokens and extracts user context. All auth flows depend on this.
- `db/models.py`: ORM definitions. Do not add new migrations without checking `db/migrations/README.md` first.
- `api/router.py`: Entry point. Routes to 12 service modules below.

## Service Modules
- `services/user_service.py`: User CRUD. Uses cached queries; mutations require cache invalidation in `cache.py`.
- `services/billing_service.py`: **HAZARD**: Contains legacy payment logic alongside newer Stripe integration. Only call `stripe_*` functions; ignore `old_payment_*`.

## Data Layer
- `db/queries/`: SQL-heavy; each query file is independently cached. Modifying a query invalidates the cache key—check `cache_keys.py`.

## Known Issues
- Session serialization in `auth/session.py` has a race condition under high concurrency. Don't modify without load testing.
- API versioning is implicit; check `api/VERSIONS.txt` before adding new endpoints.

This is machine-digestible enough that an agent can read it in 200 tokens instead of 50K, and specific enough that it actually matters. It’s not replacing the codebase—it’s a fast index to the codebase.

Second Technique: Type Signatures and Module Docstrings as Breadcrumbs

AI agents are exceptionally good at following type hints and function signatures. If your code has zero type information, add it strategically to the entry points and boundaries:

  • Every public function in a service module should have a docstring explaining what it does and what it doesn’t do
  • Type hints on parameters and returns (this is machine-readable without parsing AST)
  • Any function that an agent might call should have an example in the docstring

You don’t need to type-hint your entire codebase. Just the 50 functions that agents are likely to touch. A frontier model will read those signatures and avoid 10 wrong turns.

def update_user_profile(user_id: str, updates: Dict[str, Any]) -> User:
    """
    Updates user profile fields. Does NOT handle password changes or auth state.
    For password resets, use reset_password(). For role changes, use grant_role().
    
    Args:
        user_id: Internal user ID (UUID)
        updates: Dict of fields to update. Valid keys: name, email, avatar_url, bio.
                 Invalid keys are silently ignored.
    
    Returns:
        Updated User object.
    
    Raises:
        UserNotFound: If user_id doesn't exist.
        ValidationError: If email format is invalid.
    """

The docstring is the agent’s manual. If the manual says “do NOT”, the agent will listen—if it has frontier reasoning capability. Cheaper models may ignore it, so test.

Third Technique: Integration Tests as Agent Guides

Write a small integration test suite that exercises the happy path through your codebase in the order an agent is likely to hit it. A good agent will read passing tests as documentation of how things work.

Example: If an agent needs to add a feature that involves user creation, billing, and email notification, create one test that does all three in order and passes:

def test_new_user_onboarding_flow():
    """Complete flow: create user -> assign tier -> send welcome email."""
    user = create_user(email="test@example.com", name="Alice")
    assert user.id is not None
    
    billing = assign_tier(user.id, tier="starter")
    assert billing.status == "active"
    
    email_sent = send_onboarding_email(user.id)
    assert email_sent.success

Agents will read this and understand the interaction pattern without needing you to explain it in natural language. The test is the explanation.

Fourth Technique: Gradual Visibility for Large Services

If you have a service with 10K lines of code, you can’t hand all of it to an agent at once. Instead, structure it with a public interface and hidden implementation:

  • Create a thin __init__.py that exports only the 3-5 functions an agent should call
  • Put the implementation in private modules (subfolders, underscore prefixes)
  • Document the public API, leave the internals alone

When an agent asks “what does this service do?”, it reads the public API and stops. It won’t wander into the implementation unless it has to fix a bug, at which point it’ll have more context about why.

Fifth Technique: Codebase Archaeology

Run a model over your git history to identify which files changed together frequently. Files that are mutated together are likely coupled—if an agent changes one without the other, it’ll break tests. Create a lightweight dependency map:

## Common Mutation Patterns
- Changing `payment.py` usually requires updating `billing_service.py` (6 out of last 10 commits)
- Adding a new API endpoint in `router.py` requires adding integration tests in `tests/api/test_routes.py`
- Modifying ORM models in `models.py` requires a migration in `db/migrations/` AND a cache key update

This is something you can generate semi-automatically and then curate. Agents will check this before refactoring.

Monitoring and Iteration

Track which files agents touch most and which files they break most. If agents keep breaking billing_service.py, add a tighter docstring or mark it as read-only. If they keep ignoring cache_keys.py, move it higher in the navigation document.

The best navigation document isn’t written once—it evolves as agents interact with your codebase and you learn what guidance actually sticks.

Bottom line: Don’t try to make your legacy codebase perfect for agents. Instead, create a lightweight navigation layer (semantic file map, entry-point docstrings, happy-path tests, coupling metadata) that lets agents understand the architecture without reading the entire thing. This cuts agent context waste by 50-70% and dramatically improves reliability on real codebases.

Question via Hacker News