Office Hours — What are the best practices for formatting and structuring comments in LLM prompts? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-10-06T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — What are the best practices for formatting and structuring comments in LLM prompts?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

What are the best practices for formatting and structuring comments in LLM prompts?

The way you comment your prompts matters more than most developers realize. Comments aren’t just documentation—they’re instructions that influence how the model interprets context, prioritizes tasks, and makes decisions. The difference between a prompt that consistently works and one that’s flaky often comes down to comment structure, not the core instruction.

Comments Are Instructions, Not Just Notes

An LLM doesn’t treat comments as passive metadata. When you write a comment, you’re adding semantic weight that the model weighs during inference. A well-placed comment can disambiguate a task that would otherwise be ambiguous. A poorly placed one can introduce noise that dilutes the primary directive.

The key insight: use comments to explain why a constraint exists, not just what the constraint is. “Do not use external libraries” is weaker than “Do not use external libraries—the environment has no pip access and imports will fail silently.” The second version gives the model actual reasoning to anchor on.

Structure Comments by Responsibility Layer

Organize comments vertically by their scope and responsibility. Top-level comments explain the task’s purpose and success criteria. Mid-level comments clarify specific steps or conditional logic. Inline comments flag edge cases and exceptions.

Here’s a working structure:

# TASK: Extract structured billing records from unformatted PDF text
# SUCCESS: Return JSON array with at least 95% accuracy on amount and date fields
# FAILURE MODE: Missing or hallucinated invoice numbers are common—validate against sequence

## STEP 1: Parse raw invoice text
# Look for patterns like "Invoice #12345" or "INV-2026-001"
# Some PDFs use "Bill #" instead—treat as equivalent

## STEP 2: Extract amount
# Accept formats: $1,234.56 or 1234.56 or 1,234.56 EUR
# If amount appears in multiple places, use the largest value (usually subtotal or total)
# Never invent amounts if parsing fails

## STEP 3: Return JSON
# Include source_confidence: "high" if all three fields parsed cleanly, "medium" if one required reconstruction

The vertical structure lets the model understand task hierarchy. Top comments set stakes. Middle comments explain logic. Inline comments catch edge cases the model might otherwise miss.

Use Comments to Encode Domain Knowledge Without Over-Explaining

Domain-specific context belongs in comments, but keep it tight. Don’t explain what the model already knows, but do flag domain-specific gotchas.

Bad: “In Python, dictionaries are unordered collections…” Good: “Maintain dict insertion order—output should preserve the sequence from the input PDF, not alphabetical.”

Bad: “CSV files have commas…” Good: “Some invoice PDFs export with semicolon delimiters when the locale is EU—handle both comma and semicolon.”

The difference is actionable specificity. The second version tells the model something it wouldn’t guess from training data alone.

Separate Guardrails from Task Logic

Guardrails (safety, security, compliance constraints) should be visually distinct from task logic. Use a separate section.

# TASK LOGIC
# Extract customer names from the email subject line
# Use the format: "Last, First" if available

# GUARDRAILS
# Do not extract or process email addresses (PII)
# Do not hallucinate customer names if the subject line is ambiguous
# If you cannot extract a name with confidence, return null instead of guessing

This prevents guardrails from being accidentally overwritten by downstream task instructions. Models are sensitive to order and visual hierarchy. Guardrails at the top signal “these are non-negotiable.”

Comment Density Matters

Too many comments cause dilution. The model has to parse signal from noise. Aim for one comment per 2-3 lines of instruction, not one per line.

Bad (over-commented):

# Create a list
result = []
# Loop through items
for item in items:
    # Check if valid
    if is_valid(item):
        # Add to result
        result.append(item)

Good (signal-to-noise ratio):

# Filter to valid items only—invalid items must be skipped entirely, not logged
for item in items:
    if is_valid(item):
        result.append(item)

The second version gives the model the constraint (invalid items skipped, not logged) without drowning it in play-by-play narration.

Use Comments to Flag Model Limitations

If you know the model struggles with a specific pattern, say so explicitly.

“Note: Models often confuse invoice date with payment due date. Look for keywords like ‘Due:’ or ‘Due Date:’ to distinguish. If both appear, due date comes after invoice date in the document.”

This transforms a failure mode into a solvable problem. The model can now attend to the specific signal that distinguishes the two.

Conditional Comments for Branches

When your prompt has multiple pathways, use comments to label branches clearly.

# IF the input is a scanned PDF:
#   - Run OCR preprocessing
#   - Expect 20-30% more noise in extracted text
#   - Lower confidence thresholds accordingly

# IF the input is already digitized:
#   - Parse directly
#   - Expect near-perfect accuracy on structured fields
#   - Reject any output with missing amounts

This prevents the model from accidentally blending strategies or applying the wrong confidence threshold to the wrong input type.

Test Comment Changes Independently

Comments influence output quality, and the effect is often invisible in logs. If your structured outputs start diverging, don’t just blame the model version. Test your comments in isolation by running the same prompt with and without a specific comment section.

I’ve seen teams spend weeks optimizing the task instruction when the real problem was a single missing comment flagging an edge case.

Bottom line:

Structure comments to encode task hierarchy (top-level purpose, mid-level steps, inline edge cases), separate guardrails visually from task logic, and use comments to flag domain-specific gotchas and model failure modes rather than restating what the model should already know. Comment strategically, not verbosely—each comment should add actionable information or resolve an ambiguity.

Question via Hacker News