Paper of the Week
SeriesSeptember 2026
- Paper of the Week — Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models
4-bit quantization choice leaks fine-tuned PII at different rates — GPTQ-style methods outperform GGUF on privacy, independent of bit width.
- Paper of the Week — Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
KV-cache repair after edits: a new paper shows patching only contiguous changed blocks beats importance-scoring by 2–3× on latency with near-zero quality loss.
- Paper of the Week — Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
Token-level SFT masking on math reasoning: a small-team paper showing which tokens actually hurt training, with released analysis tools.
- Paper of the Week — HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
KV cache memory hogs in hybrid LLMs traced to per-head residency imbalance — a new budget allocator cuts waste without accuracy loss.
August 2026
- Paper of the Week — Function-Level Execution Feedback for Code Preference Optimization
Coding agents using function-level execution signals for process supervision—patch-level test feedback outperforms line and trace approaches by 8+ points on HumanEval variants.
- Paper of the Week — Looped Language Models Improve Compositional Tool Calling
Looped LLMs cut tool-calling errors by maintaining hidden state across API calls — a quiet fix for multi-step agent reliability that costs zero extra parameters.
- Paper of the Week — EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
Enterprise RAG hits an 80%→27% accuracy cliff when all constraints must hold simultaneously — here's the benchmark that quantifies it.
July 2026
- Paper of the Week — Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
Self-repair in code agents adds feedback overhead but rarely beats simply sampling again — a controlled study quantifies exactly when the extra tokens are wasted.
- Paper of the Week — Structured Output Collapses Answer Diversity Across 44 Language Models
Structured output silently collapses answer diversity across 44 LLMs — JSON-only prompts skew which answer a model picks, not just how it formats the response.
- Paper of the Week — From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
AI agent tool sets that ossify at deployment lose compounding value — this paper shows how to grow them automatically from task execution logs.
- Paper of the Week — SWE-Router: Routing in Multi-Turn Agentic Software Engineering Tasks
SWE-Router cuts LLM agent costs by routing easy GitHub issues to cheaper models mid-trajectory, not just at task start, using execution signals as routing features.
June 2026
- Paper of the Week — Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models
Quantization inflates reasoning model token counts by up to 4×, a hidden cost accuracy benchmarks miss entirely.
- Paper of the Week — Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery
Routing accuracy for 110+ agents drops sharply past 50 tools — new study maps the degradation curve and identifies three recovery strategies that work today.
- Paper of the Week — When Poison Fails After Retrieval: Revisiting Corpus Poisoning under Chunking and Reranking Pipelines
RAG corpus poisoning drops 60-80% when you add chunking+reranking — but most attack evals skip these standard pipeline stages entirely.
- Paper of the Week — Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents
Activation probes detect credential exfiltration *before* the LLM outputs any tokens — combined with honeytokens and multi-turn leakage tracking, with no model changes required.
May 2026
- Paper of the Week — Exploring the Emerging Threats of the Agent Skill Ecosystem
76 malicious skills confirmed in 3,984 audited AI agent marketplaces — credential theft, backdoor installation, and data exfiltration found hiding in plain sight.
- Paper of the Week — Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study
Code cleanliness measurably changes how well coding agents complete tasks — a controlled minimal-pair study with released dataset and reproducible methodology.
- Paper of the Week — Useful Memories Become Faulty When Continuously Updated by LLMs
LLM agent memory degrades when consolidated: episodic traces outperform summarized lessons across 5 agentic tasks.
- Paper of the Week — TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments
TSCG shows small LLMs (4–14B) drop tool-call failures by compiling JSON schemas into natural-language descriptions before inference.
April 2026
- Paper of the Week — Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
Rewriting tool descriptions at deployment time—not training time—can recover 20-40% of function-calling accuracy lost to poorly written API docs.
- Paper of the Week — Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations
Visualizing LLM output distributions reveals hidden modes, edge cases, and prompt sensitivity that single-sample evaluation completely misses.
- Paper of the Week — Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
Lossless prompt compression via dictionary encoding lets LLMs analyze repeated data at a fraction of token cost — no external tools, just in-context learning.
- Paper of the Week — Mamba-Based State Space Models for Long-Context Retrieval-Augmented Generation
Structured state-space models finally beat transformers at document retrieval — here's what the Mamba-based RAG benchmark actually shows.
- Paper of the Week — SnapKV: LLM Knows What You are Looking for Before Generation
KV cache compression that cuts memory 40–60% with under 1% accuracy loss — here's the technique your inference stack probably isn't using yet.