Library of the Week — Weights & Biases Weave A weekly teardown of one open-source AI/ML library: what it does, why it stands out, and when to use it. 2026-10-02T12:00:00.000Z Library of the Week Library of the Week open-sourcelibrariestoolsdeveloper-tools

Library of the Week — Weights & Biases Weave

A weekly teardown of one open-source AI/ML library: what it does, why it stands out, and when to use it.

Weekly One open-source library you should know about.

Weights & Biases Weave — already covered. Let me pick a fresh one.


Docling — a fast, accurate document parser built for LLM ingestion pipelines

GitHub · Language: Python · License: MIT

What it does

Docling converts PDFs, Word docs, PowerPoints, images, and HTML into clean, structured Markdown or JSON — preserving tables, figures, reading order, and document hierarchy. It’s built specifically for the “get real documents into your RAG pipeline” problem, where generic parsers produce garbage that degrades retrieval quality. The target audience is developers building document-heavy LLM applications who need something beyond pdfminer noise.

Why it stands out

  • Layout-aware parsing at inference time: Docling runs its own document layout analysis and table structure recognition models locally — no cloud API call required. This matters for sensitive enterprise documents where sending to a third-party parser isn’t acceptable.
  • First-class table handling: Tables are exported as proper Markdown or structured JSON, not flattened into a stream of whitespace-separated tokens. Most PDF parsers completely mangle tables; Docling treats them as a primary output type.
  • Chunking built in: The HierarchicalChunker respects document structure (headings, sections, captions) when splitting, so your chunks don’t arbitrarily cut across logical boundaries — a real pain point when using naive character splitters with LlamaIndex or similar.
  • Integrations ship in the box: Official connectors for LlamaIndex and LangChain mean you can drop it into an existing pipeline without glue code.

Quick start

from docling.document_converter import DocumentConverter

converter = DocumentConverter()
result = converter.convert("quarterly_report.pdf")

# Get clean Markdown for LLM context
markdown = result.document.export_to_markdown()
print(markdown[:500])

# Or export structured JSON
doc_dict = result.document.export_to_dict()

When to use it

  • You’re building a RAG system over enterprise documents (PDFs, PPTX, DOCX) and standard parsers are destroying table structure or reading order.
  • You need local, offline document processing — Docling’s layout models run on CPU without phoning home.
  • You want chunking that respects section boundaries rather than splitting on character count.

When to skip it

  • If your documents are clean, text-only PDFs with no tables or complex layouts, simpler tools like pypdf or pdfplumber are faster with less overhead.
  • First-pass conversion of large document batches can be slow on CPU; if throughput is critical at scale, benchmark against managed cloud parsers like AWS Textract before committing.

The verdict

Docling is the most complete open-source answer to the “my PDFs look terrible in RAG” complaint. Its layout models plus native chunking make it a meaningful upgrade over the patchwork of unstructured + custom splitters that most teams end up duct-taping together. If you’re ingesting anything more complex than plain text into a retrieval pipeline, start here.