What is Docling?
Docling is an open-source toolkit that converts PDFs, Office files, HTML, images, and audio into structured data for AI apps. It detects layout, tables, formulas, and reading order, then exports Markdown, HTML, or JSON. Started at IBM Research and hosted by the LF AI and Data Foundation, it runs fully offline.
Top Features:
- Layout understanding: recovers tables, formulas, figures, and reading order from multi-column pages.
- Page provenance: every element keeps its page number and bounding box for citations.
- Flexible setup: use the Python library, CLI, REST service, or MCP server.
Use Cases:
- RAG pipelines: chunk documents with structure intact so answers trace to exact pages.
- Data extraction: pull defined fields from invoices or reports into typed database rows.
- Bulk conversion: convert whole folders of PDFs into Markdown or JSON from the terminal.
Who Can Use Docling?
- AI engineers: feed structured document content into retrieval, agent, and extraction pipelines.
- Data teams: turn archives of scanned reports and spreadsheets into searchable, reusable data.
- Regulated enterprises: process sensitive files on private infrastructure without sending them outside.
Pricing
- Open source (free): the full MIT licensed library, CLI, server, and MCP tools.
- Managed trial (free, 30 days): hosted on IBM watsonx with 5,000 pages and every feature included.
- Managed ($4 per 1,000 pages): pay as you go on IBM watsonx, with annual subscriptions available.
Pros and Cons
Pros:
- Free and local: no account, upload, or proprietary SDK is needed to get started.
- Structure survives: tables, headings, and reading order stay intact instead of flat text.
- Wide ecosystem: plugs into LangChain, LlamaIndex, Haystack, Crew AI, and similar frameworks.
Cons:
- Developer focused: you need Python 3.10 or newer and some coding comfort.
- Hardware load: large batches can run slowly on laptops without a capable GPU.
- Managed tier: hosted scale and private deployments come through a paid IBM service.
FAQs:
1) Is it really free?
Yes, the library is MIT licensed, while IBM sells an optional managed version.
2) Which formats does it read?
PDF, Office files, HTML, images, and audio, exported as Markdown or JSON.
3) Does it need internet access?
No, models download once and then every conversion runs offline.
4) Can AI agents use it?
Yes, an MCP server gives agents typed tools to read and convert documents.
5) Does it handle scans?
Yes, built-in OCR reads scanned PDFs and images alongside digital files.