#4 · Primary category: Knowledge Base & RAG
docling
Get your documents ready for gen AI
Project last updated:08/28/26
GitHub Stars
65.7K
Forks
4.7K
Contributors
289
License
MIT
Why we included this project
Docling is a document parsing library that handles the files people actually have: PDFs, DOCX, PPTX, XLSX, HTML, EPUB, and even audio and video. It converts them into a single structured document model that exports to Markdown, HTML, or lossless JSON. Reading order, table structure, layout, and formulas survive the conversion, so you feed a vector store or an agent real structure instead of a flat text dump. It runs locally for sensitive or air-gapped data, includes OCR for scanned pages, and has ready-made integrations with LangChain, LlamaIndex, and Haystack, which keeps the glue code minimal. For teams building knowledge bases, document Q&A, or agent workflows over real-world files, that structured output is usually the difference between a working pipeline and one that needs constant cleanup.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
meilisearch
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.