#53 · Primary category: Knowledge Base & RAG
chonkie
🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines
Project last updated:08/26/26
GitHub Stars
4.7K
Forks
352
Contributors
49
License
MIT
Why we included this project
If you have built a retrieval pipeline by hand, you know the pain of writing text splitters that behave oddly at paragraph boundaries. Chonkie packages the common chunking strategies, from simple token and sentence splits up to semantic and LLM-driven ones, into one lightweight Python library, so you can swap approaches without maintaining splitter code yourself. It can also run the whole ingestion path, fetching a document, cleaning it, chunking, refining, embedding, and handing the result to a vector store, and an optional self-hosted REST API exposes the same logic to non-Python services. The default install stays small and integrations cover dozens of tokenizers, embedding providers, and vector databases, which keeps splitting consistent and reproducible across different retrieval setups.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
docling
Get your documents ready for gen AI
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience