#119 · Primary category: Knowledge Base & RAG
crw
Fast, lightweight Rust web scraper, crawler & search API with MCP server, drop-in Firecrawl-compatible, 2.3x faster than Tavily.
Project last updated:08/27/26
GitHub Stars
876
Forks
57
Contributors
8
License
AGPL-3.0
Why we included this project
If your RAG pipeline or agent needs current web content, the setup cost of a scraper often ends up bigger than the actual work. crw is a small Rust binary that handles search, scraping, crawling, and structured extraction in one place, returning clean markdown or typed JSON you can feed straight into an ingestion job. It ships a Firecrawl-compatible REST API, so teams already wired to that service can swap it in without rewriting code, plus an MCP server that Claude Code, Cursor, and similar tools pick up natively. A one-command install gets it running locally with no account, and you can also self-host it behind your own network or call it through the typed Python and Node SDKs. For a lightweight tool it is genuinely fast, about 2.3x faster than Tavily and 1.5x faster than Firecrawl in 1K-URL benchmarks while sitting at roughly 6 MB of RAM.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
docling
Get your documents ready for gen AI
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience