#228 · Primary category: Knowledge Base & RAG
WEB_KG
Crawl Baidu Baike Chinese pages, extract triples, and build a Chinese knowledge graph.
Project last updated:07/20/20
GitHub Stars
958
Forks
195
Contributors
3
License
Other
Why we included this project
WEB_KG is a complete pipeline for turning Chinese encyclopedia content into a queryable knowledge graph rather than a set of one-off scripts. A Scrapy spider pulls Baidu Baike pages, the parser extracts subject-predicate-object triples, and the raw pages and triples are stored in MongoDB before being pushed into Neo4j, where you can browse the finished graph through a web interface. The code is compact enough to read end to end, and the README covers deployment on both Linux and Windows. Because the extraction logic is written against Baidu Baike's page structure, it works best as a reference or starting point for Chinese-language scraping and graph pipelines; expect to adjust the parsing if you point it at another site.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
docling
Get your documents ready for gen AI
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience