#39 · Primary category: Knowledge Base & RAG
PixelRAG
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
Project last updated:08/29/26
GitHub Stars
9.8K
Forks
836
Contributors
20
License
Apache-2.0
Why we included this project
Standard retrieval pipelines often lose the meaning locked in a page's visual layout. PixelRAG takes a different route: it renders documents into screenshot tiles and indexes those images, so searches can match on charts, typography, and structure that plain text extraction would throw away. The tool itself stays small, with two main commands, one to render a page into tiles and another to query a visual index, plus a hosted index of 8.28 million Wikipedia pages that lets you try the approach without building your own pipeline. It grew out of Berkeley research that compared screenshots against text for retrieval and found the visual approach holds up in practice. For a small team wondering whether visual retrieval.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
docling
Get your documents ready for gen AI
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience