#139 · Primary category: LLM Application Frameworks
ExtractThinker
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Project last updated:08/27/25
GitHub Stars
1.6K
Forks
154
Contributors
7
License
Apache-2.0
Why we included this project
If you've ever had to pull invoice numbers or contract dates out of scanned PDFs, you know the usual dance: OCR, then some fragile parsing, then a prayer that the LLM gets the format right. ExtractThinker wraps that into a more predictable flow. You define a Pydantic schema for the fields you want, attach a loader like Tesseract or Textract, and pick an LLM provider to fill in the values. It also handles document classification and can split large files page by page, with async support that keeps long jobs from blocking your pipeline. The result is a library that feels less like gluing tools together and more like writing a query against a document.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
langchain
The agent engineering platform.
dify
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
litellm
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
llama_index
LlamaIndex is the leading document agent and OCR platform