#139 · Primary category: LLM Application Frameworks

ExtractThinker

ai document-image-analysis document-intelligence document-parsing document-processing langchain llm machine-learning nlp ocr openai pdf pdf-to-text python

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

Project last updated:08/27/25

GitHub Stars

1.6K

Forks

154

Contributors

7

License

Apache-2.0

Why we included this project

If you've ever had to pull invoice numbers or contract dates out of scanned PDFs, you know the usual dance: OCR, then some fragile parsing, then a prayer that the LLM gets the format right. ExtractThinker wraps that into a more predictable flow. You define a Pydantic schema for the fields you want, attach a loader like Tesseract or Textract, and pick an LLM provider to fill in the values. It also handles document classification and can split large files page by page, with async support that keeps long jobs from blocking your pipeline. The result is a library that feels less like gluing tools together and more like writing a query against a document.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category