#55 · Primary category: NLP Tools & Text Processing

docext

document document-analysis document-data-extraction document-information-extraction extraction llm-ocr llms machine-learning nlp ocr ocr-benchmark ocr-onpremise onprem onprem-ocr onprem-vision onpremise rag table-extraction unstructured-data vlms

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

Project last updated:03/17/26

GitHub Stars

2.1K

Forks

155

Contributors

10

License

Apache-2.0

Why we included this project

docext is a sensible pick for teams that need structured fields out of invoices, passports, receipts, and long scanned PDFs without standing up a heavy OCR stack. It is an on-premises toolkit that uses vision-language models to convert documents and images to markdown, picking out tables, LaTeX equations, signatures, watermarks, and form checkboxes along the way. The same extraction pipeline returns fields with confidence scores, so a database or retrieval system gets clean, queryable output instead of raw OCR text. A bundled benchmarking harness also lets you compare how different models handle key information extraction, table parsing, and document classification on your own data before you commit to one.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category