#55 · Primary category: NLP Tools & Text Processing
docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
Project last updated:03/17/26
GitHub Stars
2.1K
Forks
155
Contributors
10
License
Apache-2.0
Why we included this project
docext is a sensible pick for teams that need structured fields out of invoices, passports, receipts, and long scanned PDFs without standing up a heavy OCR stack. It is an on-premises toolkit that uses vision-language models to convert documents and images to markdown, picking out tables, LaTeX equations, signatures, watermarks, and form checkboxes along the way. The same extraction pipeline returns fields with confidence scores, so a database or retrieval system gets clean, queryable output instead of raw OCR text. A bundled benchmarking harness also lets you compare how different models handle key information extraction, table parsing, and document classification on your own data before you commit to one.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
flair
A very simple framework for state-of-the-art Natural Language Processing (NLP)
compromise
modest natural-language processing
tokenizers
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
CoreNLP
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
Chinese-Word-Vectors
100+ Chinese Word Vectors 上百种预训练中文词向量