#14 · Primary category: AI Data Infrastructure & Storage

unstract

ai-agents data-engineering document-ai generative-ai idp json-extraction llm mcp-server ocr pdf-extraction prompt-engineering structured-output

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

Project last updated:08/29/26

GitHub Stars

7.2K

Forks

709

Contributors

30

License

AGPL-3.0

Why we included this project

Unstract tackles the messy middle of document automation: getting usable structured data out of PDFs, scanned files, contracts, and other formats that refuse to come in a clean layout. Teams that see invoices, forms, or reports in constantly shifting arrangements will find the no-code workflow builder most useful, since it lets non-specialists set up extraction without writing parsing logic per document type. Under the hood the project leans on LLMs and OCR to read layouts and emit consistent JSON, and because it ships as an API with ETL pipeline support, it fits into an existing data engineering stack rather than needing its own little world. If your pain point is repeatable extraction from real-world documents, run it against your own files before deciding.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category