#14 · Primary category: AI Data Infrastructure & Storage
unstract
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
Project last updated:08/29/26
GitHub Stars
7.2K
Forks
709
Contributors
30
License
AGPL-3.0
Why we included this project
Unstract tackles the messy middle of document automation: getting usable structured data out of PDFs, scanned files, contracts, and other formats that refuse to come in a clean layout. Teams that see invoices, forms, or reports in constantly shifting arrangements will find the no-code workflow builder most useful, since it lets non-specialists set up extraction without writing parsing logic per document type. Under the hood the project leans on LLMs and OCR to read layouts and emit consistent JSON, and because it ships as an API with ETL pipeline support, it fits into an existing data engineering stack rather than needing its own little world. If your pain point is repeatable extraction from real-world documents, run it against your own files before deciding.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ClickHouse
ClickHouse® is a real-time analytics database management system
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks
gun
An open source cybersecurity protocol for syncing decentralized graph data.
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.