#75 · Primary category: LLM Application Frameworks

docetl

agents data data-pipelines document-analysis document-processing elt etl llm python semantic-data unstructured-data unstructured-data-analysis workflow

A system for agentic LLM-powered data processing and ETL

Project last updated:08/10/26

GitHub Stars

4.1K

Forks

438

Contributors

29

License

MIT

Why we included this project

DocETL is for teams that need to run LLM processing over messy real-world data without hand-wiring thousands of model calls. You describe each operation in plain language, like pulling every complaint out of a ticket, and the tool handles the operators, orchestration, and parallelization across your collection. The optimizer is what sets it apart: it swaps models, rewrites prompts, breaks down complex operations, and replaces subtasks with plain code when possible, all to improve accuracy and cut costs. Results come back as queryable tables that slot into your existing database or analytics stack. Data engineers get both a Python API and a low-code YAML interface, plus a visual UI for building and inspecting pipelines.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category