#11 · Primary category: LLM Application Frameworks

langextract

gemini gemini-ai gemini-api gemini-flash gemini-pro information-extration large-language-models llm nlp python structured-data

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

Project last updated:08/27/26

GitHub Stars

38.5K

Forks

2.7K

Contributors

24

License

Apache-2.0

Why we included this project

LangExtract solves a specific problem: turning free-form documents into structured data you can query, with every extracted value traced back to the exact character span it came from in the source text. You define the task with a short prompt and a few examples, and the library runs an LLM over the document, chunking long texts and processing them in parallel. That grounding is what makes it useful in practice: instead of trusting the model's output, you can open an interactive HTML report and see where each of thousands of items was pulled from. It holds up on real jobs like structuring clinical notes or radiology reports, and it works with cloud models such as Gemini and OpenAI as well as local ones through Ollama. Since no fine-tuning is required, a small team can get dependable structured output without building a custom pipeline.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category