#13 · Primary category: Data Annotation & Labeling Tools
autolabel
Label, clean and enrich text datasets with LLMs.
Project last updated:03/05/25
GitHub Stars
2.3K
Forks
159
Contributors
19
License
MIT
Why we included this project
If your team is still hand-labeling text samples to train a supervised model, this Python library hands that job to an LLM of your choosing. You point it at your own dataset and pick the model, whether GPT-4, Claude, or a local one; it handles the labeling, cleans noisy rows, and enriches the text through your own pipeline. Confidence scores and error analysis come built in, so you can see which rows the model likely got wrong instead of accepting every label at face value. The bundled benchmark tooling is useful if you run annotation jobs repeatedly or maintain training sets that keep changing, since it lets you compare how different models perform before you commit to one. For data engineers and ML practitioners who want programmatic control over a labeling pipeline rather than a hosted annotation UI, that is the appeal.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
doccano
Open source annotation tool for machine learning practitioners.
X-AnyLabeling
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
snorkel
A system for quickly generating training data with weak supervision
argilla
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
anylabeling
Effortless AI-assisted data labeling with AI support from YOLO, Segment Anything (SAM+SAM2/2.1+SAM3), MobileSAM!!