#17 · Primary category: Data Annotation & Labeling Tools
refinery
The data scientist's open-source choice to scale, assess and maintain natural language data. Treat training data like a software artifact.
Project last updated:12/09/24
GitHub Stars
1.5K
Forks
73
Contributors
11
License
Apache-2.0
Why we included this project
For data scientists building NLP models, the labeled data is often the real bottleneck, not the algorithm. Refinery gives you a web-based workspace to annotate text, review and correct label quality, and keep a record of how each dataset was assembled. It builds human-in-the-loop into the flow, with active-learning suggestions pointing to the examples that most need labeling and automated checks that flag disagreements between annotators. If you manage multiple annotators or iterate on dataset versions, this turns label management from exported CSVs into a repeatable, auditable process.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
doccano
Open source annotation tool for machine learning practitioners.
X-AnyLabeling
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
snorkel
A system for quickly generating training data with weak supervision
argilla
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
anylabeling
Effortless AI-assisted data labeling with AI support from YOLO, Segment Anything (SAM+SAM2/2.1+SAM3), MobileSAM!!