#17 · Primary category: Data Annotation & Labeling Tools

refinery

active-learning annotations artificial-intelligence data-centric-ai data-labeling data-science deep-learning human-in-the-loop labeling labeling-tool machine-learning natural-language-processing neural-search nlp python spacy supervised-learning text-annotation text-classification transformers

The data scientist's open-source choice to scale, assess and maintain natural language data. Treat training data like a software artifact.

Project last updated:12/09/24

GitHub Stars

1.5K

Forks

73

Contributors

11

License

Apache-2.0

Why we included this project

For data scientists building NLP models, the labeled data is often the real bottleneck, not the algorithm. Refinery gives you a web-based workspace to annotate text, review and correct label quality, and keep a record of how each dataset was assembled. It builds human-in-the-loop into the flow, with active-learning suggestions pointing to the examples that most need labeling and automated checks that flag disagreements between annotators. If you manage multiple annotators or iterate on dataset versions, this turns label management from exported CSVs into a repeatable, auditable process.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category