#11 · Primary category: Data Quality & Cleaning
lazynlp
Library to scrape and clean web pages to create massive datasets.
Project last updated:11/11/20
GitHub Stars
2.3K
Forks
324
Contributors
7
License
Other
Why we included this project
Most people building language models discover that the hard part isn't the model; it's getting a clean, large enough corpus to train on. lazynlp attacks that problem directly: it gives you a simple pipeline that crawls web pages, strips them into plain text, and removes duplicates, so what's left is a genuinely usable monolingual dataset. It includes helpers to fetch URL lists from sources like Reddit dumps and Project Gutenberg, along with utilities for boilerplate removal and content cleanup. That makes it a sensible starting point for researchers and engineers who need to assemble training data without writing their own scraping and cleaning code from scratch.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
qsv
Blazing-fast Data-Wrangling toolkit
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.