#7 · Primary category: Data Quality & Cleaning
DataProfiler
What's in your data? Extract schema, statistics and entities from datasets
Project last updated:08/26/26
GitHub Stars
1.6K
Forks
187
Contributors
42
License
Apache-2.0
Why we included this project
DataProfiler answers a question every data team hits early: before you can trust a dataset, you need to know its shape and what sensitive fields it carries. Point it at a file and it auto-detects the format, loads CSV, Parquet, JSON, or AVRO into a pandas DataFrame, and returns a profile covering the schema, per-column statistics, and recognized entities. The real differentiator is the pre-trained deep learning labeler, which flags names, email addresses, phone numbers, and other PII or NPI in structured and unstructured text alike. You can also retrain the recognizer or drop in a custom entity pipeline, so you are not stuck with the default labels. For data engineering, governance, and pipeline monitoring, that combination of quick profiling and sensitive-data detection saves a lot of guesswork.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
qsv
Blazing-fast Data-Wrangling toolkit
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.