#26 · Primary category: Business Intelligence & Analytics
fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Project last updated:04/22/26
GitHub Stars
13.7K
Forks
1.8K
Contributors
142
License
MIT
Why we included this project
For anyone doing exploratory analysis in Jupyter, fg-data-profiling (the renamed ydata-profiling) turns a raw DataFrame into a readable report with a single call. It produces per-column statistics, type inference, missing-value and duplicate checks, correlation matrices, and alerts for common quality problems such as skewness or constant values. The report exports as HTML for stakeholders or JSON for downstream tooling, so the same output serves people and machines. Because it handles Pandas and Spark DataFrames as well as time-series and text, it is a fast way to sanity-check incoming data before modeling starts.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
spark
Apache Spark - A unified analytics engine for large-scale data processing
metabase
The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:
streamlit
Streamlit — A faster way to build and share data apps.
duckdb
DuckDB is an analytical in-process SQL database management system
ToolJet
ToolJet is the open-source foundation of ToolJet AI - the enterprise app generation platform for building internal tools, dashboard, business applications, workflows and AI agents 🚀