#174 · Primary category: MLOps & Evaluation

feathr

apache-spark artificial-intelligence azure data-engineering data-quality data-science feature-engineering feature-governance feature-management feature-marketplace feature-metadata feature-platform feature-store machine-learning mlops

Feathr – A scalable, unified data and AI engineering platform for enterprise

Project last updated:04/04/24

GitHub Stars

1.9K

Forks

246

Contributors

45

License

Apache-2.0

Why we included this project

Hand-building training datasets from raw logs, transactions, or event streams is a recurring chore for many ML teams, and Feathr is built to take that pain away. It's a feature store where you define transformations once in Python, register them by name, and reuse the same features for both offline training and online serving, with point-in-time-correct joins that stop data leakage. Because it runs on Spark and plugs into Databricks and Azure Synapse, it handles the billion-row workloads that homegrown pipelines often choke on. Data scientists get a Pythonic API with UDF support, while platform teams get a registry, governance, and materialization for production. If your org is trying to standardize how features are built, shared, and versioned across models, this is a solid reference to study.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category