#11 · Primary category: Data Quality & Cleaning
Units_of_Measure_Harmonization-intelligence-platform
Production-Grade ML System for Automated Unit of Measure Error Detection | 88-92% Accuracy | 94% Autonomy | KNIME Workflow
Project last updated:05/21/26
GitHub Stars
817
Forks
753
Contributors
1
License
Apache-2.0
Why we included this project
Unit of measure errors are a quiet source of real losses in manufacturing and procurement, and this KNIME workflow goes after them directly. It reads CSV or Excel files, runs an XGBoost classifier over 60+ engineered features to flag suspicious values, then checks each one against NIST-compliant physics-based conversion rules before correcting it. A Q-learning agent handles the routine fixes on its own, while the built-in dashboard shows confidence scores, root-cause analytics, and anything that still needs a human reviewer. That pairing of ML detection with rule-based verification is what makes the results defensible, which matters when you need auditable data cleaning rather than a black-box model. It runs in KNIME Analytics Platform 4.5+, so teams already comfortable with visual workflow tooling can pick it up quickly, and the approach carries over to domains beyond manufacturing.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
qsv
Blazing-fast Data-Wrangling toolkit
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.