MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 141 |
beir
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets. |
2.3K | 251 | 10/16/25 | Apache-2.0 |
| 142 |
cml
♾️ CML - Continuous Machine Learning | CI/CD for ML |
4.2K | 344 | 06/02/25 | Apache-2.0 |
| 143 |
optimate
A collection of libraries to optimise AI model performances |
8.3K | 616 | 07/22/24 | Apache-2.0 |
| 144 |
xai
XAI - An eXplainability toolbox for machine learning |
1.3K | 187 | 11/29/25 | MIT |
| 145 |
langtrace
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vectorDBs and more.. Integrate using Typescript, Python. 🚀💻📊 |
1.2K | 127 | 11/17/25 | AGPL-3.0 |
| 146 |
Firefly
Firefly: 大模型训练工具,支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、Llama、Qwen、Baichuan、ChatGLM2、InternLM、Ziya2、Vicuna、Bloom等大模型 |
6.7K | 583 | 10/24/24 | Other |
| 147 |
alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast. |
2.0K | 315 | 08/09/25 | Apache-2.0 |
| 148 |
yellowbrick
Visual analysis and diagnostic tools to facilitate machine learning model selection. |
4.4K | 569 | 02/19/25 | Apache-2.0 |
| 149 |
VisualDL
Deep Learning Visualization Toolkit(『飞桨』深度学习可视化工具 ) |
4.9K | 631 | 01/22/25 | Apache-2.0 |
| 150 |
nannyml
nannyml: post-deployment data science in python |
2.2K | 192 | 07/12/25 | Apache-2.0 |
| 151 |
featureform
The Virtual Feature Store. Turn your existing data infrastructure into a feature store. |
2.0K | 107 | 07/03/25 | MPL-2.0 |
| 152 |
DB-GPT-Hub
A repository that contains models, datasets, and fine-tuning techniques for DB-GPT, with the purpose of enhancing model performance in Text-to-SQL |
2.0K | 250 | 07/02/25 | MIT |
| 153 |
determined
Determined is an open-source machine learning platform that simplifies distributed training, hyperparameter tuning, experiment tracking, and resource management. Works with PyTorch and TensorFlow. |
3.2K | 373 | 03/20/25 | Apache-2.0 |
| 154 |
clean-fid
PyTorch - FID calculation with proper image resizing and quantization steps [CVPR 2022] |
1.2K | 81 | 08/02/25 | MIT |
| 155 |
labml
🔎 Monitor deep learning model training and hardware usage from your mobile phone 📱 |
2.3K | 152 | 04/10/25 | MIT |
| 156 |
Deep-Learning-in-Production
In this repository, I will share some useful notes and references about deploying deep learning-based models in production. |
4.4K | 685 | 11/09/24 | Other |
| 157 |
RLHF-Reward-Modeling
Recipes to train reward model for RLHF. |
1.5K | 110 | 04/24/25 | Apache-2.0 |
| 158 |
whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈 |
2.8K | 144 | 01/10/25 | Apache-2.0 |
| 159 |
llm-colosseum
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM |
1.5K | 182 | 03/21/25 | MIT |
| 160 |
hopsworks
Hopsworks - Data-Intensive AI platform with a Feature Store |
1.3K | 160 | 02/10/25 | AGPL-3.0 |