MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 121 |
skillsbench
SkillsBench evaluates how well skills work and how effective agents are at using them. |
1.7K | 363 | 07/23/26 | Apache-2.0 |
| 122 |
AngelSlim
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency. |
1.6K | 174 | 08/07/26 | Other |
| 123 |
waza
CLI / Framework for Agent Skills - create, test, measure and improve skill quality and effectiveness |
1.3K | 79 | 08/29/26 | MIT |
| 124 |
tensorwatch
Debugging, monitoring and visualization for Python Machine Learning and Data Science |
3.5K | 361 | 03/30/26 | MIT |
| 125 |
gemma-tuner-multimodal
Fine-tune Gemma 4 and 3n with audio, images and text on Apple Silicon, using PyTorch and Metal Performance Shaders. |
1.5K | 103 | 08/13/26 | MIT |
| 126 |
Skills
A project to improve skills of large language models |
1.0K | 198 | 08/28/26 | Apache-2.0 |
| 127 |
Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0. |
1.2K | 104 | 08/29/26 | MIT |
| 128 |
GraphGen
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation |
1.2K | 97 | 08/17/26 | Apache-2.0 |
| 129 |
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring. |
1.1K | 96 | 08/24/26 | Apache-2.0 |
| 130 |
AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) |
3.7K | 278 | 02/08/26 | Apache-2.0 |
| 131 |
eli5
A library for debugging/inspecting machine learning classifiers and explaining their predictions |
2.8K | 325 | 04/08/26 | MIT |
| 132 |
deepchecks
Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and models from research to production. |
4.0K | 304 | 12/28/25 | Other |
| 133 |
sacred
Sacred is a tool to help you configure, organize, log and reproduce experiments developed at IDSIA. |
4.4K | 393 | 10/22/25 | MIT |
| 134 |
plexe
✨ Build a machine learning model from a prompt |
2.6K | 255 | 03/06/26 | Apache-2.0 |
| 135 |
finetrainers
Scalable and memory-optimized training of diffusion models |
1.4K | 142 | 05/26/26 | Apache-2.0 |
| 136 |
foolbox
A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX |
3.0K | 442 | 12/03/25 | MIT |
| 137 |
keras-tuner
A Hyperparameter Tuning Library for Keras |
2.9K | 404 | 12/01/25 | Apache-2.0 |
| 138 |
evaluation-guidebook
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval! |
2.1K | 126 | 12/03/25 | Other |
| 139 |
alibi
Algorithms for explaining machine learning models |
2.6K | 266 | 10/17/25 | Other |
| 140 |
GongBU
Paper accepted by CIKM 2024. Codes of GongBU, a LLM fine-tuning platform for domain-specific adaptation. |
1.2K | 90 | 01/22/26 | MIT |