MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 21 |
dvc
🦉 Data Versioning and ML Experiments |
15.8K | 1.3K | 08/24/26 | Apache-2.0 |
| 22 |
kubeflow
Machine Learning Toolkit for Kubernetes |
15.8K | 2.7K | 08/21/26 | Apache-2.0 |
| 23 |
optuna
A hyperparameter optimization framework |
14.7K | 1.4K | 08/27/26 | MIT |
| 24 |
scalene
Scalene: a high-performance, high-precision CPU, GPU, and memory profiler for Python with AI-powered optimization proposals |
13.5K | 436 | 08/27/26 | Apache-2.0 |
| 25 |
ragas
Supercharge Your LLM Application Evaluations 🚀 |
15.5K | 1.7K | 02/24/26 | Apache-2.0 |
| 26 |
RagaAI-Catalyst
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view |
16.2K | 3.6K | 02/11/26 | Apache-2.0 |
| 27 |
easy-dataset
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval |
14.9K | 1.5K | 05/01/26 | Other |
| 28 |
great_expectations
Always know what to expect from your data. |
11.7K | 1.8K | 08/28/26 | Apache-2.0 |
| 29 |
bisheng
Open-source LLM DevOps platform for enterprise AI applications, with GenAI workflow, RAG, agents, model management, evaluation, and observability. |
11.9K | 2.0K | 08/28/26 | Apache-2.0 |
| 30 |
pest
The elegant testing framework for PHP developers and AI agents. |
11.7K | 529 | 08/26/26 | MIT |
| 31 |
ai-toolkit
The ultimate training toolkit for finetuning diffusion models |
11.8K | 1.5K | 08/29/26 | MIT |
| 32 |
wandb
The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production. |
11.2K | 891 | 08/29/26 | MIT |
| 33 |
phoenix
AI Observability & Evaluation |
11.2K | 1.1K | 08/29/26 | Other |
| 34 |
kedro
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular. |
11.0K | 1.1K | 08/28/26 | Other |
| 35 |
iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds. |
11.4K | 1.2K | 08/28/26 | Apache-2.0 |
| 36 |
inspector
Visual testing tool for MCP servers |
10.8K | 1.5K | 08/29/26 | Other |
| 37 |
autogluon
Fast and Accurate ML in 3 Lines of Code |
10.6K | 1.2K | 08/28/26 | Apache-2.0 |
| 38 |
visdom
Tool for real-time visualization, monitoring and collaborative analysis of AI/ML experiments and live data. Supports Python, PyTorch/Torch, NumPy, TensorFlow/Keras https://visdom.dev |
10.3K | 1.2K | 08/29/26 | Apache-2.0 |
| 39 |
metaflow
Build, Manage and Deploy AI/ML Systems |
10.3K | 1.3K | 08/27/26 | Apache-2.0 |
| 40 |
oha
Ohayou(おはよう), HTTP load generator, inspired by rakyll/hey with tui animation. |
10.5K | 296 | 08/23/26 | MIT |