MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 101 |
elyra
Elyra extends JupyterLab with an AI centric approach. |
2.0K | 369 | 08/19/26 | Apache-2.0 |
| 102 |
evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets. |
2.5K | 338 | 07/06/26 | Apache-2.0 |
| 103 |
katib
Automated Machine Learning on Kubernetes |
1.7K | 538 | 08/27/26 | Apache-2.0 |
| 104 |
AIF360
A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models. |
2.9K | 912 | 06/15/26 | Apache-2.0 |
| 105 |
responsible-ai-toolbox
A suite of tools for model and data exploration and assessment to enable responsible AI. |
1.8K | 494 | 08/28/26 | MIT |
| 106 |
mlrun
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications. |
1.7K | 317 | 08/28/26 | Apache-2.0 |
| 107 |
expect
Expect tests your agent's code in a real browser |
3.6K | 156 | 05/06/26 | Other |
| 108 |
tau2-bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains |
1.9K | 480 | 08/27/26 | MIT |
| 109 |
WFGY
WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg. |
1.8K | 165 | 08/29/26 | Other |
| 110 |
InferenceX
Open-source continuous inference benchmark research platform comparing leading LLMs across NVIDIA, AMD, and future TPU hardware with live dashboards. |
1.6K | 277 | 08/29/26 | Apache-2.0 |
| 111 |
training
Reference implementations of MLPerf® training benchmarks |
1.8K | 593 | 08/17/26 | Apache-2.0 |
| 112 |
hallucination-leaderboard
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents |
3.3K | 109 | 05/11/26 | Apache-2.0 |
| 113 |
claimed
The goal of CLAIMED is to enable low-code/no-code rapid prototyping style programming to seamlessly CI/CD into production. |
2.3K | 3.9K | 07/07/26 | Apache-2.0 |
| 114 |
VBench
[CVPR2024 Highlight] VBench - We Evaluate Video Generation |
1.8K | 133 | 08/21/26 | Apache-2.0 |
| 115 |
openinference
OpenTelemetry Instrumentation for AI Observability |
1.2K | 301 | 08/29/26 | Apache-2.0 |
| 116 |
evals-skills
Skills for AI Evals to compliment the course: AI Evals For Engineers & PMs |
1.7K | 170 | 08/16/26 | MIT |
| 117 |
CLUE
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard |
4.3K | 544 | 02/06/26 | Other |
| 118 |
VectorDBBench
Benchmark for vector databases. |
1.2K | 432 | 08/27/26 | MIT |
| 119 |
pycm
Multi-class confusion matrix library in Python |
1.5K | 126 | 08/17/26 | MIT |
| 120 |
mlops-python-package
A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines. |
1.4K | 200 | 08/24/26 | MIT |