MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 41 |
cookiecutter-data-science
A logical, reasonably standardized, but flexible project structure for doing and sharing data science work. |
10.0K | 2.6K | 08/07/26 | MIT |
| 42 |
pycaret
Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane. |
9.8K | 1.8K | 07/23/26 | Other |
| 43 |
oumi
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM! |
9.4K | 787 | 08/28/26 | Apache-2.0 |
| 44 |
mage-ai
🧙 Build, run, and manage data pipelines for integrating and transforming data. |
8.8K | 991 | 08/13/26 | Apache-2.0 |
| 45 |
feast
The Open Source Feature Store for AI/ML |
7.2K | 1.4K | 08/28/26 | Apache-2.0 |
| 46 |
h2o-3
H2O is an open-source, distributed in-memory machine learning platform with AutoML, supporting Python, R, Java, and big data ecosystems like Hadoop and Spark. |
7.5K | 2.0K | 08/26/26 | Apache-2.0 |
| 47 |
flyte
Dynamic, resilient AI orchestration. Coordinate data, models, and compute as you build AI workflows. |
7.3K | 879 | 08/29/26 | Apache-2.0 |
| 48 |
opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. |
7.4K | 854 | 08/27/26 | Apache-2.0 |
| 49 |
evidently
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics. |
7.9K | 904 | 08/05/26 | Apache-2.0 |
| 50 |
tensorboardX
tensorboard for pytorch (and chainer, mxnet, numpy, ...) |
8.0K | 852 | 07/14/26 | MIT |
| 51 |
clearml
ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution |
6.8K | 797 | 08/23/26 | Apache-2.0 |
| 52 |
interpret
Fit interpretable models. Explain blackbox machine learning. |
6.9K | 788 | 08/24/26 | MIT |
| 53 |
aim
Aim 💫 — An easy-to-use & supercharged open-source experiment tracker. |
6.2K | 409 | 08/29/26 | Apache-2.0 |
| 54 |
chinese-llm-benchmark
A live Chinese LLM benchmark evaluating 395+ models across 300 dimensions, offering leaderboards and a 2M+ defect library for research. |
6.4K | 262 | 08/29/26 | Other |
| 55 |
lora-scripts
SD-Trainer. LoRA & Dreambooth training scripts & GUI use kohya-ss's trainer, for diffusion model. |
6.1K | 697 | 08/21/26 | AGPL-3.0 |
| 56 |
giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents |
5.8K | 527 | 08/28/26 | Apache-2.0 |
| 57 |
ClawWork
"ClawWork: OpenClaw as Your AI Coworker - 💰 $15K earned in 11 Hours" |
8.5K | 1.1K | 03/03/26 | MIT |
| 58 |
rllm
Democratizing Reinforcement Learning for LLMs |
5.8K | 613 | 08/24/26 | Apache-2.0 |
| 59 |
zenml
ZenML 🙏: One AI Platform from Pipelines to Agents. https://zenml.io. |
5.6K | 655 | 08/29/26 | Apache-2.0 |
| 60 |
coze-loop
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring. |
5.7K | 795 | 08/29/26 | Apache-2.0 |