MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more. |
75.2K | 6.8K | 08/29/26 | Apache-2.0 |
| 2 |
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024) |
74.4K | 9.1K | 08/27/26 | Apache-2.0 |
| 3 |
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows |
46.6K | 17.7K | 08/29/26 | Apache-2.0 |
| 4 |
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23 |
33.9K | 3.7K | 08/29/26 | Other |
| 5 |
netron
Visualizer for neural network, deep learning and machine learning models |
33.4K | 3.2K | 08/29/26 | MIT |
| 6 |
mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data. |
27.7K | 6.2K | 08/29/26 | Apache-2.0 |
| 7 |
heretic
Fully automatic censorship removal for language models |
28.7K | 3.2K | 08/17/26 | AGPL-3.0 |
| 8 |
label-studio
Label Studio is a multi-type data labeling and annotation tool with standardized output format |
28.2K | 3.7K | 08/29/26 | Apache-2.0 |
| 9 |
shap
A game theoretic approach to explain the output of any machine learning model. |
25.7K | 3.7K | 08/29/26 | MIT |
| 10 |
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic. |
24.7K | 2.2K | 08/29/26 | MIT |
| 11 |
open-r1
Fully open reproduction of DeepSeek-R1 |
26.4K | 2.4K | 04/02/26 | Apache-2.0 |
| 12 |
datasets
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools |
21.9K | 3.4K | 08/28/26 | Apache-2.0 |
| 13 |
opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. |
21.7K | 1.7K | 08/29/26 | Apache-2.0 |
| 14 |
openobserve
Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment. |
21.6K | 1.1K | 08/29/26 | AGPL-3.0 |
| 15 |
taipy
Turns Data and AI algorithms into production-ready web applications in no time. |
19.4K | 2.0K | 08/10/26 | Apache-2.0 |
| 16 |
argo-workflows
Workflow Engine for Kubernetes |
16.9K | 3.6K | 08/29/26 | Apache-2.0 |
| 17 |
deepeval
The LLM Evaluation Framework |
18.0K | 1.9K | 08/29/26 | Apache-2.0 |
| 18 |
evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. |
19.3K | 3.1K | 04/14/26 | Other |
| 19 |
ComfyUI-Manager
ComfyUI extension for managing custom nodes with install, remove, disable, enable, and hub features. |
16.0K | 2.4K | 08/18/26 | GPL-3.0 |
| 20 |
dagster
An orchestration platform for the development, production, and observation of data assets. |
16.1K | 2.3K | 08/28/26 | Apache-2.0 |