MLOps & Evaluation
Experiment tracking, deployment, monitoring, and model evaluation.
187 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 81 |
Soup
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU. |
3.6K | 551 | 08/29/26 | Apache-2.0 |
| 82 |
langwatch
The platform for LLM evaluations and AI agent testing |
3.5K | 356 | 08/29/26 | Apache-2.0 |
| 83 |
evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. |
3.3K | 471 | 08/28/26 | Apache-2.0 |
| 84 |
sagemaker-python-sdk
A library for training and deploying machine learning models on Amazon SageMaker |
2.3K | 1.3K | 08/28/26 | Apache-2.0 |
| 85 |
shapash
🔅 Shapash: User-friendly Explainability and Interpretability to Develop Reliable and Transparent Machine Learning Models |
3.3K | 388 | 08/28/26 | Apache-2.0 |
| 86 |
lit
The Learning Interpretability Tool: Interactively analyze ML models to understand their behavior in an extensible and framework agnostic interface. |
3.7K | 369 | 07/29/26 | Apache-2.0 |
| 87 |
lmnr
Laminar - open-source observability platform purpose-built for AI agents. YC S24. |
3.2K | 228 | 08/29/26 | Apache-2.0 |
| 88 |
OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments |
3.1K | 529 | 08/21/26 | Apache-2.0 |
| 89 |
llm-datasets
Curated list of datasets and tools for post-training. |
4.8K | 394 | 04/29/26 | Other |
| 90 |
openlit
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs. |
2.7K | 366 | 08/29/26 | Apache-2.0 |
| 91 |
trainer
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes |
2.2K | 1.0K | 08/29/26 | Apache-2.0 |
| 92 |
hamilton
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does. |
2.6K | 213 | 08/19/26 | Apache-2.0 |
| 93 |
maestro
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL |
2.7K | 223 | 08/24/26 | Apache-2.0 |
| 94 |
inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps. |
2.2K | 271 | 08/29/26 | Other |
| 95 |
fairlearn
A Python package to assess and improve fairness of machine learning models. |
2.3K | 515 | 08/24/26 | MIT |
| 96 |
tfx
TFX is an end-to-end platform for deploying production ML pipelines |
2.2K | 725 | 08/17/26 | Apache-2.0 |
| 97 |
llm-foundry
LLM training code for Databricks foundation models |
4.4K | 590 | 03/25/26 | Apache-2.0 |
| 98 |
EvalAI
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI |
2.0K | 983 | 08/29/26 | Other |
| 99 |
cube-studio
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持 |
2.5K | 198 | 08/17/26 | Other |
| 100 |
elyra
Elyra extends JupyterLab with an AI centric approach. |
2.0K | 369 | 08/19/26 | Apache-2.0 |