MLOps & Evaluation

Experiment tracking, deployment, monitoring, and model evaluation.

187 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

121–140 of 187

Rank Project Stars Forks
121 skillsbench

SkillsBench evaluates how well skills work and how effective agents are at using them.

1.7K 363
122 AngelSlim

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

1.6K 174
123 waza

CLI / Framework for Agent Skills - create, test, measure and improve skill quality and effectiveness

1.3K 79
124 tensorwatch

Debugging, monitoring and visualization for Python Machine Learning and Data Science

3.5K 361
125 gemma-tuner-multimodal

Fine-tune Gemma 4 and 3n with audio, images and text on Apple Silicon, using PyTorch and Metal Performance Shaders.

1.5K 103
126 Skills

A project to improve skills of large language models

1.0K 198
127 Tracely-ai

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

1.2K 104
128 GraphGen

GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

1.2K 97
129 judgeval

The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

1.1K 96
130 AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3.7K 278
131 eli5

A library for debugging/inspecting machine learning classifiers and explaining their predictions

2.8K 325
132 deepchecks

Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and models from research to production.

4.0K 304
133 sacred

Sacred is a tool to help you configure, organize, log and reproduce experiments developed at IDSIA.

4.4K 393
134 plexe

✨ Build a machine learning model from a prompt

2.6K 255
135 finetrainers

Scalable and memory-optimized training of diffusion models

1.4K 142
136 foolbox

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

3.0K 442
137 keras-tuner

A Hyperparameter Tuning Library for Keras

2.9K 404
138 evaluation-guidebook

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

2.1K 126
139 alibi

Algorithms for explaining machine learning models

2.6K 266
140 GongBU

Paper accepted by CIKM 2024. Codes of GongBU, a LLM fine-tuning platform for domain-specific adaptation.

1.2K 90