MLOps & Evaluation

Experiment tracking, deployment, monitoring, and model evaluation.

187 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

101–120 of 187

Rank Project Stars Forks
101 elyra

Elyra extends JupyterLab with an AI centric approach.

2.0K 369
102 evaluate

🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

2.5K 338
103 katib

Automated Machine Learning on Kubernetes

1.7K 538
104 AIF360

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

2.9K 912
105 responsible-ai-toolbox

A suite of tools for model and data exploration and assessment to enable responsible AI.

1.8K 494
106 mlrun

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.

1.7K 317
107 expect

Expect tests your agent's code in a real browser

3.6K 156
108 tau2-bench

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

1.9K 480
109 WFGY

WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.

1.8K 165
110 InferenceX

Open-source continuous inference benchmark research platform comparing leading LLMs across NVIDIA, AMD, and future TPU hardware with live dashboards.

1.6K 277
111 training

Reference implementations of MLPerf® training benchmarks

1.8K 593
112 hallucination-leaderboard

Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents

3.3K 109
113 claimed

The goal of CLAIMED is to enable low-code/no-code rapid prototyping style programming to seamlessly CI/CD into production.

2.3K 3.9K
114 VBench

[CVPR2024 Highlight] VBench - We Evaluate Video Generation

1.8K 133
115 openinference

OpenTelemetry Instrumentation for AI Observability

1.2K 301
116 evals-skills

Skills for AI Evals to compliment the course: AI Evals For Engineers & PMs

1.7K 170
117 CLUE

中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard

4.3K 544
118 VectorDBBench

Benchmark for vector databases.

1.2K 432
119 pycm

Multi-class confusion matrix library in Python

1.5K 126
120 mlops-python-package

A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.

1.4K 200