#127 · Primary category: MLOps & Evaluation
Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
Project last updated:08/29/26
GitHub Stars
1.2K
Forks
104
Contributors
7
License
MIT
Why we included this project
Most agent-observability tools stop at showing you what broke; Tracely acts on that result instead. It ingests agent traces over OTLP, scores each one with LLM-as-judge and structural checks, clusters repeated failures, and turns a bad run into a reproducible regression case that replays in CI with no API keys and no model spend. That case then blocks the pull request that would reintroduce the failure, and alerts land in Slack, email, or your own webhook. The conversation-grouped waterfall and per-agent metrics help teams running multi-agent setups see exactly which agent misfired and whether a fix held. The stack self-hosts with Postgres, ClickHouse, Redis, and MinIO, so teams that want evaluation and telemetry data to stay in their own infrastructure can keep it that way.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models