#127 · Primary category: MLOps & Evaluation

Tracely-ai

agent agent-observability ai-agents ci-cd clickhouse evals evaluation llm llm-as-judge llm-evaluation llm-observability llm-ops llmops mcp monitoring opentelemetry python self-hosted tracing

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

Project last updated:08/29/26

GitHub Stars

1.2K

Forks

104

Contributors

7

License

MIT

Why we included this project

Most agent-observability tools stop at showing you what broke; Tracely acts on that result instead. It ingests agent traces over OTLP, scores each one with LLM-as-judge and structural checks, clusters repeated failures, and turns a bad run into a reproducible regression case that replays in CI with no API keys and no model spend. That case then blocks the pull request that would reintroduce the failure, and alerts land in Slack, email, or your own webhook. The conversation-grouped waterfall and per-agent metrics help teams running multi-agent setups see exactly which agent misfired and whether a fix held. The stack self-hosts with Postgres, ClickHouse, Redis, and MinIO, so teams that want evaluation and telemetry data to stay in their own infrastructure can keep it that way.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category