#33 · Primary category: MLOps & Evaluation

phoenix

agents ai-monitoring ai-observability aiengineering anthropic datasets evals langchain llamaindex llm-eval llm-evaluation llmops llms openai prompt-engineering smolagents

AI Observability & Evaluation

Project last updated:08/29/26

GitHub Stars

11.2K

Forks

1.1K

Contributors

223

License

Other

Why we included this project

Most teams shipping LLM applications burn their time on the same problem: a prompt tweak, a model swap, or a retrieval change quietly breaks behavior, and nobody can see why. Phoenix fills in that gap by pairing OpenTelemetry-based tracing with an evaluation workflow, so you can pull up real runtime traces and then score the responses and retrieved context against criteria you define. It keeps versioned datasets and an experiment runner that tracks how prompt, model, and retrieval changes behave over time, plus a playground for comparing models and parameters before you commit. A built-in assistant can help debug individual traces, and a remote MCP endpoint lets tools like Claude Code and Cursor query your traces and datasets. Everything runs locally from a pip install, so a small team gets a full observability stack without standing up cloud infrastructure, and it stays agnostic to language and vendor across the usual frameworks and providers.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category