#82 · Primary category: MLOps & Evaluation

langwatch

ai analytics datasets dspy evaluation gpt llm llm-ops llmops low-code observability openai prompt-engineering

The platform for LLM evaluations and AI agent testing

Project last updated:08/29/26

GitHub Stars

3.5K

Forks

356

Contributors

37

License

Apache-2.0

Why we included this project

LangWatch is one of the few open-source platforms that handles LLM evaluation, observability, and agent testing in one place, so teams can catch regressions before release and have somewhere concrete to look when something fails in production. Rather than wiring separate tracing, dataset, and evaluation tools together, you get a self-hostable workspace where traces feed datasets, offline and online evals run against them, and prompt or model changes get compared and optimized in place. Its agent simulator is the standout: it runs realistic multi-step scenarios against your full stack, including tools, state, a simulated user, and a judge, then points to the exact decision where an agent broke. That gives engineering teams shipping agents a repeatable, evidence-based way to test behavior instead of smoke-testing, and it gives anyone running production LLM workloads a concrete trail when a prompt silently degrades.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category