#82 · Primary category: MLOps & Evaluation
langwatch
The platform for LLM evaluations and AI agent testing
Project last updated:08/29/26
GitHub Stars
3.5K
Forks
356
Contributors
37
License
Apache-2.0
Why we included this project
LangWatch is one of the few open-source platforms that handles LLM evaluation, observability, and agent testing in one place, so teams can catch regressions before release and have somewhere concrete to look when something fails in production. Rather than wiring separate tracing, dataset, and evaluation tools together, you get a self-hostable workspace where traces feed datasets, offline and online evals run against them, and prompt or model changes get compared and optimized in place. Its agent simulator is the standout: it runs realistic multi-step scenarios against your full stack, including tools, state, a simulated user, and a judge, then points to the exact decision where an agent broke. That gives engineering teams shipping agents a repeatable, evidence-based way to test behavior instead of smoke-testing, and it gives anyone running production LLM workloads a concrete trail when a prompt silently degrades.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models