#166 · Primary category: MLOps & Evaluation
uptrain
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
Project last updated:08/18/24
GitHub Stars
2.4K
Forks
205
Contributors
40
License
Apache-2.0
Why we included this project
Teams shipping production LLM applications often have no reliable way to tell whether a model is quietly degrading until users complain. UpTrain gives them a structured answer: more than 20 preconfigured checks that grade responses on correctness, relevance, tone, code quality, and hallucination risk, plus 40 operators for building custom evals. When a check fails, the platform digs into the failure and points at likely root causes, so you are not stuck with a single low score and no idea what went wrong. That suits teams that want evaluation running continuously through development and monitoring instead of manual spot checks. It runs locally with a lightweight setup, which helps when you want evaluation close to your pipeline and data.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models