#166 · Primary category: MLOps & Evaluation

uptrain

autoevaluation evaluation experimentation hallucination-detection jailbreak-detection llm-eval llm-prompting llm-test llmops machine-learning monitoring openai-evals prompt-engineering root-cause-analysis

UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

Project last updated:08/18/24

GitHub Stars

2.4K

Forks

205

Contributors

40

License

Apache-2.0

Why we included this project

Teams shipping production LLM applications often have no reliable way to tell whether a model is quietly degrading until users complain. UpTrain gives them a structured answer: more than 20 preconfigured checks that grade responses on correctness, relevance, tone, code quality, and hallucination risk, plus 40 operators for building custom evals. When a check fails, the platform digs into the failure and points at likely root causes, so you are not stuck with a single low score and no idea what went wrong. That suits teams that want evaluation running continuously through development and monitoring instead of manual spot checks. It runs locally with a lightweight setup, which helps when you want evaluation close to your pipeline and data.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category