#10 · Primary category: MLOps & Evaluation
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Project last updated:08/29/26
GitHub Stars
24.7K
Forks
2.2K
Contributors
325
License
MIT
Why we included this project
Promptfoo is a test harness for LLM apps. You define declarative test cases in a config file, run them from the CLI or as a library, and get side-by-side output comparisons across prompts and providers. It also handles security: automated red-teaming and vulnerability scanning flag jailbreaks, prompt injection, and compliance risks. Because it runs locally and talks directly to the model APIs, you can wire it into CI/CD without sending your data anywhere. If you're comparing GPT, Claude, Gemini, or open-weight models, it's a straightforward way to spot regressions as your prompts change.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models