#10 · Primary category: MLOps & Evaluation

promptfoo

ci ci-cd cicd evaluation evaluation-framework llm llm-eval llm-evaluation llm-evaluation-framework llmops pentesting prompt-engineering prompt-testing prompts rag red-teaming testing vulnerability-scanners

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Project last updated:08/29/26

GitHub Stars

24.7K

Forks

2.2K

Contributors

325

License

MIT

Why we included this project

Promptfoo is a test harness for LLM apps. You define declarative test cases in a config file, run them from the CLI or as a library, and get side-by-side output comparisons across prompts and providers. It also handles security: automated red-teaming and vulnerability scanning flag jailbreaks, prompt injection, and compliance risks. Because it runs locally and talks directly to the model APIs, you can wire it into CI/CD without sending your data anywhere. If you're comparing GPT, Claude, Gemini, or open-weight models, it's a straightforward way to spot regressions as your prompts change.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category