Promptfoo
A CLI for testing and red teaming LLM applications: compares models, runs automated CI checks and finds vulnerabilities
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Makes network calls to model APIs and can consume paid tokens
- Red team mode sends adversarial test attacks against your LLM application
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add cli/promptfooDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Install the tool
npm install -g promptfooInstall promptfoo: npm install -g promptfoo (or brew install promptfoo, or pip install promptfoo). Scaffold an example: promptfoo init --example getting-started, then run promptfoo eval and promptfoo view.
Other ways from the author
npm install -g promptfooAlso available via brew install promptfoo, pip install promptfoo, or a one-off npx promptfoo@latest run.
This is third-party code. Review the repository files before installing.
What it does
Promptfoo runs test suites against prompts, agents and RAG pipelines and compares outputs from different models against defined criteria. A separate red team mode simulates attacks on an LLM application and looks for issues like prompt leakage or guardrail bypass. Test suites are described in a declarative YAML config, and results open in a local web viewer or plug into CI. A Claude Code plugin ships a skill that helps write and maintain eval suites directly in conversation with the agent.
Who it is for. For developers and engineers who verify prompt quality and LLM application security before release.
Good fit when
- You need to compare several models' answers on the same prompt set
- You need regression checks for answer quality in CI
- You need to test an application for prompt injection and other vulnerabilities
Not a fit when
- You just need a one-off check of a single prompt without keeping history
- You don't have API access to the models you want to compare
Example request
Set up a promptfoo config comparing GPT and Claude on my system prompts and run the evalLimitations
Comparing models needs separate provider API keys (OpenAI, Anthropic and others), some of which are unavailable from Russia without a VPN. Red teaming and some checks require target configuration and can consume a noticeable amount of tokens on large test suites.
How to disable. Remove the promptfoo package (npm uninstall -g promptfoo or the brew/pip equivalent) and delete promptfooconfig.yaml files from the project. For the Claude Code plugin, remove it via /plugin.
Security check
- Makes network calls to model APIs and can consume paid tokens
- Red team mode sends adversarial test attacks against your LLM application
README in short
The README describes promptfoo as a CLI and library for evaluating and red teaming LLM applications. Main scenarios: comparing models and providers, regression checks in CI, vulnerability scanning and security reports. It installs via npm, brew or pip, and evaluations run locally without sending prompts to promptfoo's servers. The project recently joined OpenAI but stays open source under the MIT license.
FAQ
How is promptfoo different from manually testing prompts?
It runs the same test suite across multiple models and providers automatically and keeps a comparison history, instead of one-off checks in a chat window.
Do I need a promptfoo account to use it?
No, evaluations run locally with your own model API keys, no registration required.
Related
Playwright CLI
playwright-cli
Microsoft's official Playwright CLI with an agent skill: drive a browser through short commands without heavy MCP schemas
The official MCP server debugger: web UI, automation CLI and terminal UI in one package
SonarSource's official MCP server: quality and security issues, quality gates and code analysis from SonarQube Server and Cloud
Open-source AI agent for pull request review: descriptions, feedback and improvement suggestions on GitHub, GitLab and more