OPFOR
Open-source adversary emulation for AI agents, LLM apps and MCP servers, checked against OWASP catalogs
High risk
We rate an entry high when the tool writes to external systems, handles money, production databases or secrets, or runs arbitrary commands. The CLI installs it only with your consent.
Why this level
- Generates and fires attacks at a target, use only on your own systems and with permission
- Needs a model provider key, keep it out of chats and repositories
- Consumes provider tokens and money on runs
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add cli/opforDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Install the tool
npm install -g @keyvaluesystems/agent-opfor-cliInstall the CLI: npm install -g @keyvaluesystems/agent-opfor-cli, set the provider key as an environment variable, then run opfor run. Test only targets you own.
Other ways from the author
npm install -g @keyvaluesystems/agent-opfor-cliAfter install set a provider key via an environment variable and run opfor run. Test only targets you own.
This is third-party code. Review the repository files before installing.
What it does
OPFOR tests an AI agent the way an attacker would: prompts, tools, MCP servers, memory and multi-turn reasoning. It generates targeted attacks against the OWASP LLM Top 10, Agentic AI Top 10, MCP Top 10 and API Security catalogs, fires them at your target and judges each response with a separate model. Every attack prompt, request, response and verdict is logged, so a run is reproducible and auditable. It runs from the CLI, as an MCP server, through IDE skills, a browser extension or an SDK. All modes share the same evaluators and judge logic.
Who it is for. For teams shipping AI agents who want to test their resilience to attacks on systems they own.
Good fit when
- You need to test your own agent or LLM app for attack resilience
- You need an MCP server checked against OWASP catalogs
- You need a reproducible report with attack logs and verdicts for an audit
Not a fit when
- You lack written permission from the target owner, testing others' systems is off limits
- You have no model provider key, which is needed to generate attacks and judge
Example request
Run a red-team check of my agent against the OWASP catalogs and produce a report with verdictsLimitations
The tool runs only against systems you own and with explicit permission, using it on others' resources is not acceptable. It needs a model provider key, by default OpenAI, Gemini or Anthropic, whose access from some regions is not guaranteed. Judging is done by a model and does not replace a manual security review.
How to disable. Remove the package: npm uninstall -g @keyvaluesystems/agent-opfor-cli, delete the .opfor folder and the MCP entry if you added one.
Security check
- Generates and fires attacks at a target, use only on your own systems and with permission
- Needs a model provider key, keep it out of chats and repositories
- Consumes provider tokens and money on runs
README in short
The README describes OPFOR as an open adversary-emulation tool for AI agents, LLM apps and MCP servers. It covers the OWASP LLM, Agentic AI, MCP and API Security catalogs plus bias suites, generates attacks, fires them at a target and judges with a model. Five run modes: CLI, browser extension, MCP server, skills and SDK, all on shared evaluators. Every step is logged for reproducibility and audit, with trace integration through Langfuse and Netra. Apache 2.0 licensed.
FAQ
What run modes are there?
CLI, an MCP server for running from Cursor or Claude Desktop, IDE skills, a browser extension and an SDK. Evaluators and judge logic are shared.
Do I need a model key?
Yes, a provider key such as OpenAI, Gemini or Anthropic. It is set through an environment variable.
Related
A code security audit skill by Cloudflare: the agent runs recon, coverage-led hunting and independent verification of findings, then produces a structured repor
NVIDIA's open stack for running OpenClaw, Hermes and LangChain Deep Agents in OpenShell sandboxes with network policy and managed inference
Security scanner for agent skills and MCP servers: finds prompt injection, data exfiltration and supply chain risks before install
Static code analysis with rules that look like source code, plus a built-in MCP server for AI agents