skill-up
A CLI to evaluate and evolve agent skills: declarative YAML cases, runs across engines and structured reports
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- The CLI runs skills on agent engines and executes grading scripts
- Runs and the agent judge call the model and spend tokens
Install
Manual install
curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bashInstall the skill-up CLI directly. Usually skill-upper installs the CLI itself on first run.
This is third-party code. Review the repository files before installing.
What it does
skill-up is a CLI that makes skill quality measurable and repeatable. Test cases are declared in YAML, run across multiple agent engines and graded by rules, scripts or an agent judge, with results collected into reports locally or in CI. It ships the skill-upper skill: through conversation it reads failures, repairs and expands the case set, reruns skill-up and repeats the loop. It imports Anthropic's evals.json format and produces compatible grading.json and benchmark reports.
Who it is for. For developers and QA engineers who build agent skills and want to test them with regressions.
Good fit when
- You need to measure skill quality with tests, not by eye
- You need regression cases for a skill in CI
- You want to move an Anthropic evals.json into a repeatable run
Not a fit when
- You do not build skills, you only use them
- You have no access to an agent engine to run the cases
Example request
Write an eval for my skill and run it through skill-up, then show the failure reportLimitations
Running cases needs an agent engine: Claude Code, Codex and Qoder CLI are built in, plus custom engines. Agent-judge grading and the runs themselves spend model tokens of whichever engine you connect.
How to disable. Remove the skill-up binary from PATH and the skill-upper folder from the agent's skills directory.
Security check
- The CLI runs skills on agent engines and executes grading scripts
- Runs and the agent judge call the model and spend tokens
README in short
The README describes skill-up as an evaluation and evolution tool for skills, written in Go. Evaluation is set through eval.yaml and case files, runs on Claude Code, Codex and Qoder CLI engines plus custom ones, and grading can be rule, script or agent based. Reports are Anthropic compatible, with evals.json import and CI modes. The recommended path is the skill-upper skill, which repairs and expands cases and reruns the loop through conversation. Apache-2.0 licensed.
FAQ
Do I install skill-up before skill-upper?
Usually no. skill-upper checks for the CLI at runtime and guides the agent through installing it if missing.
How are answers graded?
By three strategies: rule based, script and agent judge, with results collected into structured reports.
Related
A CLI for testing and red teaming LLM applications: compares models, runs automated CI checks and finds vulnerabilities
Playwright CLI
playwright-cli
Microsoft's official Playwright CLI with an agent skill: drive a browser through short commands without heavy MCP schemas
The official MCP server debugger: web UI, automation CLI and terminal UI in one package
SonarSource's official MCP server: quality and security issues, quality gates and code analysis from SonarQube Server and Cloud