skill-up

A CLI to evaluate and evolve agent skills: declarative YAML cases, runs across engines and structured reports

CLI

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • The CLI runs skills on agent engines and executes grading scripts
  • Runs and the agent judge call the model and spend tokens
All reasons and checks

alibaba/skill-up

Install

Manual install

curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash

Install the skill-up CLI directly. Usually skill-upper installs the CLI itself on first run.

This is third-party code. Review the repository files before installing.

What it does

skill-up is a CLI that makes skill quality measurable and repeatable. Test cases are declared in YAML, run across multiple agent engines and graded by rules, scripts or an agent judge, with results collected into reports locally or in CI. It ships the skill-upper skill: through conversation it reads failures, repairs and expands the case set, reruns skill-up and repeats the loop. It imports Anthropic's evals.json format and produces compatible grading.json and benchmark reports.

Who it is for. For developers and QA engineers who build agent skills and want to test them with regressions.

Good fit when

  • You need to measure skill quality with tests, not by eye
  • You need regression cases for a skill in CI
  • You want to move an Anthropic evals.json into a repeatable run

Not a fit when

  • You do not build skills, you only use them
  • You have no access to an agent engine to run the cases

Example request

Write an eval for my skill and run it through skill-up, then show the failure report

Limitations

Running cases needs an agent engine: Claude Code, Codex and Qoder CLI are built in, plus custom engines. Agent-judge grading and the runs themselves spend model tokens of whichever engine you connect.

How to disable. Remove the skill-up binary from PATH and the skill-upper folder from the agent's skills directory.

Security check

  • The CLI runs skills on agent engines and executes grading scripts
  • Runs and the agent judge call the model and spend tokens

README in short

The README describes skill-up as an evaluation and evolution tool for skills, written in Go. Evaluation is set through eval.yaml and case files, runs on Claude Code, Codex and Qoder CLI engines plus custom ones, and grading can be rule, script or agent based. Reports are Anthropic compatible, with evals.json import and CI modes. The recommended path is the skill-upper skill, which repairs and expands cases and reruns the loop through conversation. Apache-2.0 licensed.

FAQ

Do I install skill-up before skill-upper?

Usually no. skill-upper checks for the CLI at runtime and guides the agent through installing it if missing.

How are answers graded?

By three strategies: rule based, script and agent judge, with results collected into structured reports.

Editors’ pick

A CLI for testing and red teaming LLM applications: compares models, runs automated CI checks and finds vulnerabilities

CLIMedium risk25.5KRepository stars

Playwright CLI

playwright-cli

Editors’ pick

Microsoft's official Playwright CLI with an agent skill: drive a browser through short commands without heavy MCP schemas

CLIMedium riskNo VPN needed13.6KRepository stars
Editors’ pick

The official MCP server debugger: web UI, automation CLI and terminal UI in one package

CLIMedium riskNo VPN needed11KRepository stars
Official

SonarSource's official MCP server: quality and security issues, quality gates and code analysis from SonarQube Server and Cloud

MCP serverMedium risk655Repository stars
Foxx AIskill-up

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.