Codex Autoresearch
A Codex skill that runs a loop of measurable experiments over a repository toward a numeric target: change, verify, revert failures, repeat
High risk
We rate an entry high when the tool writes to external systems, handles money, production databases or secrets, or runs arbitrary commands. The CLI installs it only with your consent.
Why this level
- Autonomously changes code and creates and reverts git commits
- Runs arbitrary metric commands and defaults to full access
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add skills/codex-autoresearchDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a claude-code -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a cursor -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a github-copilot -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a codex -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a gemini-cli -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a cline -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a roo -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
A fork of Roo Code, same .roo folders.
Run in a terminal in the project folder
npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a opencode -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
In Codex run $skill-installer install https://github.com/leo-lilinxiao/codex-autoresearch, open a clean git branch with full access, then invoke $codex-autoresearch and name the metric, target, edit scope and mode.
Other ways from the author
$skill-installer install https://github.com/leo-lilinxiao/codex-autoresearchThe command is typed inside Codex through the built in skill installer. Then open the target repository and invoke $codex-autoresearch.
This is third-party code. Review the repository files before installing.
What it does
The skill gives Codex a discipline for autonomous optimization of a repository toward a chosen numeric metric. The user names a measurable result, for example zero test failures, fewer warnings, lower latency or binary size, and the skill repeats a loop: change one thing, measure with a verify command, keep the improvement or revert with git revert. A control script owns commits, verification, rollback and the event history, while Codex owns hypotheses and code changes. Before the first write the skill confirms the goal, edit scopes, the metric command, an optional regression guard and the run mode, foreground or background. A complete status is set only when the retained metric reaches the confirmed target, and the full history is written to autoresearch-results.
Who it is for. For developers and quality engineers who need autonomous iterative optimization of code toward a measurable goal.
Good fit when
- You need to drive a metric to a target: zero failing tests, fewer warnings, lower latency
- You want a long autonomous run of edits with verification and rollback at each step
- You need a reproducible experiment history with commits and a log
Not a fit when
- You need a one-off edit rather than a loop of measurable experiments
- The result cannot be expressed by a single numeric verify command
- There is no clean git branch or git cannot be used
Example request
Reduce error_count from python3 scripts/score.py to zero, only touch src, run in backgroundLimitations
It works only with Codex and needs Python 3.11 or newer, git and a configured author identity. One run is one repository, one metric and one clean named branch; without git the skill does not work, since git is the experiment memory and the rollback boundary. The metric command must exit successfully and print one finite number or a JSON object with an explicitly named key. Background runs default to full access because each step creates or reverts a commit.
How to disable. Remove the codex-autoresearch folder from the skills directory, usually .agents/skills in the project or ~/.agents/skills. The autoresearch-results directory in the project repository can be removed separately.
Security check
- Autonomously changes code and creates and reverts git commits
- Runs arbitrary metric commands and defaults to full access
README in short
The README describes an autonomous system for Codex that iteratively improves a repository toward a numeric target, inspired by Karpathy's autoresearch idea. The user names a measurable result, Codex inspects the repository, confirms the experiment, changes one thing, verifies and keeps a success or reverts a failure. Installation goes through $skill-installer or manual copying into the skills directory, and a run opens on a clean git branch with full access. Run artifacts live in autoresearch-results: an immutable config, an event log, logs and an optional HTML report. Every trial is a commit, any failed check is reverted with git revert, and a strict safety model stops the run on out-of-scope edits or failures. MIT licensed.
SKILL.md
--- name: codex-autoresearch description: "Run repeated, measured Git experiments toward a numeric target; keep improvements and revert failures. Use for autonomous optimization or managing an autoresearch run, not one-shot edits." metadata: short-description: "Run measurable autonomous experiments" --- # Codex Autoresearch Improve a repository through repeated, reversible experiments: `hypothesize -> change -> measure -> learn -> keep or revert -> repeat` Codex chooses hypotheses and makes code changes. The control script owns measurement, commits, rollback, and the event history. ## Respond To The Request Resolve `<control>` to this skill's own `scripts/autoresearch.py`; do not assume it is installed in the target repository. For a status or results request, run the corresponding command directly: | Request | Command | |---|---| | Status | `python3 <control> status --repo <repo>` | | History | `python3 <control> history --repo <repo>` | | TSV export | `python3 <control> history --repo <repo> --format tsv` | | HTML report | `python3 <control> report --repo <repo>` | ## Experiment Boundary - One run owns one Git repository, one numeric metric, one confirmed target, and approved path scopes. - Use `finish` to finalize each coherent experiment. - Only a verified target can mean `complete`; an iteration limit, error, or external blocker has its own status.
FAQ
Does installation change Codex settings?
No, installation only copies the skill files. Foreground runs need a current Codex release with goal support.
Can I stop and resume a run?
Yes. A foreground run pauses through a Codex goal, and a background run is controlled through $codex-autoresearch: status, stop or resume the confirmed experiment.
Related
A skills library that gives coding agents a development process: brainstorming, planning, TDD, subagents and code review
Skills for real engineers by Matt Pocock
Skills For Real Engineers
Small composable skills for engineering with agents: plan grilling, TDD, bug diagnosis, code review and architecture
GitHub toolkit for spec-driven development: the specify CLI adds agent commands and skills to a project, from principles to implementation
Reference MCP servers
Model Context Protocol servers
Official reference MCP servers: Filesystem, Fetch, Git, Memory, Sequential Thinking, Time and Everything