Codex Autoresearch

A Codex skill that runs a loop of measurable experiments over a repository toward a numeric target: change, verify, revert failures, repeat

Skill

High risk

We rate an entry high when the tool writes to external systems, handles money, production databases or secrets, or runs arbitrary commands. The CLI installs it only with your consent.

Why this level

  • Autonomously changes code and creates and reverts git commits
  • Runs arbitrary metric commands and defaults to full access
All reasons and checks

leo-lilinxiao/codex-autoresearch

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add skills/codex-autoresearch

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a claude-code -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a cursor -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a github-copilot -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a codex -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a gemini-cli -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a cline -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a roo -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

A fork of Roo Code, same .roo folders.

Run in a terminal in the project folder

npx skills add leo-lilinxiao/codex-autoresearch --skill codex-autoresearch -a opencode -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

You will need: Node.js

Checked against the repository on Sep 24, 2026, commit 0f54c57.

Text for your agent

In Codex run $skill-installer install https://github.com/leo-lilinxiao/codex-autoresearch, open a clean git branch with full access, then invoke $codex-autoresearch and name the metric, target, edit scope and mode.

Other ways from the author
$skill-installer install https://github.com/leo-lilinxiao/codex-autoresearch

The command is typed inside Codex through the built in skill installer. Then open the target repository and invoke $codex-autoresearch.

This is third-party code. Review the repository files before installing.

What it does

The skill gives Codex a discipline for autonomous optimization of a repository toward a chosen numeric metric. The user names a measurable result, for example zero test failures, fewer warnings, lower latency or binary size, and the skill repeats a loop: change one thing, measure with a verify command, keep the improvement or revert with git revert. A control script owns commits, verification, rollback and the event history, while Codex owns hypotheses and code changes. Before the first write the skill confirms the goal, edit scopes, the metric command, an optional regression guard and the run mode, foreground or background. A complete status is set only when the retained metric reaches the confirmed target, and the full history is written to autoresearch-results.

Who it is for. For developers and quality engineers who need autonomous iterative optimization of code toward a measurable goal.

Good fit when

  • You need to drive a metric to a target: zero failing tests, fewer warnings, lower latency
  • You want a long autonomous run of edits with verification and rollback at each step
  • You need a reproducible experiment history with commits and a log

Not a fit when

  • You need a one-off edit rather than a loop of measurable experiments
  • The result cannot be expressed by a single numeric verify command
  • There is no clean git branch or git cannot be used

Example request

Reduce error_count from python3 scripts/score.py to zero, only touch src, run in background

Limitations

It works only with Codex and needs Python 3.11 or newer, git and a configured author identity. One run is one repository, one metric and one clean named branch; without git the skill does not work, since git is the experiment memory and the rollback boundary. The metric command must exit successfully and print one finite number or a JSON object with an explicitly named key. Background runs default to full access because each step creates or reverts a commit.

How to disable. Remove the codex-autoresearch folder from the skills directory, usually .agents/skills in the project or ~/.agents/skills. The autoresearch-results directory in the project repository can be removed separately.

Security check

  • Autonomously changes code and creates and reverts git commits
  • Runs arbitrary metric commands and defaults to full access

README in short

The README describes an autonomous system for Codex that iteratively improves a repository toward a numeric target, inspired by Karpathy's autoresearch idea. The user names a measurable result, Codex inspects the repository, confirms the experiment, changes one thing, verifies and keeps a success or reverts a failure. Installation goes through $skill-installer or manual copying into the skills directory, and a run opens on a clean git branch with full access. Run artifacts live in autoresearch-results: an immutable config, an event log, logs and an optional HTML report. Every trial is a commit, any failed check is reverted with git revert, and a strict safety model stops the run on out-of-scope edits or failures. MIT licensed.

SKILL.md

---
name: codex-autoresearch
description: "Run repeated, measured Git experiments toward a numeric target; keep improvements and revert failures. Use for autonomous optimization or managing an autoresearch run, not one-shot edits."
metadata:
  short-description: "Run measurable autonomous experiments"
---

# Codex Autoresearch

Improve a repository through repeated, reversible experiments:

`hypothesize -> change -> measure -> learn -> keep or revert -> repeat`

Codex chooses hypotheses and makes code changes. The control script owns measurement, commits, rollback, and the event history.

## Respond To The Request

Resolve `<control>` to this skill's own `scripts/autoresearch.py`; do not assume it is installed in the target repository.

For a status or results request, run the corresponding command directly:

| Request | Command |
|---|---|
| Status | `python3 <control> status --repo <repo>` |
| History | `python3 <control> history --repo <repo>` |
| TSV export | `python3 <control> history --repo <repo> --format tsv` |
| HTML report | `python3 <control> report --repo <repo>` |

## Experiment Boundary

- One run owns one Git repository, one numeric metric, one confirmed target, and approved path scopes.
- Use `finish` to finalize each coherent experiment.
- Only a verified target can mean `complete`; an iteration limit, error, or external blocker has its own status.

FAQ

Does installation change Codex settings?

No, installation only copies the skill files. Foreground runs need a current Codex release with goal support.

Can I stop and resume a run?

Yes. A foreground run pauses through a Codex goal, and a background run is controlled through $codex-autoresearch: status, stop or resume the confirmed experiment.

Editors’ pick

A skills library that gives coding agents a development process: brainstorming, planning, TDD, subagents and code review

PluginMedium riskNo VPN needed292.5KRepository stars
Editors’ pick

Small composable skills for engineering with agents: plan grilling, TDD, bug diagnosis, code review and architecture

SkillLow risk271.4KRepository stars
Editors’ pick

GitHub toolkit for spec-driven development: the specify CLI adds agent commands and skills to a project, from principles to implementation

CLIMedium riskNo VPN needed139.3KRepository stars

Reference MCP servers

Model Context Protocol servers

Official

Official reference MCP servers: Filesystem, Fetch, Git, Memory, Sequential Thinking, Time and Everything

MCP serverMedium risk90.6KRepository stars
Foxx AICodex Autoresearch

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.