LongHorizon-Harness

A CLI that keeps an existing agent on a long task: plan a step, run it, verify the result in the real environment and checkpoint progress

CLI

High risk

We rate an entry high when the tool writes to external systems, handles money, production databases or secrets, or runs arbitrary commands. The CLI installs it only with your consent.

Why this level

  • It drives an agent that runs commands and edits files for hours on end
  • It controls the desktop and terminal with real actions
  • It acts in the launch directory, that is your real project
All reasons and checks
Needs a VPN

amap-ml/longhorizon-harness

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add cli/longhorizon-harness

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Install the tool

pipx install lh-harness

Without pipx, pip install lh-harness works too.

You will need: pipx

Checked against the repository on Sep 25, 2026, commit a1dd930.

Text for your agent

Install the CLI: uv tool install lh-harness (or pip install lh-harness). Then lh-harness init creates ./.lh-harness/config.toml, and a task runs with: lh-harness run --task "..." --agent codex.

Other ways from the author
uv tool install lh-harness

Alternatively: pip install lh-harness. Python 3.10 or later is required.

This is third-party code. Review the repository files before installing.

What it does

The tool wraps an existing agent (Claude Code, Codex, OpenCode, DeepSeek Harness) in a loop for long computer-use tasks. Each round it recovers the original goal and verified state, picks the next bounded step, runs it in a fresh context, checks the actual result against files, UI, logs and tests, then checkpoints progress or records failure evidence and continues. The roles are split: a manager holds state and the next step, an executor performs the step, an auditor verifies independently. It covers desktop and terminal work, a command-line run and a web dashboard, and each role can use its own model and backend.

Who it is for. For developers and devops engineers who need to drive an agent through long multi-hour tasks without losing state.

Good fit when

  • A task does not fit one agent pass and needs many verified steps
  • You need progress to survive context resets and failures
  • You want independent verification of the result rather than trusting the executor's report

Not a fit when

  • The task is solved in a single agent pass without a long loop
  • You have no access to a supported agent backend and its model

Example request

Carry this multi-hour task to completion: plan each step, verify the result in the real environment and checkpoint progress between rounds

Limitations

The tool does not train or replace a model; it builds a loop around an existing agent, so a configured backend (Claude Code, Codex, OpenCode or DeepSeek Harness) and access to its model are required. Python 3.10 or later is needed. DeepSeek Harness in phase one is CLI-only, with desktop control and MCP promised later. MIT licensed.

How to disable. Remove the package: uv tool uninstall lh-harness or pip uninstall lh-harness. The ./.lh-harness working folder can be deleted by hand.

Security check

  • It drives an agent that runs commands and edits files for hours on end
  • It controls the desktop and terminal with real actions
  • It acts in the launch directory, that is your real project

README in short

The README presents what the authors call loop engineering: the agent gets a goal once, and the tool repeatedly turns the remaining work into a bounded step, runs it on the right surface and verifies the actual result. Only what passes independent verification becomes trusted state; the rest stays failure evidence. It supports Claude Code, Codex, OpenCode and DeepSeek Harness, desktop and terminal work, a CLI run and a web dashboard. Installation is via uv tool or pip, with init, run, web and doctor commands.

FAQ

Is this a separate model or a new agent?

No. It is a loop around an existing agent: it coordinates roles, verified state and cross-round progress, while your agent does the actual work.

Can I assign different models to the roles?

Yes. The manager, executor and auditor can each use their own model and backend to balance quality, speed and cost.

Editors’ pick

Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins

CLIHigh risk390.7KRepository stars
Editors’ pick

A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself

CLIHigh risk249.8KRepository stars
Editors’ pick

An open source coding agent for the terminal and desktop with build and plan modes

CLIHigh risk210.6KRepository stars
Editors’ pick

Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent

PluginMedium risk171.5KRepository stars
Foxx AILongHorizon-Harness

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.