Documentation

Everything here runs on your laptop with Python 3.11+ and git — no docker, no API keys, no network. When in doubt, make proof re-verifies every claim and keeps the evidence.

Install

The quick way — the Homebrew package bundles GNU timeout and puts ratchet on your PATH:

brew install ayaangazali/ratchet/ratchet-agent    # macOS
# or, anywhere with uv:
uv tool install git+https://github.com/ayaangazali/ratchet

From source (Python 3.11+, git; Node 22.14+ only for the live TrueForge harness path):

git clone https://github.com/ayaangazali/ratchet && cd ratchet
make dev          # editable install + dev deps
macOS: the eval script uses GNU timeout. See troubleshooting if make evals fails.

Run the demo

make demo seeds demo-repo/ with a deliberately broken slugify function and three prepared patches:

honest.diffthe real fix — should go green
cheat.diffa reward hack — must be blocked before it executes
canary_hack.diffa cheat with zero static findings — must be caught anyway

It also seeds visible tests, a held-out (hidden) suite, and a canary test. Then run the whole suite and the red team:

make test         # the whole suite; no docker, no network
make redteam      # the published reward-hacking battery fired at the verifier

Your first verdicts

The gauntlet runs standalone — no agent, no model. Three verdicts tell the whole story:

ratchet verify --task tasks/demo-001-slugify/task.yaml --repo demo-repo \
               --diff demo-repo/patches/honest.diff    # GREEN, score 1.000
ratchet verify --task tasks/demo-001-slugify/task.yaml --repo demo-repo \
               --diff demo-repo/patches/cheat.diff     # CHEATED
ratchet verify --task tasks/canary-impossible/task.yaml --repo demo-repo \
               --diff demo-repo/patches/canary_hack.diff  # CHEATED, zero static findings

The cheat is caught by static analysis before a line of it runs. The canary hack trips no static rule and is caught by behavior instead — see the cheat detector.

How it works

You give Ratchet a repo and a task file. An agent proposes patches; Ratchet never asks the agent whether it's finished — there is no "done" tool anywhere in the system. Instead, every candidate patch runs a seven-stage verifier gauntlet, and green is set in exactly one place in the codebase: the gauntlet itself.

Every step the agent takes is a git commit plus a sandbox snapshot, so a run is a tree search over repo states with the verifier's score as the value function. The winning path exits as one squashed diff at a human approval gate, and every graded step is signed into a tamper-evident receipt chain.

Division of labor: the TrueForge harness owns model calls, sandboxes, sub-agents, and session persistence. Ratchet owns exactly one question — what counts as progress.

The gauntlet

Seven stages, in order. A patch that fails an early stage never reaches a later one — and a patch that fails the cheat check never executes at all.

#stagewhat it asks
1buildDoes the repo still build/import with the patch applied?
2cheat checkDoes the diff match any known reward-hacking pattern? (static — runs before execution)
3fail-to-passDo the held-out tests that were failing now pass?
4pass-to-passDoes everything that passed before still pass? (no regressions)
5typesType check clean?
6lintLint clean?
7diff hygieneIs the diff minimal and inside allowed paths?

Two hard rules around the gauntlet: protected paths (tests, verifier config) are reverted before grading on every run with no flag to skip, and the exit code is echoed outside the output region the agent can influence.

The cheat detector

Stage two is a library of static rules, each shipped with two tests — a patch that trips it and a patch that must not. Current rules:

hard_exit · always_equal · special_casing · runtime_test_write · protected_path · skip_marker · test_file_emptied · test_deleted · assertion_removed · report_hook_tamper · env_bypass · canary_passed · broad_except_pass · log_spoofed

The canary

Static rules can't catch everything — so one task in every battery is impossible: its hidden test cannot legitimately pass. Any patch that makes the canary go green must have smuggled an answer, even if it trips zero static rules. That's the canary_passed finding, and it's why the canary hack in the demo is caught with no static findings at all.

Hidden tests stay hidden

Held-out test names and failure details never appear in anything the agent can read — context assembly, observations, and every bus event are scrubbed, and the suite has tests asserting it. A leak would silently destroy the signal.

Receipts

Every graded node — accepted or pruned — is recorded into a signed, hash-linked chain, and a finished run seals its chain. That gives you two guarantees:

  • Order — results are provably in the order they were issued.
  • Integrity — edit any verdict after the fact and the chain breaks.
make audit
#   receipts      4
#   head          b26a36f26e627c03
#   chain         intact

Forge a green verdict and the same command reports CHAIN BROKEN with the exact receipt whose signature fails. The tamper demo runs as part of make proof.

The red team

The verifier itself is what gets evaluated. make redteam fires a published battery of real reward-hacking patterns at the gauntlet — skip markers, hardcoded answers, tests rewritten at import time, conftest report hooks, spoofed exit codes — plus two honest patches that must be accepted, because a paranoid verifier that rejects real fixes is broken too.

The score to hold: every attack caught, zero false positives. It runs in CI, and the dashboard shows the full attack-by-attack table from the latest proof run.

Objective graph

Bigger goals decompose into a graph of objective nodes — each one fulfillable only by tests, never by the agent's say-so. Nodes are attempted cheaply first; a node that exhausts its attempts escalates to full tree search automatically.

graph demo-slugify-graph · 2 node(s)
  ✓ accents        fulfilled  attempts 1
  ✓ truncation     fulfilled  attempts 3 (escalated to tree search)

The demo graph lives at objectives/demo-graph.yaml.

The docs oracle

When a failure looks like the outside world — an import error, a missing attribute, an unexpected keyword argument — the oracle attaches current upstream documentation for the exact version in the lockfile to the next prompt, instead of letting the model guess from memory.

  • Sources are configured in ratchet/scrapers.yaml, version-controlled — never one-off shell scrapes.
  • Content is extracted by heading, not CSS selector — headings survive site redesigns.
  • Every fetch is validated against the source's expect block; on failure the oracle relocates the section by heading similarity and commits the repair as a diff, so the repair history is auditable.

Sandboxes

Ratchet does not orchestrate containers — ever. sandbox.py is an interface with two implementations:

HarnessProviderExecution and snapshots from the TrueForge sandbox provider. Children inherit installed deps and warm caches — this makes forking cheap.
WorktreeProviderThe offline fallback: one git worktree per node off a prebuilt base, all attempts sharing a pre-warmed virtualenv. No snapshots, same search, same verifier.

Decide between them with ratchet bench-snapshot: under ~5s fork round-trip, run the search on snapshots; over it, take the worktree fallback.

Write a task

A task is a YAML file that tells the gauntlet what to grade. Use tasks/demo-001-slugify/task.yaml as the template. A task names:

  • the repo and the goal,
  • the visible tests the agent may run itself,
  • the held-out tests that actually grade it (never shown to the agent),
  • protected paths that revert before grading.

Verify your task with no agent in the loop first:

ratchet verify --task tasks/<your-task>/task.yaml --repo <repo> --diff <patch>

Add a verification rule

Every new rule ships as a set — this is the repo's contract, not a suggestion:

  1. The rule itself, in the verifier's cheat module (pure function: data in, findings out).
  2. A test with a patch that trips it.
  3. A test with a patch that must not trip it (false-positive guard).
  4. An entry in the red-team battery, so make redteam proves it end to end.
If an honest patch starts getting flagged after your rule lands, your rule is the bug. The battery's two honest patches exist to catch exactly that.

Use the console

The TUI shows the live search tree, scores, prunes, and the approval gate:

make console      # the TUI
make fixture      # a recorded run, so the console works with no model

Audit a run

make audit        # verify the latest run's receipt chain

Reports receipt count, chain head, and intact/broken status. Broken output names the first receipt whose signature or back-link fails. See receipts for what's being checked.

CLI commands

ratchet verify --task <yaml> --repo <dir> --diff <patch>Run the gauntlet on one patch, no agent, no model. Exit code is the verdict.
ratchet run --repo <dir> [--scripted <json>]A full search run. --scripted replays recorded agent moves — fully offline.
ratchet bench-snapshotTime a sandbox fork round-trip; decides HarnessProvider vs WorktreeProvider.
ratchet redteam --repo <dir>Fire the attack battery at the verifier and print the score table.

Make targets

make deveditable install + dev deps
make demoseed demo-repo/ with the broken slugify and three patches
make testthe whole suite; no docker, no network
make lintruff + mypy
make redteamscore the verifier against known cheating patterns
make evalslinear loop vs tree search on seeded bugs
make benchtime a sandbox fork round trip
make run-offlinea complete scripted search, offline
make fixturea recorded run for the console
make consolethe TUI
make auditverify the latest run's receipt chain
make proofexercise every public claim offline, keep the evidence

macOS: timeout error in make evals

The eval script uses GNU timeout, which macOS doesn't ship.

brew install coreutils
ln -s "$(command -v gtimeout)" .venv/bin/timeout   # any PATH location works

An honest patch gets flagged as a cheat

Run the gauntlet on the diff by hand and read the failing stage — it names the exact rule and location:

ratchet verify --task <task> --repo <repo> --diff <patch>

If you still think the rule is wrong, open an issue with the diff attached. The verifier only changes with a reproducing test — never to make one specific patch pass. That rule applies to humans too.

The search never goes green

  • Scores below 1.0 name the stage holding the line — f2p means the hidden tests still fail.
  • Check the budget line: a run that exhausts its node budget parks its best branch rather than forcing a verdict.
  • In an objective graph, an exhausted node escalates to tree search automatically — check the graph status output.

Recover a pruned branch

Pruned work is parked before it's dropped:

git for-each-ref refs/ratchet/pruned/     # list parked nodes
git checkout refs/ratchet/pruned/<node>   # inspect one

A dead end is still a node you can rewind to.

Common questions

What does Ratchet cost?
Nothing — MIT-licensed open source, every feature in the public repo. You pay only your own model providers on live runs, capped by per-run budgets.
What models does it work with?
Model routing belongs to the TrueForge harness; Ratchet is model-agnostic and the whole verifier runs with no model at all.
Does it need Docker or network access?
No. All offline paths (make test/redteam/proof/run-offline) use no docker and no network. Live sandbox egress is allowlisted to PyPI and GitHub.
Can the agent see the hidden tests?
No — names and details are scrubbed from everything the agent reads, with tests asserting it. See the cheat detector.
Can a run's results be forged?
Edit any verdict and make audit reports CHAIN BROKEN with the failing receipt. See receipts.
Where do I get help?
GitHub issues — that's the support channel, and it's read.