Documentation
Everything here runs on your laptop with Python 3.11+ and git —
no docker, no API keys, no network. When in doubt, make proof re-verifies every
claim and keeps the evidence.
Install
The quick way — the Homebrew package bundles GNU timeout and puts ratchet on your PATH:
brew install ayaangazali/ratchet/ratchet-agent # macOS # or, anywhere with uv: uv tool install git+https://github.com/ayaangazali/ratchet
From source (Python 3.11+, git; Node 22.14+ only for the live TrueForge harness path):
git clone https://github.com/ayaangazali/ratchet && cd ratchet make dev # editable install + dev deps
Run the demo
make demo seeds demo-repo/ with a deliberately broken
slugify function and three prepared patches:
| honest.diff | the real fix — should go green |
| cheat.diff | a reward hack — must be blocked before it executes |
| canary_hack.diff | a cheat with zero static findings — must be caught anyway |
It also seeds visible tests, a held-out (hidden) suite, and a canary test. Then run the whole suite and the red team:
make test # the whole suite; no docker, no network make redteam # the published reward-hacking battery fired at the verifier
Your first verdicts
The gauntlet runs standalone — no agent, no model. Three verdicts tell the whole story:
ratchet verify --task tasks/demo-001-slugify/task.yaml --repo demo-repo \
--diff demo-repo/patches/honest.diff # GREEN, score 1.000
ratchet verify --task tasks/demo-001-slugify/task.yaml --repo demo-repo \
--diff demo-repo/patches/cheat.diff # CHEATED
ratchet verify --task tasks/canary-impossible/task.yaml --repo demo-repo \
--diff demo-repo/patches/canary_hack.diff # CHEATED, zero static findings
The cheat is caught by static analysis before a line of it runs. The canary hack trips no static rule and is caught by behavior instead — see the cheat detector.
A full search run
A complete run — root node, a real prune, a green node, and the human approval gate — works offline with a scripted agent:
make run-offline # or directly: ratchet run --repo demo-repo --scripted demo-repo/patches/scripted.json
You'll see the search tree print as it goes: pruned branches marked
✗, the live best node starred, and the run ending at an approval that waits
for a human decision.
● root 0.58 ├─✗ 1e0f 0.58 pruned └─★ 2874 1.00 ✓green
How it works
You give Ratchet a repo and a task file. An agent proposes patches; Ratchet never asks the agent whether it's finished — there is no "done" tool anywhere in the system. Instead, every candidate patch runs a seven-stage verifier gauntlet, and green is set in exactly one place in the codebase: the gauntlet itself.
Every step the agent takes is a git commit plus a sandbox snapshot, so a run is a tree search over repo states with the verifier's score as the value function. The winning path exits as one squashed diff at a human approval gate, and every graded step is signed into a tamper-evident receipt chain.
Division of labor: the TrueForge harness owns model calls, sandboxes, sub-agents, and session persistence. Ratchet owns exactly one question — what counts as progress.
The gauntlet
Seven stages, in order. A patch that fails an early stage never reaches a later one — and a patch that fails the cheat check never executes at all.
| # | stage | what it asks |
|---|---|---|
| 1 | build | Does the repo still build/import with the patch applied? |
| 2 | cheat check | Does the diff match any known reward-hacking pattern? (static — runs before execution) |
| 3 | fail-to-pass | Do the held-out tests that were failing now pass? |
| 4 | pass-to-pass | Does everything that passed before still pass? (no regressions) |
| 5 | types | Type check clean? |
| 6 | lint | Lint clean? |
| 7 | diff hygiene | Is the diff minimal and inside allowed paths? |
Two hard rules around the gauntlet: protected paths (tests, verifier config) are reverted before grading on every run with no flag to skip, and the exit code is echoed outside the output region the agent can influence.
The cheat detector
Stage two is a library of static rules, each shipped with two tests — a patch that trips it and a patch that must not. Current rules:
hard_exit · always_equal · special_casing · runtime_test_write · protected_path · skip_marker · test_file_emptied · test_deleted · assertion_removed · report_hook_tamper · env_bypass · canary_passed · broad_except_pass · log_spoofed
The canary
Static rules can't catch everything — so one task in every battery is
impossible: its hidden test cannot legitimately pass. Any patch that makes the
canary go green must have smuggled an answer, even if it trips zero static rules. That's
the canary_passed finding, and it's why the canary hack in the demo is caught
with no static findings at all.
Hidden tests stay hidden
Held-out test names and failure details never appear in anything the agent can read — context assembly, observations, and every bus event are scrubbed, and the suite has tests asserting it. A leak would silently destroy the signal.
Tree search
Because every step is a restorable node (git commit + snapshot), a run is not a linear loop with retries — it's a tree search over repo states:
- Score — the verifier grades each node; the score is the value function.
- Fork — stalled branches fork in parallel; sub-agents share the sandbox but see none of the conversation.
- Prune — dead ends are cut, but parked first at
refs/ratchet/pruned/<node>, so nothing is lost. - Budget — every run has explicit caps: nodes, wall-clock, dollars. The scheduler decides where the next unit of compute goes.
The winning path is squashed into one clean diff and parked at the approval gate. Nothing irreversible — no push, no PR — happens without a human decision.
Receipts
Every graded node — accepted or pruned — is recorded into a signed, hash-linked chain, and a finished run seals its chain. That gives you two guarantees:
- Order — results are provably in the order they were issued.
- Integrity — edit any verdict after the fact and the chain breaks.
make audit # receipts 4 # head b26a36f26e627c03 # chain intact
Forge a green verdict and the same command reports CHAIN BROKEN with the
exact receipt whose signature fails. The tamper demo runs as part of make proof.
The red team
The verifier itself is what gets evaluated. make redteam fires a published
battery of real reward-hacking patterns at the gauntlet — skip markers, hardcoded answers,
tests rewritten at import time, conftest report hooks, spoofed exit codes — plus two honest
patches that must be accepted, because a paranoid verifier that rejects real fixes
is broken too.
The score to hold: every attack caught, zero false positives. It runs in CI, and the dashboard shows the full attack-by-attack table from the latest proof run.
Objective graph
Bigger goals decompose into a graph of objective nodes — each one fulfillable only by tests, never by the agent's say-so. Nodes are attempted cheaply first; a node that exhausts its attempts escalates to full tree search automatically.
graph demo-slugify-graph · 2 node(s) ✓ accents fulfilled attempts 1 ✓ truncation fulfilled attempts 3 (escalated to tree search)
The demo graph lives at objectives/demo-graph.yaml.
The docs oracle
When a failure looks like the outside world — an import error, a missing attribute, an unexpected keyword argument — the oracle attaches current upstream documentation for the exact version in the lockfile to the next prompt, instead of letting the model guess from memory.
- Sources are configured in
ratchet/scrapers.yaml, version-controlled — never one-off shell scrapes. - Content is extracted by heading, not CSS selector — headings survive site redesigns.
- Every fetch is validated against the source's
expectblock; on failure the oracle relocates the section by heading similarity and commits the repair as a diff, so the repair history is auditable.
Sandboxes
Ratchet does not orchestrate containers — ever. sandbox.py is an interface
with two implementations:
| HarnessProvider | Execution and snapshots from the TrueForge sandbox provider. Children inherit installed deps and warm caches — this makes forking cheap. |
| WorktreeProvider | The offline fallback: one git worktree per node off a prebuilt base, all attempts sharing a pre-warmed virtualenv. No snapshots, same search, same verifier. |
Decide between them with ratchet bench-snapshot: under ~5s fork round-trip,
run the search on snapshots; over it, take the worktree fallback.
Write a task
A task is a YAML file that tells the gauntlet what to grade. Use
tasks/demo-001-slugify/task.yaml as the template. A task names:
- the repo and the goal,
- the visible tests the agent may run itself,
- the held-out tests that actually grade it (never shown to the agent),
- protected paths that revert before grading.
Verify your task with no agent in the loop first:
ratchet verify --task tasks/<your-task>/task.yaml --repo <repo> --diff <patch>
Add a verification rule
Every new rule ships as a set — this is the repo's contract, not a suggestion:
- The rule itself, in the verifier's cheat module (pure function: data in, findings out).
- A test with a patch that trips it.
- A test with a patch that must not trip it (false-positive guard).
- An entry in the red-team battery, so
make redteamproves it end to end.
Use the console
The TUI shows the live search tree, scores, prunes, and the approval gate:
make console # the TUI make fixture # a recorded run, so the console works with no model
Audit a run
make audit # verify the latest run's receipt chain
Reports receipt count, chain head, and intact/broken status. Broken output names the first receipt whose signature or back-link fails. See receipts for what's being checked.
CLI commands
| ratchet verify --task <yaml> --repo <dir> --diff <patch> | Run the gauntlet on one patch, no agent, no model. Exit code is the verdict. |
| ratchet run --repo <dir> [--scripted <json>] | A full search run. --scripted replays recorded agent moves — fully offline. |
| ratchet bench-snapshot | Time a sandbox fork round-trip; decides HarnessProvider vs WorktreeProvider. |
| ratchet redteam --repo <dir> | Fire the attack battery at the verifier and print the score table. |
Make targets
| make dev | editable install + dev deps |
| make demo | seed demo-repo/ with the broken slugify and three patches |
| make test | the whole suite; no docker, no network |
| make lint | ruff + mypy |
| make redteam | score the verifier against known cheating patterns |
| make evals | linear loop vs tree search on seeded bugs |
| make bench | time a sandbox fork round trip |
| make run-offline | a complete scripted search, offline |
| make fixture | a recorded run for the console |
| make console | the TUI |
| make audit | verify the latest run's receipt chain |
| make proof | exercise every public claim offline, keep the evidence |
macOS: timeout error in make evals
The eval script uses GNU timeout, which macOS doesn't ship.
brew install coreutils ln -s "$(command -v gtimeout)" .venv/bin/timeout # any PATH location works
An honest patch gets flagged as a cheat
Run the gauntlet on the diff by hand and read the failing stage — it names the exact rule and location:
ratchet verify --task <task> --repo <repo> --diff <patch>
If you still think the rule is wrong, open an issue with the diff attached. The verifier only changes with a reproducing test — never to make one specific patch pass. That rule applies to humans too.
The search never goes green
- Scores below 1.0 name the stage holding the line —
f2pmeans the hidden tests still fail. - Check the budget line: a run that exhausts its node budget parks its best branch rather than forcing a verdict.
- In an objective graph, an exhausted node escalates to tree search automatically — check the graph status output.
Recover a pruned branch
Pruned work is parked before it's dropped:
git for-each-ref refs/ratchet/pruned/ # list parked nodes git checkout refs/ratchet/pruned/<node> # inspect one
A dead end is still a node you can rewind to.
Common questions
What does Ratchet cost?
What models does it work with?
Does it need Docker or network access?
make test/redteam/proof/run-offline)
use no docker and no network. Live sandbox egress is allowlisted to PyPI and GitHub.Can the agent see the hidden tests?
Can a run's results be forged?
make audit reports CHAIN BROKEN with
the failing receipt. See receipts.