## autonomous coding harness
Given a spec and a goal, it keeps on chugging.
design rule #1 — the loop is code, not conversation. the model never decides whether to continue: the driver does.
## 01 · how it works
A run takes a spec file with a check: line and a goal, and drives a fixed loop until the check exits 0. Goal and progress live in files on disk, re-read every iteration — so context trimming can never kill the run.
SPEC.md with a check: line, plus --goal. Plain files — the driver re-reads them every iteration.
Continue/stop is the driver's call, never the model's. Anti-stall kick, stuck tripwire, budget-low warnings, same-cwd driver lock.
read_file, write_file, edit_file, bash, grep, tgrep, glob, list_dir, update_ledger, decision_log, delegate, web_fetch, goal_complete — sandboxed to --cwd.
Verified, not taken on faith: the spec's check: re-runs and the claim is accepted only on exit 0.
The failure goes back to the model with the loop still in charge. It fixes, re-runs, re-claims. Rejected claims never end the run — the budget does, and a budget death prints a resume line and leaves the ledger intact.
Updated every iteration, injected into every turn, so transcript trimming never loses progress. A fresh run archives and reseeds it; --resume keeps it.
## 02 · proof it runs itself
chug improves chug. The loopd supervisor runs the self-improvement cycle back to back; every number below is quoted from the repo's own ledgers (cycle 51, 2026-09-27) — never invented, never rounded up.
When the loop built its own token-budgeted context search, the adversarial validator refused to pass it — four times. R1 FAIL merge-radius contradiction, R2 FAIL symbols-mode dropping declarations, R3 FAIL vacuous pins, R4 FAIL one flaky perf pin (21/21 mutants red). Each FAIL forced a fix-up round; R5 PASS closed it with 8 fresh mutants all dead. Sixteen artifacts harvested; post-merge gates 627/627.
Building its own hooks system, the implementer passed gates — and the validator still failed it. R1 FAIL: PostToolUse was firing on calls that PreToolUse had vetoed, a real semantic bug no test had pinned. The fix-up swept the whole class (the sibling risk-gate-block path had the same hole) with two RED-proven killing tests. R2 PASS, zero blocking findings.
Its interactive chat mode: written for it, by it. Its delegate tool — 887 lines launching child chug runs in git worktrees: written by a child run, validated by another model family, then grown through wait_secs long-polling, resume, collect, and terminal-wait mode across later cycles. Its decision_log, hooks, plan mode, TUI, MCP client, Langfuse tracing: same story. And this very page — the one you are reading — was written by chug running its own loop, with this site's check gate re-running on every claim.
Honest failure books: cycle 45 hit its 160-iteration ceiling mid-wrap and the wrap was lost — cycle 46 reconstructed it from git history and ledgers alone. Failed validation rounds are the norm, not the exception; the tgrep arc alone logged four before its PASS.
## 03 · the timeline
Every date and hash below is quoted straight from git history — the harness repo and this site's own repo. The loop wrote this page too, so this timeline is the loop reading its own commit log back to you.
First commit at 17:42: a driver-in-code loop, LEDGER.md as external memory, and goal completion that only counts when the check exits 0. Thirty-one commits land before midnight.
A ratatui dashboard at 17:50 — live activity stream, LEDGER panel, steering input, operator abort.
glob, list_dir and replace_all join the tools, and laya starts classifying every bash command before it runs — destructive calls are blocked with an error the model can route around.
Interactive mode lands with the claim in the commit message itself: "dogfooded by chug itself".
META-SPEC writes down chug-orchestrating-chug; SELF-SPEC makes improvement continuous. That evening "muse implements + kimi validates" becomes the default round policy (0611c33).
SPEC-7 merges: config discovery, handshake, tools/list registry, seven review fixes. The streamable-HTTP/SSE transport (SPEC-9) lands the same day — remote servers appear as mcp__name__tool too.
SPEC-8 merges: a trace per run, a generation per LLM call, spans and scores per tool call. Fire-and-forget — telemetry never changes run behavior.
Mutation testing enters the validation template — every landing must now prove its tests can fail. LOOP-SPEC writes the whole evaluate→queue→implement→validate→merge→push loop into one file. glm-5-3-flash takes the implementer seat; kimi-k3 keeps the verdicts.
The supervisor lands at 14:02 and starts its first cycle the same minute (18:02:52Z): cycles run back-to-back, each wrap the next cycle's input. It is still running today.
T23: launch and status for bounded child runs in git worktrees. Later cycles grow wait_secs long-polling, zombie reaping, resume, collect, and terminal-wait mode.
chug plan (T73): read-only exploration under a five-tool contract — read_file, grep, glob, list_dir — with submit_plan as the single write and exit path.
PreToolUse veto + PostToolUse advisory (T83). Validation round 1 FAIL caught PostToolUse firing on calls that never executed — a real bug — and the class sweep fixed the sibling risk-gate path with it.
The deny-list lands (T90): .chug/permissions.json, first-match-wins, fail-closed on match, first in the policy chain before hooks and the risk gate. FEATURES.md F4 flips to LANDED.
T81: glm orchestrates routine cycles, kimi takes the judgment calls — a mechanical freshness predicate in loopd.sh, never model vibes.
loopd moves under launchd on the operator's host (K7HC2K125R): the plist wraps loopd.sh in caffeinate -dims with RunAtLoad. The improvement loop no longer lives in a terminal.
chug builds its own public site: eight commits in one evening, every section landed against the verify.sh check gate, GitHub Pages serving a CNAME. You are reading the dogfood.
The quiet patch stays on the chart: zero commits Sep 23–24, between the last cycle-3 landing and loopd's first start. Honest books include the blank pages.
## 04 · feature grid
Every capability below is quoted from the repo's README and FEATURES roadmap — including the phases still deferred with written reasons, because a harness that keeps honest books about what it hasn't done yet is the whole point.
read_file write_file edit_file bash grep tgrep glob list_dir update_ledger decision_log delegate web_fetch goal_complete — all paths sandboxed to the run's working directory.
Launch, observe, and collect bounded child chug runs in git worktrees: launch / status / collect, long-poll wait_secs, terminal-wait mode, and resume of an aborted child.
chug plan explores the repo and drafts an implementation plan with a read-only contract: exactly five tools plus submit_plan, its only write and exit path.
.chug/hooks.json policy-as-config: PreToolUse hooks can veto a tool call before it executes; PostToolUse hooks advise. Fails open, per-checkout, never fires a hook from a hook.
.chug/permissions.json deny rules evaluated first in the policy chain — before hooks and the risk gate; fail-closed on match, first-match-wins. Phase 1 landed 2026-09-27; ask-mode and settings.json unification are deferred phases with written reasons.
Every bash command is classified destructive / risky / safe by a local Laya judge server before it runs; destructive is blocked with an error the model can route around. Verdicts logged to .chug/risk_verdicts.jsonl.
Consumes tools from MCP servers over stdio or streamable HTTP with Claude Code-compatible config. Server tools appear as mcp__name__tool alongside the builtins; per-server fail-soft.
Optional tracing to self-hosted Langfuse v3: a trace per run, a generation per LLM call, a span per tool call, outcome and iteration scores. Fire-and-forget — off means zero cost, telemetry never changes run behavior.
--tui: live activity stream, LEDGER.md panel, and a status bar with model, iteration/budget, elapsed time, and cumulative tokens. i steers, q aborts gracefully.
Read-only HTTP(S) GET with hard bounds: 5 redirects, 10s connect / 30s total, output capped and HTML stripped to text, binary types refused. Bounded and audited where raw curl is neither.
Token-budgeted ranked context search: ranked match clusters with ±3-line windows, quoted phrases, a symbols mode for Rust signatures. Deterministic scoring — no embeddings, no LLM.
Interactive TUI sessions: @file attachments, tab autocomplete for commands and paths, slash commands, esc to interrupt. Steering notes pass into autonomous runs at iteration boundaries.
Every loop judgment — validation verdicts, routing calls, recoveries — emits a structured record with class, inputs, options, choice, and a confidence score, building the corpus for confidence-gated routing.
Driver lock against concurrent runs, transcript trimming that keeps the cached prefix byte-stable, jq-mineable events.jsonl, budget-low warnings, a stuck tripwire, and an anti-stall kick.
## 05 · the loop doctrine
The evaluator reads the corpus — events, ledgers, specs, code — and writes the assessment plus the next queue rows, each with its own spec before any child is dispatched. Children implement in git worktrees, one at a time, under iteration and token budgets; glm-5-3-flash implements. Then validation is always kimi-k3 — a different model family, so the verdict is an independent second opinion, and it is adversarial by design: gates re-run from a clean checkout, mutants are injected to prove the tests can actually fail, weak pins get called out. A FAIL verdict is not a crisis; it is the pipeline working. Only then: merge, flip the queue row, push. The doctrine is written down, versioned, and pinned by tests — the loop edits its own rules in the open, never by vibes.
The full protocol — phases, budgets, recovery recipes, hard rules — is one file: LOOP-SPEC.md.
## 06 · get started
Everything below is the repo's own quickstart, verbatim. Put a check: line in your spec — goal_complete is only accepted when the check exits 0.
$ git clone https://github.com/tampajohn/chug $ cd chug $ cargo build $ cargo install --path . # puts the chug binary on PATH (~/.cargo/bin)
Zero setup if you have Claude Code configured: chug reads its auth from ~/.claude/settings.json when the process env doesn't have it.
$ chug run --spec SPEC.md --goal "Build X and make the check pass" \ --model anthropic-system.ai.kimi-k3 --max-iters 40 --max-minutes 120 $ chug run --tui ... # same, with the live dashboard $ chug run --resume # continue an aborted run from .chug/transcript.jsonl $ chug chat # interactive mode (TUI) $ chug ledger # print current LEDGER.md
$ nohup ./loopd.sh > /dev/null 2>&1 & # start (detached) $ ./loopd.sh status # liveness + recent cycle activity $ ./loopd.sh stop # exits after the current cycle
LOOP-SPEC.md is the one-command self-improvement loop; the loopd supervisor runs it continuously — cycle after cycle, no human per-phase prompting. It has been doing so for 51 cycles.
$ cargo build && cargo clippy --all-targets -- -D warnings && cargo test
All three must stay green — the loop's own gates enforce the same bar on every change it lands.