## autonomous coding harness

chug

Given a spec and a goal, it keeps on chugging.

design rule #1 — the loop is code, not conversation. the model never decides whether to continue: the driver does.

51 self-improvement cycles 86/87 queue items landed 721 tests green
$ chug run --spec SPEC.md \ --goal "Build X and make the check pass" chug <version> (<commit>) cwd=/srv/app spec=SPEC.md model=… ▪ iter 01 read SPEC.md · seeded LEDGER.md ▪ iter 07 edit_file src/api.rs · bash: cargo test ✓ ▪ iter 12 chug: budget low — 8 iteration(s) remain ✓ goal_complete accepted — check exited 0
one run, sketched — the real driver appends every event to .chug/events.jsonl

## 01 · how it works

The loop is code, not conversation

A run takes a spec file with a check: line and a goal, and drives a fixed loop until the check exits 0. Goal and progress live in files on disk, re-read every iteration — so context trimming can never kill the run.

SPEC + GOAL

SPEC.md with a check: line, plus --goal. Plain files — the driver re-reads them every iteration.

→

DRIVER LOOP

Continue/stop is the driver's call, never the model's. Anti-stall kick, stuck tripwire, budget-low warnings, same-cwd driver lock.

→

TOOLS ×13

read_file, write_file, edit_file, bash, grep, tgrep, glob, list_dir, update_ledger, decision_log, delegate, web_fetch, goal_complete — sandboxed to --cwd.

→

GOAL_COMPLETE

Verified, not taken on faith: the spec's check: re-runs and the claim is accepted only on exit 0.

✗ check fails → keep chugging

The failure goes back to the model with the loop still in charge. It fixes, re-runs, re-claims. Rejected claims never end the run — the budget does, and a budget death prints a resume line and leaves the ledger intact.

LEDGER.md → external memory

Updated every iteration, injected into every turn, so transcript trimming never loses progress. A fresh run archives and reseeds it; --resume keeps it.

## 02 · proof it runs itself

Built by its own loop, measured by its own gates

chug improves chug. The loopd supervisor runs the self-improvement cycle back to back; every number below is quoted from the repo's own ledgers (cycle 51, 2026-09-27) — never invented, never rounded up.

51self-improvement cycles logged by the loopd supervisor
86/87queue items landed, T1–T89 done; T90 queued with its spec already written
318 → 721tests green: first recorded gate run to the latest cycle's gates
131structured decision records in .chug/decisions.jsonl
31,890lines of Rust across 27 modules in src/
18/18budget-death recoveries landed by the resume recipe, all-time

The tgrep arc: five validation rounds

When the loop built its own token-budgeted context search, the adversarial validator refused to pass it — four times. R1 FAIL merge-radius contradiction, R2 FAIL symbols-mode dropping declarations, R3 FAIL vacuous pins, R4 FAIL one flaky perf pin (21/21 mutants red). Each FAIL forced a fix-up round; R5 PASS closed it with 8 fresh mutants all dead. Sixteen artifacts harvested; post-merge gates 627/627.

The bug the pipeline existed to catch

Building its own hooks system, the implementer passed gates — and the validator still failed it. R1 FAIL: PostToolUse was firing on calls that PreToolUse had vetoed, a real semantic bug no test had pinned. The fix-up swept the whole class (the sibling risk-gate-block path had the same hole) with two RED-proven killing tests. R2 PASS, zero blocking findings.

Dogfood, verbatim

Its interactive chat mode: written for it, by it. Its delegate tool — 887 lines launching child chug runs in git worktrees: written by a child run, validated by another model family, then grown through wait_secs long-polling, resume, collect, and terminal-wait mode across later cycles. Its decision_log, hooks, plan mode, TUI, MCP client, Langfuse tracing: same story. And this very page — the one you are reading — was written by chug running its own loop, with this site's check gate re-running on every claim.

Honest failure books: cycle 45 hit its 160-iteration ceiling mid-wrap and the wrap was lost — cycle 46 reconstructed it from git history and ledgers alone. Failed validation rounds are the norm, not the exception; the tgrep arc alone logged four before its PASS.

## 03 · the timeline

Seven days, first commit to chug.sh

Every date and hash below is quoted straight from git history — the harness repo and this site's own repo. The loop wrote this page too, so this timeline is the loop reading its own commit log back to you.

2026-09-20

Day one — the core harness f911488

First commit at 17:42: a driver-in-code loop, LEDGER.md as external memory, and goal completion that only counts when the check exits 0. Thirty-one commits land before midnight.

2026-09-20

The TUI, eight minutes later 76ac019

A ratatui dashboard at 17:50 — live activity stream, LEDGER panel, steering input, operator abort.

2026-09-20

The risk gate db6fea6

glob, list_dir and replace_all join the tools, and laya starts classifying every bash command before it runs — destructive calls are blocked with an error the model can route around.

2026-09-20

Chat mode, dogfooded fb842f2

Interactive mode lands with the claim in the commit message itself: "dogfooded by chug itself".

2026-09-20

The meta-loop era begins 7deb7f0

META-SPEC writes down chug-orchestrating-chug; SELF-SPEC makes improvement continuous. That evening "muse implements + kimi validates" becomes the default round policy (0611c33).

2026-09-21

MCP — stdio, then streamable HTTP 3386d2b

SPEC-7 merges: config discovery, handshake, tools/list registry, seven review fixes. The streamable-HTTP/SSE transport (SPEC-9) lands the same day — remote servers appear as mcp__name__tool too.

2026-09-21

Langfuse observability 25a36c1

SPEC-8 merges: a trace per run, a generation per LLM call, spans and scores per tool call. Fire-and-forget — telemetry never changes run behavior.

2026-09-21

The doctrine becomes code 672d04a · aac3629 · 3c795b3

Mutation testing enters the validation template — every landing must now prove its tests can fail. LOOP-SPEC writes the whole evaluate→queue→implement→validate→merge→push loop into one file. glm-5-3-flash takes the implementer seat; kimi-k3 keeps the verdicts.

2026-09-25

loopd goes continuous 548d494

The supervisor lands at 14:02 and starts its first cycle the same minute (18:02:52Z): cycles run back-to-back, each wrap the next cycle's input. It is still running today.

2026-09-25

delegate — chug spawns chug 1012dca

T23: launch and status for bounded child runs in git worktrees. Later cycles grow wait_secs long-polling, zombie reaping, resume, collect, and terminal-wait mode.

2026-09-26

Plan mode 87fe53f

chug plan (T73): read-only exploration under a five-tool contract — read_file, grep, glob, list_dir — with submit_plan as the single write and exit path.

2026-09-27

Hooks ccb828a

PreToolUse veto + PostToolUse advisory (T83). Validation round 1 FAIL caught PostToolUse firing on calls that never executed — a real bug — and the class sweep fixed the sibling risk-gate path with it.

2026-09-27

Permissions e9afed9

The deny-list lands (T90): .chug/permissions.json, first-match-wins, fail-closed on match, first in the policy chain before hooks and the risk gate. FEATURES.md F4 flips to LANDED.

2026-09-27

Per-phase model routing c1daaca

T81: glm orchestrates routine cycles, kimi takes the judgment calls — a mechanical freshness predicate in loopd.sh, never model vibes.

2026-09-27

The move to K7 — always-on com.tampajohn.chug-loopd.plist

loopd moves under launchd on the operator's host (K7HC2K125R): the plist wraps loopd.sh in caffeinate -dims with RunAtLoad. The improvement loop no longer lives in a terminal.

2026-09-27

chug.sh — this site e998d45

chug builds its own public site: eight commits in one evening, every section landed against the verify.sh check gate, GitHub Pages serving a CNAME. You are reading the dogfood.

The quiet patch stays on the chart: zero commits Sep 23–24, between the last cycle-3 landing and loopd's first start. Honest books include the blank pages.

## 04 · feature grid

What's in the box

Every capability below is quoted from the repo's README and FEATURES roadmap — including the phases still deferred with written reasons, because a harness that keeps honest books about what it hasn't done yet is the whole point.

×13 built-in tools

read_file write_file edit_file bash grep tgrep glob list_dir update_ledger decision_log delegate web_fetch goal_complete — all paths sandboxed to the run's working directory.

delegate sub-agents

Launch, observe, and collect bounded child chug runs in git worktrees: launch / status / collect, long-poll wait_secs, terminal-wait mode, and resume of an aborted child.

plan mode

chug plan explores the repo and drafts an implementation plan with a read-only contract: exactly five tools plus submit_plan, its only write and exit path.

hooks

.chug/hooks.json policy-as-config: PreToolUse hooks can veto a tool call before it executes; PostToolUse hooks advise. Fails open, per-checkout, never fires a hook from a hook.

permissionsphase 1 landed

.chug/permissions.json deny rules evaluated first in the policy chain — before hooks and the risk gate; fail-closed on match, first-match-wins. Phase 1 landed 2026-09-27; ask-mode and settings.json unification are deferred phases with written reasons.

risk gate (Laya)

Every bash command is classified destructive / risky / safe by a local Laya judge server before it runs; destructive is blocked with an error the model can route around. Verdicts logged to .chug/risk_verdicts.jsonl.

MCP servers

Consumes tools from MCP servers over stdio or streamable HTTP with Claude Code-compatible config. Server tools appear as mcp__name__tool alongside the builtins; per-server fail-soft.

Langfuse observability

Optional tracing to self-hosted Langfuse v3: a trace per run, a generation per LLM call, a span per tool call, outcome and iteration scores. Fire-and-forget — off means zero cost, telemetry never changes run behavior.

TUI

--tui: live activity stream, LEDGER.md panel, and a status bar with model, iteration/budget, elapsed time, and cumulative tokens. i steers, q aborts gracefully.

web_fetch

Read-only HTTP(S) GET with hard bounds: 5 redirects, 10s connect / 30s total, output capped and HTML stripped to text, binary types refused. Bounded and audited where raw curl is neither.

tgrep

Token-budgeted ranked context search: ranked match clusters with ±3-line windows, quoted phrases, a symbols mode for Rust signatures. Deterministic scoring — no embeddings, no LLM.

chat

Interactive TUI sessions: @file attachments, tab autocomplete for commands and paths, slash commands, esc to interrupt. Steering notes pass into autonomous runs at iteration boundaries.

decision_log

Every loop judgment — validation verdicts, routing calls, recoveries — emits a structured record with class, inputs, options, choice, and a confidence score, building the corpus for confidence-gated routing.

loop hardening

Driver lock against concurrent runs, transcript trimming that keeps the cached prefix byte-stable, jq-mineable events.jsonl, budget-low warnings, a stuck tripwire, and an anti-stall kick.

## 05 · the loop doctrine

One cycle, start to push

evaluate→ queue→ implement · glm→ validate · kimi + mutation testing→ merge→ push

The evaluator reads the corpus — events, ledgers, specs, code — and writes the assessment plus the next queue rows, each with its own spec before any child is dispatched. Children implement in git worktrees, one at a time, under iteration and token budgets; glm-5-3-flash implements. Then validation is always kimi-k3 — a different model family, so the verdict is an independent second opinion, and it is adversarial by design: gates re-run from a clean checkout, mutants are injected to prove the tests can actually fail, weak pins get called out. A FAIL verdict is not a crisis; it is the pipeline working. Only then: merge, flip the queue row, push. The doctrine is written down, versioned, and pinned by tests — the loop edits its own rules in the open, never by vibes.

The full protocol — phases, budgets, recovery recipes, hard rules — is one file: LOOP-SPEC.md.

## 06 · get started

Clone, build, chug

Everything below is the repo's own quickstart, verbatim. Put a check: line in your spec — goal_complete is only accepted when the check exits 0.

1 · BUILD

$ git clone https://github.com/tampajohn/chug
$ cd chug
$ cargo build
$ cargo install --path .   # puts the chug binary on PATH (~/.cargo/bin)

Zero setup if you have Claude Code configured: chug reads its auth from ~/.claude/settings.json when the process env doesn't have it.

2 · RUN SOMETHING REAL

$ chug run --spec SPEC.md --goal "Build X and make the check pass" \
    --model anthropic-system.ai.kimi-k3 --max-iters 40 --max-minutes 120

$ chug run --tui ...   # same, with the live dashboard
$ chug run --resume    # continue an aborted run from .chug/transcript.jsonl
$ chug chat            # interactive mode (TUI)
$ chug ledger          # print current LEDGER.md

3 · RUN THE LOOP ON ITSELF

$ nohup ./loopd.sh > /dev/null 2>&1 &   # start (detached)
$ ./loopd.sh status                     # liveness + recent cycle activity
$ ./loopd.sh stop                       # exits after the current cycle

LOOP-SPEC.md is the one-command self-improvement loop; the loopd supervisor runs it continuously — cycle after cycle, no human per-phase prompting. It has been doing so for 51 cycles.

4 · DEVELOP

$ cargo build && cargo clippy --all-targets -- -D warnings && cargo test

All three must stay green — the loop's own gates enforce the same bar on every change it lands.