Skip to content
Rohit Behera
← All work
Kind
Own open source · MIT
Year
2026
Role
Author
Stack
Python 3.10+Claude Code hooks & skillsmypy --strictGitHub Actions

Plumb

Evidence-first workflows for Claude Code. A coding agent has to name what proves the work is done, collect that evidence, and try to break its own result before it may say "done".

The hard partAsk an agent whether the feature works and it will say yes, sincerely, having verified nothing. "Done" was never defined, so the only available answer is a feeling.

  • Zero runtime dependencies, no network, no telemetry
  • Uninstall removes only files whose hash it recorded
  • Never changes the user's permission mode

The problem

Coding agents are good at producing work and bad at knowing whether it is finished. "The tests pass" and "I implemented it" are inputs to a verdict, not the verdict. Without a definition of done, an agent stops when it runs out of ideas, and reports success.

The loop

Plumb turns "done" into something checkable, in five steps the agent cannot skip.

Plumb evidence loopA run starts by naming an oracle, the observable that proves the goal, and the misfire, how the run could pass and still be wrong. The agent works, records each command and result in an evidence ledger, then tries to falsify each criterion. A broken criterion sends it back to work. A verdict must cite ledger entries, and the report lists what was not verified.1 · Oraclethe observablethat proves done2 · Misfirehow it could passand still be wrong3 · Workbuild, run, measure4 · Evidence ledgerE1 pytest: 214 passedE2 p95 186ms cold5 · Falsifycold path, empty case,unplanned inputbroke it → back to workVerdict + reportcites E-ids · lists what was not verifiedsurvivedhooks inform, never blockrules load on demand from an indexevery installed file recorded with sha256
No oracle, no run. A verdict that cites no ledger entry counts as a failed criterion.

Name an oracle. The one observable that, if true, means the goal is met: "p95 under 200 ms on the real index", not "search is fast". No oracle, no run.

Cite evidence. Each run keeps a ledger of commands and results with their conditions. A verdict that cites nothing is a failed criterion. "p95 = 41 ms" is not a measurement; "p95 = 41 ms, warm cache, 500 requests, 20 concurrent" is.

Name the misfire. Before starting: how could this pass every criterion and still be wrong? Writing it down early is what stops the agent walking into it.

Try to break it. One honest attempt to make each criterion fail — the cold path, the empty case, the unplanned input.

Say what was not checked. A report that lists only successes is a pitch.

Context is a budget

Every line that loads into an agent's context on every turn costs tokens and attention on every turn. Two design decisions follow from that.

Rules are indexed, not inlined. Coding standards live on disk. CLAUDE.md gets a short table naming each rule file and the situation in which it is worth reading. The agent loads a rule when it is about to edit a file it applies to, instead of carrying every language's standards in every turn.

Skill descriptions have a budget the build enforces. A skill's description loads every turn, so it is compiled from a JSON manifest and validated: over 1,024 characters fails the build. A step file that exists but is never referenced is flagged too, because "wrote it, forgot to register it" silently produces a skill that never runs that step.

Not touching what is not yours

Agent tooling installs into a directory people have already customised, and most of it gets that wrong. Each of these is enforced by a test:

  • The settings merge is additive. A key the user already set is reported, never overwritten.
  • Every installed file is recorded with its SHA-256. Uninstall removes files that still match and keeps the ones the user edited. Content is canonicalised before hashing so Plumb's own files always re-hash to the recorded digest.
  • Plumb's index in CLAUDE.md sits between markers; everything outside them is preserved byte for byte.
  • No runtime dependencies, so there is no supply chain to audit, and it never sets the agent's permission mode.

Outcome

Five workflows (/land, /plan, /fix, /groundwork, /distill), rules loaded on demand, hooks that inform rather than block, and an installer that can prove what it changed and undo exactly that.

loading index…Full retrieval trace →