- Kind
- Own open source · MIT
- Year
- 2026
- Role
- Author
- Stack
- Python 3.10+Claude Code hooks & skillsmypy --strictGitHub Actions
- Source
- r0h1tb/plumb ↗
Plumb
Evidence-first workflows for Claude Code. A coding agent has to name what proves the work is done, collect that evidence, and try to break its own result before it may say "done".
The hard partAsk an agent whether the feature works and it will say yes, sincerely, having verified nothing. "Done" was never defined, so the only available answer is a feeling.
- Zero runtime dependencies, no network, no telemetry
- Uninstall removes only files whose hash it recorded
- Never changes the user's permission mode
The problem
Coding agents are good at producing work and bad at knowing whether it is finished. "The tests pass" and "I implemented it" are inputs to a verdict, not the verdict. Without a definition of done, an agent stops when it runs out of ideas, and reports success.
The loop
Plumb turns "done" into something checkable, in five steps the agent cannot skip.
Name an oracle. The one observable that, if true, means the goal is met: "p95 under 200 ms on the real index", not "search is fast". No oracle, no run.
Cite evidence. Each run keeps a ledger of commands and results with their conditions. A verdict that cites nothing is a failed criterion. "p95 = 41 ms" is not a measurement; "p95 = 41 ms, warm cache, 500 requests, 20 concurrent" is.
Name the misfire. Before starting: how could this pass every criterion and still be wrong? Writing it down early is what stops the agent walking into it.
Try to break it. One honest attempt to make each criterion fail — the cold path, the empty case, the unplanned input.
Say what was not checked. A report that lists only successes is a pitch.
Context is a budget
Every line that loads into an agent's context on every turn costs tokens and attention on every turn. Two design decisions follow from that.
Rules are indexed, not inlined. Coding standards live on disk.
CLAUDE.md gets a short table naming each rule file and the situation
in which it is worth reading. The agent loads a rule when it is about to
edit a file it applies to, instead of carrying every language's
standards in every turn.
Skill descriptions have a budget the build enforces. A skill's description loads every turn, so it is compiled from a JSON manifest and validated: over 1,024 characters fails the build. A step file that exists but is never referenced is flagged too, because "wrote it, forgot to register it" silently produces a skill that never runs that step.
Not touching what is not yours
Agent tooling installs into a directory people have already customised, and most of it gets that wrong. Each of these is enforced by a test:
- The settings merge is additive. A key the user already set is reported, never overwritten.
- Every installed file is recorded with its SHA-256. Uninstall removes files that still match and keeps the ones the user edited. Content is canonicalised before hashing so Plumb's own files always re-hash to the recorded digest.
- Plumb's index in
CLAUDE.mdsits between markers; everything outside them is preserved byte for byte. - No runtime dependencies, so there is no supply chain to audit, and it never sets the agent's permission mode.
Outcome
Five workflows (/land, /plan, /fix, /groundwork, /distill), rules loaded on demand, hooks that inform rather than block, and an installer that can prove what it changed and undo exactly that.