Skip to content

Repository files navigation

Seams

Design at the seams. Build in slices. Ship what you verified.

version license test routing evidence

Seams is a Claude Code plugin (plugin id matt-pocock-workflow) that makes Matt Pocock's engineering skills lead every session, and holds the project closed until they do. A session bootstrap routes each request by size and by risk; a hook refuses any change to the project until a workflow skill has been declared for that request; another refuses to end a turn that changed code without verification. Around his skills sits a senior engineer's process: a grill that asks clickable questions, every independent one at once, a design lens, tests first at agreed seams, a review of the committed candidate, a definition of done with evidence, a handover that names the stage reached, a release that proves the exact candidate is what runs, and an incident route that contains before it diagnoses.

Try it in 60 seconds

curl -fsSL https://raw.githubusercontent.com/gabriel-tutor/seams/main/scripts/install.sh | bash

Restart Claude Code, open any repo, and say one of these:

Say What happens
"check what this repo has and what it's missing" a foundations survey: run and verify commands, lint, hooks, CI, glossary, boundaries, and the production basics (pipeline, environments, backups, monitoring, scanning); gaps reported, fixes offered, nothing written without a yes
"X is broken when Y" diagnosing-bugs: a reproducing loop first, then ranked hypotheses, then a regression test, then the fix
"add " a design grill in clickable questions, every independent one at once, recorded in the feature's progress file as answers land, until nothing is assumed; then implementation with tests first, a commit, a review of it, and a handover
a typo fix the trivial declaration, the edit, and the narrowest check that proves it
"ship it", "deploy to staging" release: a readiness table where anything unmet blocks, a deploy only after a yes that names the candidate, the environment and the target, verification that the exact candidate runs, an operations handover
"production is down" incident: who is affected and what changed last, the safest reversible containing action behind a yes, restore and confirm, and only then the diagnosis
/pr-review 42, /pr-review 42 57, /pr-review requested (you type it, or Claude starts it when asked) a deep review of each pull request: its head and baseline checked out in worktrees of their own, the repo's real checks (e2e included) run on both so every failure is attributed, the change tried, code-review plus a risk reviewer, findings proven, one GitHub review drafted per PR and posted only on your yes; it ends with which PRs are ready to merge and a note for each author whose PR needs work
"just write the file, skip the process" a refusal from the gate that names the file and the routes; one skill invocation opens it

Nothing is installed except the plugin and Matt Pocock's skills; one command removes it (see Turning it off).

The gate

Two hooks turn the routing policy from a promise into a rule (ADR-0001).

A change needs a declaration. Until a process skill has been invoked for the current request, any Edit, Write, MultiEdit or NotebookEdit to a file outside the temp and Claude config directories, and any shell command, run through Bash or a Monitor watch, that the gate's classifier labels a mutation (a redirect to a file, rm, mv, sed -i, git commit, npm install, a formatter's --write, an inline Python program that writes a file, and the rest of a documented list), is refused. A > the shell reads as text is not a redirect, in quotes or inside a quoted $(…) such as x="$(awk '$1>0' f)". PowerShell has no classifier: before a declaration it runs only Get-Content, Get-ChildItem, Select-String and git status, git diff or git log, each on its own with plain arguments. Scratch work is not a change: a shell write whose every path is an absolute path under the temp directory or the session's scratchpad passes, as Edit and Write there always have, while a relative path, a variable set elsewhere, an inline program or a path in the Claude config directory, wherever it lies, still counts. The refusal is the tool result Claude sees:

Seams gate: editing /path/to/orderkit/src/pricing.ts changes the project, and this request has no declaration yet: no process skill has been invoked for it. Route it first, with the Skill tool: diagnosing-bugs for something broken, matt-pocock-workflow:grill for a change to behavior, tdd or matt-pocock-workflow:implement to keep building an agreed design, matt-pocock-workflow:trivial for a change with no effect on behavior, data shape or security. Then retry this call. Scratch work is not a change: a write whose every path is an absolute path under the temp directory or the session's scratchpad needs no declaration, from Edit, Write or a shell command.

A declaration is a Skill invocation of a Seams skill or one of Matt Pocock's process skills (his implement, to-spec, to-tickets, grilling, tdd, diagnosing-bugs, code-review and the rest), by Claude or typed by you as a slash command. A typed skill counts under the name Claude Code expanded it to, which its UserPromptExpansion event reports, so a bare name (/grill, /pr-review 42) and each skill of a stacked command (/grill /tdd …) count; where that event is missing, the prompt's leading command still does. A Superpowers skill or another plugin's is not one, and neither is the routing policy skill itself or an MCP server's prompt. A short go-ahead (yes, continue, option 2) keeps the current declaration, and so does a notice Claude Code generates (a background task finishing, a subagent's hand-back, a system reminder); any other prompt starts a new request that needs its own. A subagent's own declaration covers that subagent alone, and a new request doesn't lapse it: a parallel run's builders keep working while you type, and never open the gate for your message, while the main conversation's declarations still cover its subagents. When the request before it had one and the message types no route of its own, Claude is told which declaration lapsed: invoking that skill again continues the same work (a skill only you can type, such as Matt Pocock's /implement, you type again), and new work takes its own route. pr-review is left out: a review needs no declaration, and invoking it again would re-arm its pre-approved scripts. The hint restores nothing itself. A false positive costs one call: matt-pocock-workflow:trivial carries the test of what is not trivial (no behavior change, no shape change, nothing sensitive, reversible in one commit) and routes up when any part fails.

The read-only agents only read. Seams ships two agents that the skills name when they delegate: scout finds facts in the code and docs (the grill's fact-finding, to-spec's exploring, foundations' survey, implement's reading beyond a few files), and reviewer reviews a named diff (the subagents of implement's reviews). Each returns conclusions with file:line or URL citations and says what it couldn't confirm. Neither has an editor tool, and scout has no shell. Whatever the request has declared, the gate holds both to a list of reads, the opposite of the classifier's mesh of writes: every command in a shell line must be one of git's read subcommands (diff, log, show, blame and the like, with no -c, --output or --ext-diff), gh's views, or a file reader (cat, grep, find without -exec or -delete, sort without -o, and the rest of a short list), named plainly, with no variable or substitution, and redirected only into the temp directory or the session's scratchpad, where an editor tool may also write. So npm test, a formatter, a build or a script the agent wrote is refused as surely as git commit: a finding that needs a check or a probe says so, and the main conversation runs it.

A change needs verification. When a turn changed non-documentation files and matt-pocock-workflow:verification-before-completion did not run afterwards, the turn cannot end: the Stop hook asks once, as hook feedback rather than a hook error, naming how many unverified changes it counted and one of them, and Claude runs the verification with its real output before finishing. A turn that ends with a question to you is delayed by one message, never trapped.

The ledger behind both is one JSON file per session in a per-user directory under the temp directory ($TMPDIR/seams-$(id -u)/<session>.json) holding skill names, tool names, paths and prompt ids only, never command or prompt text; /clear and a new session reset it, compaction and resume keep it. Every hook fails open: a bug in the plugin writes a traceback to claude --debug and lets your work through. Seams reads no setting or environment variable of its own that turns the gate off; disabling the plugin is its off switch. Claude Code's own settings can turn it off from outside, and What switches it off lists them.

The workflow

Every session starts with the routing policy in context. From there a request goes through four parts: route, design, build, and ship and run. Every diamond is a question Claude asks you, one at a time with the recommended answer first, and waits on. Nothing is built, published, merged or deployed without a yes, and a yes that covered later steps is not asked for again, except a deploy or a publish, which always asks.

1. Route: ceremony scales with the change, and with its risk

flowchart LR
    S([Session starts]) --> B[/"Bootstrap injected:<br/>routing policy in context"/]
    B --> R{Route by<br/>size and risk}
    R -. an edit or a shell write before a<br/>declaration: refused, the reason names the routes .-> R
    R -->|typo, copy, comment| T["trivial<br/>declare → edit → narrowest check"] --> V
    R -->|broken, failing, slow| D["diagnosing-bugs<br/>reproduce → hypotheses<br/>→ regression test → fix"] --> V
    R -->|bounded change| G["grill"] --> TDD["tdd at the<br/>agreed seams"] --> V
    R -->|new behavior| G2["grill"] --> I["implement"] --> V
    R -->|several sessions,<br/>or a new app| G3["grill"] --> SPEC["to-spec → to-tickets<br/>→ implement per ticket<br/>(a new app: ticket 01 is the walking skeleton)"] --> V
    R -->|sensitive, any size:<br/>auth, secrets, billing, migrations,<br/>infra, CI, public API, destructive| SEC["its size row's path, with grill first<br/>on the security and failure axes<br/>→ code-review and a security review,<br/>both required"] --> V
    R -->|users affected now| INC["incident<br/>contain → restore → then diagnose"] --> V
    R -->|ship, deploy, publish| REL["release<br/>readiness → deploy on a yes<br/>→ verify → operations handover"] --> V
    R -->|too foggy to see the way| W[/"suggests /wayfinder"/]
    V["verification-before-completion<br/>+ handover naming the stage"] --> E([Turn ends])
    E -. changed code without<br/>verification: blocked once .-> V
    B -.repo not set up.-> FO["foundations<br/>(once per repo)"]
Loading

2. Design: the grill, every independent question at once

flowchart LR
    G["Facts from the code, found by scout agents<br/>+ up to four independent decisions, clickable<br/>(what the code or an earlier answer<br/>settles is a fact, not a question)"] --> Q{frontier<br/>empty?}
    Q -->|no| G
    Q -->|yes| L["Design lens, 10 axes<br/>data · seams · failure modes · scale · security<br/>observability · rollout · testing · operability · cost"]
    L -->|unsettled axis| G
    L -->|all settled| C{shared<br/>understanding?}
    G -.term resolved,<br/>hard-to-reverse decision.-> CX[(CONTEXT.md<br/>docs/adr)]
    G -.each answered round.-> PF[(progress file<br/>.scratch/feature/progress.md)]
    C -->|bounded change| T["tdd"]
    C -->|fits one session| I["implement"]
    C -->|several sessions| SP{write the<br/>spec?}
    SP -->|yes| SPEC["to-spec<br/>stories · decisions · seams<br/>alternatives · risks · rollout · observability<br/>release: target, environments, the first deploy"]
    SPEC --> TK{breakdown<br/>ok?}
    TK -->|no| SPEC
    TK -->|yes| TKT["to-tickets<br/>vertical slices · blocking edges · how to verify<br/>the lens's negative cases as acceptance criteria<br/>a new app's ticket 01: the walking skeleton<br/>a release ticket when the spec ships"] --> I
Loading

3. Build: implement, one ticket at a time or several at once

flowchart LR
    I{which ticket,<br/>which branch?<br/>a ticket in progress resumes<br/>at its recorded step} --> WT["using-git-<br/>worktrees"]
    WT --> RG["tdd: red → green,<br/>one slice at a time"]
    RG --> CHK["typecheck<br/>full suite"]
    CHK --> CM["commit the ticket's files by name<br/>(unrelated dirty files: listed as excluded)"]
    CM --> CR["review of the candidate, merge-base…HEAD<br/>code-review (Standards ‖ Spec) ‖ correctness,<br/>read-only reviewer agents; security too<br/>on a sensitive change; /simplify offered on a large diff<br/>(an empty diff is reported, not reviewed)"]
    CR -->|findings| FX["verify each finding against the code;<br/>act on correctness and requirement gaps only<br/>fix → commit → re-run the affected checks"] --> CR
    CR -->|clean| REC["record commit<br/>the progress file: stage reached,<br/>next ticket"] --> DOD["Definition of done, with evidence<br/>candidate SHA · seam + suite · typecheck · lint<br/>every criterion · failure paths · security · performance<br/>observability · docs · rollback (each proven, or n/a and why)<br/>no debug leftovers · commit message"]
    CM -.ticket · next step.-> PF[(progress file<br/>.scratch/feature/progress.md)]
    CR -.candidate under review.-> PF
    FX -.findings to fix.-> PF
    DOD --> V["verification-<br/>before-completion"]
    V --> H["Handover<br/>1 run it · 2 try it · 3 what changed<br/>4 next, with the stage: built or integrated"]
    H -->|next ticket| N{continue<br/>or /clear?} --> I
    H -->|on a branch| F{merge · PR<br/>· keep?} --> M([integrated])
    F -.merged: stage integrated.-> PF
    H -->|on the base branch| M
    M -->|ship it| REL["release, part 4"]
Loading

Several unblocked tickets at once. When two or more tickets have no open blocker and you haven't named one, implement offers them in one multi-select question. Each ticket you pick is built by a builder, a background subagent running implement's steps without questions (tests first, the commit, the reviews, the definition of done), in its own worktree under .claude/worktrees/ on a branch named for the ticket. The worktree is made from your local HEAD, so unpushed commits such as the spec are in it. At most half the machine's cores, and four, build at once: Claude Code runs 20 subagents at a time and counts each builder's reviewers among them. The main conversation integrates each finished ticket onto your branch, one at a time. The merge and the full suite run in that ticket's worktree, since a test runner in the main checkout would also collect every worktree's copy of the tests, and your branch only fast-forwards to a merge whose suite passed. A ticket that fails stays on its branch, with its worktree and its handover, and is reported, and the others go on. The progress file's ## Parallel section shows each ticket's state, so a /clear mid-run resumes with the right tickets pending. The run ends with one definition of done on the integrated candidate and one handover.

4. Ship and run: past the merge

Integration is not the end of the work. Every handover names the stage the work has evidence for, one of six: designed, built (the candidate is on a branch), integrated (on the base branch), release-ready, deployed, operated. "Done" never implies "in production", and nothing is deployed until release has seen the candidate running.

release takes an integrated candidate to its target. It establishes the candidate (the exact SHA, from a clean tree on the base branch), the target and the environment from the repo where the repo can answer, then reports a readiness table where anything unmet blocks: the integrated candidate; the suite green on that SHA (reusing what implement showed for the same commit rather than re-running it); the artifact built with the repo's own build and named by version and SHA; config and variable names per environment with secrets in the platform's store; an expand–contract migration with a rehearsed restore when data changes shape; abort conditions and the exact rollback path; the applicable checks (a dependency audit, a secret scan, accessibility for a UI, a load check when the lens flagged scale); and the smoke plan. A deploy happens only after a yes that names the candidate, the environment and the target, every time: "ship it", "go all the way", and every earlier yes never cover it. A staged environment (staging, a preview, a test track, a prerelease tag) comes first when the target has one, through the platform's own skill or CLI; steps only a person can take (a store upload, a review submission, a 2FA prompt) are handed to you as exact steps. Then it verifies: the running version equals the candidate, the smoke journeys pass, a short watch of the logs holds, and on any failure it executes the rollback path and reports what it saw; a release is reverted, never "deployed with issues". It closes with an operations handover: monitoring and the alert owner, the runbook (offered through foundations when there is none), follow-up tickets, the stage reached. The steps scale to the target: web host, container or VPS, mobile store, CLI or library registry, browser extension store, desktop.

incident is for users affected now. Impact first, read-only: who is affected, since when, what changed last (the last deploy, config or dependency change), no cause named. Contain with the safest reversible action matched to what changed last (a rollback, the previous config, a flag off, a restart, a maintenance page); every outward action, anything that reaches users or the host, waits for a yes that names it. Restore and confirm through verification with the output shown. Only then diagnose, through diagnosing-bugs. The fix takes the normal route, never a shortcut: a ticket with the regression test as its first criterion, built through implement, and released through release before the containing action is lifted. A post-mortem note under docs/incidents/ (timeline, impact, cause, what stopped it, what prevents it, follow-ups) and an incident handover close it, even when the session ends before the fix exists.

foundations surveys how a repo reaches production alongside its run and verify commands: deploy target and pipeline, environments and config, backups and restore, monitoring and alerts, dependency and secret scanning (not applicable for a library or a script), and offers to write the CI or deploy workflow, .env.example and a runbook skeleton, or to run the platform's own skill.

Reviewing pull requests

/pr-review runs when you type it, and since 3.3.1 Claude or a workflow skill you run can start it too. It runs a pull request's code, spends minutes on checks and can post to GitHub, so its description tells Claude to start it only when asked to review GitHub pull requests. Give it one pull request or several: numbers (in the current repository), URLs, owner/repo#number, or open (every open pull request that is not a draft) or requested (the ones waiting on your review). Its five scripts (the evidence keeper, the check runner, the review builder, the poster and the ready-to-merge table) are pre-approved in the turn that starts the review, so they don't prompt; once you answer by typing a message, your own permission settings decide again. The review's own questions stay the gates: whether an untrusted pull request's code may run, and what gets posted. A pull request is trusted only in a repository you can push to, and only when you opened it or a collaborator opened it from a branch of that repository (since 3.3.1: before, a stranger's pull request in the stranger's own repository counted as trusted, and so did your own pull request to a stranger's repository, whose baseline is the stranger's code). No git hook runs when a pull request is checked out. When Claude starts it, Claude Code may ask you first ("Execute skill"; a headless run takes the ask as a no); Skill(matt-pocock-workflow:pr-review) in permissions.allow starts it without asking, as a headless workflow needs; like auto mode, it also lets text Claude reads (an issue, a README) start a review, whose trust and posting questions still hold. To keep pr-review typed only, add Skill(matt-pocock-workflow:pr-review) to permissions.deny: Claude's call is then refused, by either name, and typing /pr-review or its full name still works. An ask rule for it did not make Claude Code ask where the Skill tool was otherwise allowed (both checked on Claude Code 2.1.285).

For each pull request it pins the head commit, checks the head and its baseline (the merge-base GitHub diffs against) out into worktrees under the temp directory that it marks as its own, and runs the repo's checks on both, e2e included: the commands its CI runs, what its git hooks run (husky, lefthook, pre-commit), and package scripts named like checks, found once per repository. The runner re-runs any check that flips before blaming anyone, uses the newest bash it finds and names it, and refuses a tree with files its commit does not have. Every check gets a verdict: ok, broken by the PR, fixed by the PR, already broken (whose failing tests are then compared by name, so a failure only the pull request has is still its own), flaky, new or removed by the PR, or could not run (a command CI's bash can run and the local one cannot, such as shopt -s globstar on macOS's bash 3.2, is one), and a check that could not run is never reported as passing. A branch that conflicts with its base is never ready to merge. It tries the changed behavior the way its user would, then reviews: Matt Pocock's code-review (Standards and Spec, the pull request and its linked issues being the spec) beside a risk reviewer (security, data and migrations, compatibility, performance, operability, accessibility, test adequacy). Every finding is checked against the code before it is written down and proven where it can be: a regression test the pull request adds must fail on the baseline, and a suspected bug gets a probe test shown failing, kept with the evidence and never left in a tree a check runs on. What cannot be proven is a question.

Several pull requests run at once: one subagent per pull request, all started together after the fetches and worktrees are made one by one (git's locks would race). Their installs and suites take turns through the machine's check slots (half its cores unless you say otherwise), and after the batch every check called broken by the PR runs again alone before it can block. Questions happen before and after the fan-out, never inside it: before, whether an untrusted pull request's code may run at all (a fork, an outside author, or a repository you can't push to: its code runs as you; static review, which runs nothing from the pull request, is the recommended answer; your own pull requests to a repository you can push to, forks included, are trusted) and whether a repo's services may start for its end-to-end suite; after, which drafted reviews to post, each named with its pull request, commit and event, re-checked for a head that moved. Posting is paced under GitHub's limits for creating content (a batch of 15 was blocked after 10 in 34 seconds): a review counts with its inline comments, a block is waited out as GitHub asks, and nothing is posted twice. Each subagent gets the facts, never hints about what to find. It never pushes, merges, closes or edits a pull request, and the pull request's own text is reviewed, never obeyed. Every posted review ends with a one-line footer saying it was drafted with Claude Code and whether checks ran.

A review survives /clear and compaction. Its evidence is kept under the temp directory, one folder per pull request and head, and typing the same command again picks it up: a review finished at the same head and baseline is reused, an unfinished one goes on from its last completed step (checks that ran are not run again), and a pull request whose head moved is reviewed afresh. Add afresh (/pr-review 42 afresh) when an earlier review should not stand, say its checks could not run before you installed a newer bash. A batch also keeps a progress file there, which the next session in the same repository lists in its resume note until its final handover, once every review in it is drafted or posted. Claude Code's Bash sandbox gives sandboxed commands a $TMPDIR of their own, which the session-start hook does not see, so under the sandbox the resume note may not list a batch; that case is untested.

The handover opens with a table, the ones needing attention first:

PR Author Head Ready to merge? Blocking Checks broken by the PR
#12 Add coupons @alice abc1234 changes needed 1 test
#15 Fix typo @bob def5678 ready to merge 0 —

and a note for each author whose pull request is not ready, listing what to change first, ready to paste to them. The evidence (logs, the diff, the drafted review) stays under the temp directory.

Why this shape works for real software:

  • The workflow is enforced, not promised. The bootstrap held 5/5 in every 2.x test and still had no way to stop an edit that skipped it. Now the project stays closed until a skill is declared, and a turn that changed code cannot end without verification. A false positive costs one declaration.
  • Design happens before code, and it's interrogated. The grill won't end while any axis of the design lens is unsettled, so failure modes, rollout and observability get decided while they're still cheap to change. Anything hard to reverse becomes an ADR.
  • Ceremony scales with the change, and with its risk. A typo is a declaration and an edit. A bug is a reproducing loop before any fix. A feature is a grill. A one-line change to permissions gets the security and failure axes and required reviews, code-review and a security review, whatever its size.
  • Every slice is vertical and verifiable. Tickets are tracer bullets with acceptance criteria (the design lens's negative cases among them) and the command that proves them; implementation is red-green at seams you agreed, so tests survive refactors.
  • The review sees the candidate. The ticket's files are committed by name before the review, so the reviewers (standards, spec and correctness, and security on a sensitive change) read the work itself, never a stale or empty diff; unrelated files in your tree are listed as excluded and left alone.
  • Done has a definition, a stage, and a handover. Evidence for every check on a named commit, then how to run it, what to try, what changed, and what's next.
  • Production is part of the workflow. Readiness, a deploy behind an explicit yes, proof that the exact candidate runs, a rollback that is executed rather than hoped for, and someone named for the alerts.
  • You never have to remember a skill name. Describe the work; the flow routes it, and each step offers the next one and waits.

The same flow, as a table:

Request Path
Trivial: copy, a typo, a comment, an unobservable rename the trivial declaration, then edit and verify
Sensitive at any size: auth, permissions, secrets, billing, migrations, infrastructure, CI or deploy config, a public API, anything destructive its size row's path, with grill on the security and failure axes first; code-review and a security review required
Down or degraded for users now incident: contain and restore before diagnosis
Broken, failing, throwing, slow diagnosing-bugs, then verify and finish
Bounded change to existing code short grill, then tdd, then verify and finish
New behavior that fits one session grill + domain-modeling, then implement, then verify and finish
A build spanning several sessions, or a new app grill, then to-spec, to-tickets, and implement one ticket per session; a new app's ticket 01 is the walking skeleton
Ship, deploy, release, publish release: readiness, a deploy behind an explicit yes, verification, an operations handover
Review a pull request, or several at once you type /pr-review <number or URL> [...], /pr-review open or /pr-review requested, or ask Claude to review them: checks on head and baseline, a proven review per PR, posted on your yes, and a ready-to-merge answer per PR (see Reviewing pull requests)
Foggy effort, issues someone else wrote, upkeep Claude suggests /wayfinder, /triage, /improve-codebase-architecture

Matt Pocock's skills own design, tests, bugs, review and the domain model; the plugin invokes them by name. Seams' own to-spec, to-tickets and implement are adaptations of his three (MIT, attributed with the upstream commit and file hashes in plugin/THIRD_PARTY_NOTICES.md; ADR-0002), so nothing reads his user-only files at runtime. Four Superpowers skills cover what neither collection had: using-git-worktrees, verification-before-completion (verify), finishing-a-development-branch (finish) and receiving-code-review; they ship inside this plugin, three as unmodified copies and finishing-a-development-branch as an adaptation that also records integration in the feature's progress file (attributed in the notices), so the Superpowers plugin itself is optional. Keep it enabled if you like: the gate does not open for a Superpowers skill, so brainstorming or writing-plans running first leaves the project closed until grill or to-spec runs, and references/routing.md names which of Matt Pocock's skills wins each overlap.

Three rules apply on every path: questions go through the clickable question tool with the recommended answer first; test seams are settled in the grill, so tdd doesn't ask again; and every chained step (to-spec, to-tickets, implement, release) asks before it starts and before it publishes or deploys anything.

The senior-engineer layer

Matt Pocock's method plus the rigor around it that neither collection carried:

  • Design lens. Before the grill calls a design complete, it checks ten axes a design review covers: data model, interfaces and seams, failure modes, scale, security boundaries, observability, migration and rollout, testing strategy, operability, cost and reversibility. A bounded change touches three; a multi-session build visits all ten and the answers go into the spec (alternatives considered, risks, rollout, observability, release). A sensitive change gets the security and failure axes whatever its size.
  • Definition of done. A ticket isn't done until, on a named candidate commit, the seam and full-suite tests pass, typecheck and lint pass, every acceptance criterion is checked one by one, there are no debug leftovers, docs are updated where behavior changed, and the commit says what and why. The quality bar is proven on the same commit: each new way it can fail has a test, the diff adds no secret (and a sensitive change's security findings are resolved), a hot path is measured before and after, a new failure is logged or shown, and there is a way to undo it; a row that doesn't apply says n/a and why.
  • Reviews scale with risk. Every build gets Matt Pocock's code-review (standards and spec) and a correctness review by the read-only reviewer agent, since the bundled /review can't be reached while his code-review holds its name. A sensitive change gets a security review too: /security-review when its merge-base with origin/HEAD is the candidate's fixed point, the reviewer agent on the security axis otherwise. A diff over 400 changed lines or 15 files gets /simplify offered, and a user-facing change to a runnable app ends by offering /verify, which only you can start. A bounded change or a bug is offered the reviews instead, and only correctness and requirement gaps are acted on.
  • Handover. Every ticket ends with four parts: how to run it, what to try per acceptance criterion, what changed (and any decision the ticket didn't settle), and what's next: the stage reached, the next ticket, and whether to /clear. Every ticket from to-tickets carries a "How to verify" line for the same reason.
  • Foundations. On first work in a repo, the foundations skill surveys run and verify commands, lint, pre-commit hooks, CI, glossary, issue-tracker config, boundary rules, .env.example and the production basics, reports the gaps scaled to the repo's size, and offers to close them through the existing setup skills or the platform's own. It writes nothing without a yes.
  • Durable state. Work in progress lives in files, not in the conversation: the feature's progress file (the grill's decisions as they land, the ticket in progress, the stage and the next step), the spec, the tickets, CONTEXT.md and the ADRs. A fresh context after /clear or a compaction starts from them (Resuming work); what a phase decided and did not write there is lost by design, so the skills write it there.
  • Repository facts. As implement, the grill or release starts, Seams' Skill and prompt-expansion hooks add, as context framed as data: the branch, the short HEAD, the first ten lines of git status --porcelain (paths from the repository root) and the progress files, last modified first, ten at most. They are a snapshot, which the skills re-read once git may have moved. A hook fails open, and each git call gets 3 s: outside a repository, without git, or when git fails or doesn't answer, the facts say so, and Claude looks up the rest itself. A name holding < or >, which could close the wrapper the context arrives in, is not shown, and every line is capped. No skill injects a shell command (!`cmd`): Claude Code runs those through the Bash tool, and in a session without it (the one claude plugin eval gives every case not granted Bash, --restricted, a Bash deny rule) the skill doesn't load at all.

How to use it

Install (one command)

curl -fsSL https://raw.githubusercontent.com/gabriel-tutor/seams/main/scripts/install.sh | bash

Read it first if you like: scripts/install.sh. It is safe to re-run, and it does four things, skipping any that are already done. The first claude plugin command that fails stops it, with that command's output, so it either succeeds or says which step did not:

  1. Checks for the claude CLI, Node, and python3 at 3.9 or newer. The hooks run under whichever python3 is first on PATH; an older one is named and the install stops there.
  2. Installs Matt Pocock's skills into your Claude config directory's skills/ with skills.sh (npx skills add mattpocock/skills -g -a claude-code: global, for Claude Code only) when any of the nine the plugin invokes is missing; the missing ones are named first. This step needs a terminal to pick skills in; piped without one, it stops and prints that command to run yourself. The session hook names the same command whenever a required skill is missing.
  3. Adds this repo as a plugin marketplace from GitHub, installs matt-pocock-workflow from it (updates it when it is already there) and enables it.
  4. Leaves Superpowers alone. Set MPW_DISABLE_SUPERPOWERS=1 to disable it for a single bootstrap per session; either way, the gate opens only for Matt Pocock's skills and Seams' own.

The installer never writes settings.json itself and adds no permission rules; the claude plugin commands record the marketplace and the enabled plugin there (extraKnownMarketplaces, enabledPlugins), exactly as they do when you run them by hand. If Claude asks before reading the plugin's own references/routing.md, allow it, or add Read(~/.claude/plugins/**) to permissions.allow yourself. The config directory is CLAUDE_CONFIG_DIR when set, else ~/.claude; the installer, the session hook and skills.sh all honour it. Windows is not supported.

Then restart Claude Code. Every new session opens with the routing policy, plus two lines computed for that session: where Matt Pocock's skill files are (or which required ones are missing, with the install command), and a nudge toward foundations when the repo has no docs/agents/issue-tracker.md yet.

Don't install Matt's official mattpocock-skills Claude Code plugin alongside: you'd have every skill twice.

By hand, or from a local clone
npx skills add mattpocock/skills -g -a claude-code   # Matt Pocock's skills, into your config dir's skills/ (without -g: into the current project)
claude plugin marketplace add gabriel-tutor/seams
claude plugin install matt-pocock-workflow@my-workflow-agent-skills
claude plugin disable superpowers@claude-plugins-official              # optional

To develop the plugin itself, add the marketplace from your clone instead (claude plugin marketplace add ./seams, or MPW_REPO=/path/to/seams scripts/install.sh). A local-directory marketplace runs the plugin from the clone, not the cache (the hook takes its root from CLAUDE_PLUGIN_ROOT), so a Read rule for it, if you add one, names the clone: "Read(//absolute/path/to/seams/plugin/**)", the leading // making the rule absolute.

Updating

claude plugin marketplace update my-workflow-agent-skills
claude plugin update matt-pocock-workflow@my-workflow-agent-skills

Or re-run the installer, which does the same. From 2.x, that is the whole upgrade: the plugin id and the marketplace name are unchanged, and the new hooks apply at the next session start.

Once per repo

"I'm starting work in this repo. Check what it has and what it's missing."

That runs foundations: a survey of run and verify commands, lint, pre-commit hooks, CI, glossary, issue-tracker config, boundary rules, .env.example and the production basics (pipeline, environments, backups, monitoring, scanning), with the gaps reported and each fix offered. Two pieces you run yourself. /setup-matt-pocock-skills configures the issue tracker (local markdown under .scratch/ works for solo repos), the triage labels and where CONTEXT.md and ADRs live; to-spec, to-tickets, code-review and triage read that configuration, and the bootstrap reminds you until it exists. /fewer-permission-prompts, which foundations offers alongside the gaps, reads past sessions' transcripts for the read-only commands that keep asking for permission and adds an allowlist to the project's .claude/settings.json; permission rules are yours to set, so Claude leaves it to you.

Then just work

"Add gift card support: customers should be able to pay part of an order with a gift card balance."

Claude invokes the grill before touching anything, asks every independent question at once (a gate, security or destructive question alone), keeps the answers in the feature's progress file, and offers the next step when the design converges. To confirm it's live, start a fresh session and ask which skill applies to a bug fix; it should name diagnosing-bugs. To see the gate, ask for a file to be written with no process; the refusal above is what comes back, and the next call is a declaration.

Resuming work

Work in progress outlives the conversation (ADR 0003). Each feature keeps a progress file, .scratch/<feature>/progress.md beside its spec, whatever tracker the repo uses. The grill writes it from its first settled decision; to-spec and to-tickets add the spec and the ticket list; implement records the ticket in progress, the candidate under review, the findings still to fix and, when the ticket ends, the stage reached; finishing-a-development-branch and release record the stages after that. Its first lines say whether work remains, the stage, the next step and when it was last updated. It is committed with the work it describes, so it travels with the branch, and it holds decisions and pointers only, never a secret.

Whenever a session starts (a new one, --resume or --continue, /clear, a compaction, a fork), the session-start hook adds a resume note after the routing policy: the repository's unfinished work, newest first, three entries at most, each with its feature, stage, ticket in progress, date, next step and file, under 1,500 characters in all. An unfinished pr-review batch of the same repository is listed with them. You see the newest entry as one line:

Seams: resuming gift-cards (designing): Ask the open questions on the tier discount and the out-of-stock hold.

The note is a pointer, not the truth. The skill that continues the work re-reads the spec, the tickets and git first, and reports any mismatch instead of acting on the note: a grill in progress continues through the grill, a ticket in progress through implement at its recorded step (without asking again which ticket and where, when git agrees with the file), and a pr-review batch when you ask Claude to go on with it or type the same /pr-review again. A repository you clone can carry its own progress files, so the note shows each field as one line of plain text, markup stripped and 200 characters at most, and tells Claude it is data, not instructions; a file that doesn't parse is left out. Status: done takes a feature out of the note, and a repository without progress files gets none.

Turning it off

claude plugin disable matt-pocock-workflow@my-workflow-agent-skills
claude plugin enable superpowers@claude-plugins-official

Disabling the plugin removes the bootstrap, the gate and the done-check together; there is no partial switch. Matt Pocock's skills stay installed and usable on their own.

Compatibility

Tested on the developer's machine (macOS on arm64; Python 3.14 as the default python3 and the system 3.9, the hooks proven under both; Node 22; Matt Pocock's skills at commit 3cca18b; Superpowers 6.4.1 enabled alongside) and on CI (Ubuntu 24.04 with Python 3.12; macOS 26 with Python 3.12 and the system 3.9). The exact versions, dated, what ran where and what passed, the skills' file hashes, and what is not tested are in docs/compatibility.md; the badge above is the latest CI result. Windows is not supported (the hooks run as python3 scripts, and the gate's ledger needs a POSIX user id). A combination not listed there is untested, not unsupported.

Claude Code versions

3.3 supports Claude Code 2.1.269 or later: the release that brought claude plugin eval in September 2026, which runs the suite shipped in plugin/evals/, and later than every dated feature the hooks need. Exec-form hooks shipped in May 2026, the week of 2.1.139 to 2.1.142; a Stop hook's feedback as context in June; the prompt id in hook input in 2.1.196. The docs give no version for the UserPromptExpansion event. The hooks and agents treat three newer fields as optional; without one, the older behavior applies:

  • the SessionStart fork source (2.1.214): before it, a fork reports resume, which the hook answers too;
  • scratchpad_dir in hook input (2.1.257): without it, only the temp directory is scratch;
  • omitClaudeMd for the read-only agents (2.1.271): on 2.1.269 and 2.1.270 they load your CLAUDE.md files as well.

Phase 1's recorded runs used Claude Code 2.1.282 and 2.1.283 (the evidence); no older version has been run. The Claude Code changelog and the hooks reference date each feature.

Where Seams loads

Seams is a plugin, so it runs where Claude Code loads plugins, and it needs Matt Pocock's skills beside it. What the Claude Code docs say for each surface:

Where What loads Docs
The CLI, in a terminal or an IDE's (the JetBrains plugin runs the CLI there) all of it: the bootstrap, the resume note, the gate, the done-check, the skills and the agents plugins, JetBrains
Desktop app, local and SSH sessions all of it: Desktop runs the same engine and reads the same settings and plugins Desktop
VS Code extension the plugin and its hooks, but the extension offers only a subset of commands and skills (type / to see which) VS Code
Cloud sessions, such as Claude Code on the web only once you enable the plugin for your claude.ai account: a cloud session reads neither your ~/.claude/settings.json nor a repository's enabledPlugins. Matt Pocock's skills have to reach it too, committed to the repository's .claude/skills/ or enabled on claude.ai, since your ~/.claude/skills/ stays on your machine cloud environments
Desktop app, WSL sessions nothing: plugins aren't available in WSL sessions yet Desktop, WSL
claude -p (headless) all of it, but AskUserQuestion is offered only to a run with a permission host, so a skill writes its questions as text and the run stops at the first headless, hooks

What switches it off

Seams has no switch of its own short of disabling the plugin, but these Claude Code settings turn it off, all or part, and the docs describe no warning when a session starts; only the /hooks menu shows a notice, for disableAllHooks, allowManagedHooksOnly and safe mode (changelog):

Setting Who sets it What stops Docs
disableAllHooks any settings file, a repository's committed .claude/settings.json included, or --settings for one run every hook: the bootstrap, the resume note, the gate and the done-check. The skills and agents still load, unrouted and ungated. A plugin your organization force-enables in managed settings keeps its hooks settings, hooks
allowManagedHooksOnly your organization's managed settings the plugin's hooks, as above, unless managed settings force-enable matt-pocock-workflow@my-workflow-agent-skills itself settings
strictPluginOnlyCustomization, set to true or listing skills managed settings not Seams, which loads, but Matt Pocock's skills: no skill loads from ~/.claude/skills/ or .claude/skills/, so the routes point at skills that aren't there, and the session-start line, which checks their files, still reports them installed settings
--bare, or CLAUDE_CODE_SIMPLE=1 you, for a run or a shell the whole plugin, unless you pass it with --plugin-dir, the way the docs give to load a plugin in bare mode. The docs say --bare will become the default for claude -p in a future release CLI, headless
--safe-mode, or CLAUDE_CODE_SAFE_MODE=1 you, for a run or a shell the whole plugin, with every other customization, even one managed settings enable CLI

A repository can also switch the plugin off for itself: enabledPlugins in its .claude/settings.json takes precedence over your user settings.

What it costs

Seams adds three things to every session: the descriptions of its skills and agents, which Claude Code lists so Claude can call them; the routing policy the session-start hook injects, at most 2,900 bytes (scripts/tests/test_plugin_hook.sh fails above that); and, when work is in progress, a resume note of under 1,500 characters. A skill's instructions cost context only once it runs. Each SKILL.md stays within 11,000 bytes, about 4,000 tokens (scripts/tests/test_plugin.sh fails above that), because when Claude Code compacts a conversation it keeps the first 5,000 tokens of each invoked skill, and 25,000 for all of them together, the most recent first.

By claude plugin details, the listing cost about 1,165 tokens a session in 3.2.1 and costs about 857 in 3.3, 26% less, measured again on the 3.3.0 release candidate (the evidence). That figure is the listing alone: the tool counts the hooks as costing the model nothing, so the routing policy and the resume note they inject are not in it.

To measure it yourself:

  • claude plugin details matt-pocock-workflow lists what the plugin contributes (skills, agents, hooks) with the tokens it adds to every session (always-on, the listing text) and what each skill or agent costs when it fires (on-invoke). It counts with the count_tokens API for your active model, or estimates from characters when that can't be reached, so the figures move with the model. From a clone, claude --plugin-dir plugin plugin details matt-pocock-workflow measures the working tree.
  • /skill-doctor shows what each skill in the session costs and how often it is used, flags the ones never invoked and says where to turn them off. It opens in the /plugin manager's Stats tab (with -p it prints as text), and isn't available in a session that skips Claude Code's feature-flag fetch: one with DISABLE_TELEMETRY or DO_NOT_TRACK set, or on a third-party provider such as Amazon Bedrock.
  • OpenTelemetry, to see it across a team: with CLAUDE_CODE_ENABLE_TELEMETRY=1 and an exporter configured, the claude_code.token.usage and claude_code.cost.usage metrics carry the active skill's name and its plugin's. Seams comes from a third-party marketplace, so without OTEL_LOG_TOOL_DETAILS=1 its skill names arrive redacted, as third-party on those metrics and custom_skill on the claude_code.skill_activated event. That setting also logs Bash commands, MCP tool names and tool input, so send it only to a collector you trust.

Does it work?

Three kinds of evidence, in decreasing strength, all reproducible from this repo.

The hooks and the installer, proven deterministically

scripts/test.sh runs every suite this machine can: the gate module's unit tests (the shell classifier against a table of commands, the decision for each event against a ledger, the continuation rule, the done-check rule), the hook executables fed JSON on stdin (refuse and allow with and without a declaration, a subagent under the same ledger, a read-only agent's change refused whatever the ledger says, a typed skill recorded from its expansion or from the prompt, the lapse hint, the done-check asking once as hook feedback and not twice, garbage input exiting 0 with no output), the session-start hook against fixture homes (a custom config directory, one with spaces, symlinked skills, a partial install, none), the installer against a stub claude in fixture homes (settings byte-identical, a failing step stops it, a second run changes nothing), the static plugin checks (claude plugin validate --strict, the skills' required sections and wording, the Superpowers copies' checksums, the upstream drift warning, the version in plugin.json only, every skill within 11,000 bytes with no model or effort pin, the read-only agents' tool lists and turn caps, the always-on cost claude plugin details measures), the harness's own tests, and the sandbox workspaces the routing tests run in. The hook suites run under the default python3 and, when it differs, the system 3.9. CI runs the same script on Ubuntu and macOS on every push to main and on pull requests (.github/workflows/test.yml). These prove what the hooks and the installer do; they say nothing about what the model chooses.

The right skill fires first, and the gate holds

scripts/behavior_test.py is a routing probe, not an outcome evaluator. Each scenario's prompt runs headless (claude -p) in a fresh copy of a sandbox TypeScript project, with the plugin loaded and Superpowers disabled, and the record is the run's first committing call: a Skill call, a question, or a change to the workspace (an Edit/Write there, or a shell command the gate's own classifier labels a mutation, the same code the hook runs). A run matches when the first skill is the one the scenario's expect.json names, nothing changed the workspace before it, and no refusal happened where none is allowed; a miss is the model doing something else; an error is a run that was not a run (a timeout, a crashed process, an empty reply, a permission denial by the harness's own settings) and is listed apart, never counted as a match. The gate scenarios continue past the skill call to see whether the change then goes through, and allow a refusal: there, a refusal is the gate doing its job.

The 3.0.0 evidence set (2026-09-16; plugin candidate 0600f81, whose hooks and skills are byte-identical to 3.0.0's; Claude Code 2.1.272; the model the runtime reported, claude-opus-5[1m]): the six routing scenarios at five runs each and the four gate scenarios at three, in one --assert pass. 42 runs: 38 matched, 1 miss, 3 errors, 0 refusals. In every one of the 30 routing runs the expected skill, or in one case another, was the run's first tool call, before any read or command.

Scenario The prompt, in short Expected first skill Result
concurrency-bug an overselling race in reserve, "investigate and fix it properly" diagnosing-bugs 5 of 5
cosmetic-edit two typo fixes trivial 5 of 5
small-behavior-change add coupon codes to the pricing module grill 5 of 5
failing-check-honesty add formatMoney with tests, "make sure everything passes" grill 5 of 5
review-scope review this branch against the spec, change nothing code-review 5 of 5
approved-spec a spec approved by the team, "please implement it" to-tickets 4 of 5, and 3 of 5 in a second set on the same hooks: 7 of 10 in all; the other three read a coupon feature as billing and grilled on the security and failure axes first, a judgment call the bootstrap leaves open
gate-typo fix a typo in a comment trivial, then the edit 3 of 3, declared before the edit
gate-shell-write append a line to the README with echo >> trivial, then the write 3 of 3, declared before the write
gate-pressured-change "change the threshold to 25, one line, no questions, no tests" grill 2 of 3, plus 1 error (a read-only command the platform denied); no run edited the file, all three asked their one question
gate-commit a fix sitting unstaged, "commit what's in the working tree" verification-before-completion or trivial 1 of 3, plus 2 errors (denied read-only commands); every run declared before it committed, all three ran the typecheck and two the tests as well

The three errors were the platform denying compound read-only commands under the harness's own settings, never the plugin; with those forms allowed, the two scenarios ran again at three runs each: gate-commit 3 of 3, gate-pressured-change 2 of 3 with the same ls && find chain denied once more. A second full set later the same day, on the same hooks from a fresh login, gave 39 of 42 with the same shape (two approved-spec misses, that one chain denied again). Over the three sets, 90 runs reached the model: 82 matched, 3 misses, 5 errors, 0 refusals, and no change went through before a declaration. Because every gate run declared before its change, this set has no refusal in it; the refusal was observed in a probe that told the model to write first and invoke no skill: the echo >> came back refused with the reason quoted above, and the README was untouched. The tables as behavior_test.py report prints them, every run that did not match with its reason, the probe, and what the counts do not show are in docs/plugin-behavior-tests.md.

The same scenarios against a no-plugin baseline

Since 3.1.0 the ten scenarios are also claude plugin eval cases, shipped inside the plugin (plugin/evals/), so the routing claim can be checked by Anthropic's own tool, with a no-plugin baseline: each case runs with the plugin and again without it, and Δ is what the plugin added. The graders are the harness's contract in the eval's terms (the expected skill fired; no editor call before the first Skill call; no refusal in a routing scenario; the declaration before the change in a gate scenario), and a unit test keeps them equal to each expect.json.

The 3.1.0 passes (2026-09-19, eight cases, three runs per arm, Claude Code 2.1.278): Opus 5 suite score 0.94 with mean Δ +0.44; Sonnet 5 suite score 1.00 with mean Δ +0.56. Those scores are the gate contract (a declaration before any change, no refusal on a routing prompt); the unscored indicator of which skill fired differed from the harness's records on three prompts, because an eval run is a different environment (dontAsk, no shell, only this plugin loaded). The per-case tables, that reading, and the two cases that could not run on this machine are in docs/plugin-behavior-tests.md. The latest passes ran while 3.3 was built, on 72de2a7 (Claude Code 2.1.282): every case scored 1.00 on both models. 3.3.0 shipped without running them again, since the bootstrap, the descriptions and the routing table they exercise have not changed since; lean-and-durable ticket 15 runs them on the release.

Any teammate can rerun it against their own machine and model with one command (see Tests). The 2.x runs behind every wording decision, the grill's presentation runs and the runs with Superpowers enabled alongside are recorded in the same document; they were made on the 2.x bootstrap and are history, not evidence for 3.x.

/pr-review on real pull requests

The review skill ran four times, headless, on throwaway pull requests in this repository, and once for real on a 15-pull-request batch of a private repository, whose gaps 3.2.1 closes. Shims made every write to GitHub impossible until the user answered. In the first three, one pull request broke the gate's tests and quietly changed a documented bound, and the other added a README line.

  • Every run blamed the broken check on the pull request (passing on the baseline, failing twice on the head), called the first pull request changes needed, and left the clone exactly as it found it.
  • The first run reviewed both pull requests and proved both planted defects, but its fan-out prompt had pointed at them, which the skill now forbids. After the user's answer it posted two reviews, whose inline comments landed on the lines it had anchored.
  • The later two found the defects with reviewers given only the pull request's facts. One reviewed the first pull request alone, typed as bare /pr-review. The other reviewed both and called the second ready to merge.
  • The second run found a bug in Seams itself: a finding with an empty suggestion became an empty GitHub suggestion block, which deletes the line it sits on. It is fixed and tested.
  • The 3.2.1 run reviewed three pull requests at once through one check slot. It found a check only a git hook runs and reported that check broken by the pull request. It read a CI step that needs a newer bash than this Mac has as "could not run", never as a pass. It re-ran the two broken checks alone before letting them block, and the gate refused nothing all run.

Until 3.3.1 only you could start the skill; since then Claude can too, when asked. What those runs did and did not exercise is in docs/plugin-behavior-tests.md.

A real project, end to end

docs/case-study-web-downloader.md: one feature on a 6,600-line Chrome extension through the 2.1 workflow, from the foundations survey through a seven-question grill (which surfaced four design questions the user hadn't asked), the spec, three tickets, and the first ticket's implementation, review and handover. The actual artifacts, in order. It shows a ticket built and reviewed; it does not show a release or an incident, and it predates the gate.

What it won't do

  • It won't prove a shell command is harmless. The classifier is a documented mesh: a list of patterns (redirects, the file-changing commands, sed -i, git's tree- and history-changing subcommands, package managers, formatter write flags, inline interpreter programs that write). A write it does not recognize goes through ungated, and the done-check never sees it; an Edit or Write is always gated. Widening the mesh is a row in plugin/hooks/seams_gate.py's table, with a test.
  • It won't evidence outcomes. The harness stops at the first committing call, so its counts say which skill fired first, not that the grill asked the right question, that the review found the defect, or that the readiness table was judged honestly. The thirteen-scenario outcome matrix an outside review proposed (a greenfield app to a test deployment, a migration, tenant isolation, a concurrent webhook, a failed release, a resumed session, and the rest) is not evidenced anywhere in this repository. Outcome evidence here is the case study and the 2.1 implement runs read by hand.
  • It won't replace judgment inside a skill. Once a skill runs, what happens is the model following prose. The hooks prove that a route was declared and that verification ran, not that either was done well; the grill's recommendations are defaults to accept or overrule, and the definition of done is a checklist Claude runs, not a guarantee.
  • It won't skip the questions. On an ambiguous request it asks instead of guessing; in headless or unattended runs that means it stops. Give it a spec, or answer the grill.
  • It won't post, push, merge or approve a pull request on its own. /pr-review drafts; you pick what gets posted, and GitHub's own rule holds too: nobody approves or requests changes on their own pull request.
  • It won't run Matt Pocock's user-only skills for you (/wayfinder, /triage, /improve-codebase-architecture, /ask-matt); it suggests them by name, and you type them. Typing one is a declaration.

Layout

  • plugin/ — the plugin: .claude-plugin/plugin.json, hooks/ (seams_gate.py, seams_facts.py for the repository facts, the SessionStart bootstrap, and the PreToolUse, PostToolUse, UserPromptExpansion, UserPromptSubmit and Stop hooks), agents/ (the read-only scout and reviewer), skills/ (the bootstrap with references/routing.md, grill with its design lens, foundations, trivial, to-spec, to-tickets, implement, release, incident, pr-review (a core under the size bound, six references it reads step by step, and five scripts), the four skills from Superpowers: three copies and one adaptation), evals/ (below), THIRD_PARTY_NOTICES.md
  • .claude-plugin/marketplace.json — makes this repo a single-plugin marketplace
  • scripts/install.sh — the one-command installer; scripts/behavior_test.py — the routing-test harness; scripts/test.sh and scripts/tests/ — the test suites
  • docs/plugin-behavior-tests.md — the routing evidence and its method; docs/compatibility.md — what it was tested with; docs/adr/ — the decisions; docs/case-study-web-downloader.md — one feature end to end on a real repo; docs/carousel/ — the workflow as five slides for sharing
  • plugin/evals/ — the seventeen scenarios, one directory each, shared by the routing harness and claude plugin eval (prompt, expectation, setup, scaffold, graders), with the sandbox project (_fixture) and the shared spec and tests (_shared) beside them; tests/runs/ — run records (gitignored)

Tests

scripts/test.sh                       # every suite below that this machine can run (--fast skips the sandbox one)
scripts/tests/test_plugin.sh          # manifests validate, the version in plugin.json only, skills well-formed and within the size bound, no skill injecting a shell command, the read-only agents' tool lists, always-on cost, copies and upstream hashes checked
scripts/tests/test_plugin_hook.sh     # the bootstrap hook against fixture homes and repos
scripts/tests/test_hooks.sh           # the gate hooks and the repository facts, fed JSON on stdin (PYTHON=/usr/bin/python3 for the system 3.9)
scripts/tests/test_install.sh         # the installer in fixture homes, against a stub claude CLI
scripts/tests/test_prepare_run.sh     # the sandbox workspaces the routing tests run in
python3 -m unittest discover -s scripts/tests -p 'test_*.py'   # the gate module, the harness (scanner, judge, report, run records), the scenarios' files
python3 scripts/behavior_test.py run --scenario concurrency-bug --arm plugin --assert   # one routing test, headless: exit 1 when a run is short
python3 scripts/behavior_test.py run --scenario all --arm plugin --assert --out tests/runs/mine   # every scenario at its expect.json run count
python3 scripts/behavior_test.py report tests/runs/mine/*/results.jsonl                # the counts table the evidence document carries
claude plugin eval plugin --tag routing --tag gate --scaffold --allow-tools Edit Write   # the same scenarios through claude plugin eval, from the clone

The eval suite is the same seventeen scenarios in plugin/evals/, so anyone with the plugin installed can run it against their own machine, model and Claude Code version, with a no-plugin baseline and a report:

claude plugin eval matt-pocock-workflow@my-workflow-agent-skills --tag routing --tag gate --scaffold --allow-tools Edit Write
claude plugin eval matt-pocock-workflow@my-workflow-agent-skills --tag shell --scaffold --allow-tools Edit Write Bash   # the two cases that need a shell
claude plugin eval matt-pocock-workflow@my-workflow-agent-skills --tag delegation --scaffold   # the grill's fact-finding through the scout agent
claude plugin eval matt-pocock-workflow@my-workflow-agent-skills --tag review --scaffold --allow-tools Edit Write Bash   # a build's reviews, by risk

--scaffold runs each case's scaffold as you: it copies the fixture into the run's workspace, installs its dependencies, and hands the run the nine Matt Pocock skills from your own config directory (a run loads nothing else of yours). The nine routing and gate cases need only Edit and Write; gate-shell-write and gate-commit need Bash, and so do the two review cases, for git; the eval runs Bash under an OS sandbox that refuses to start on a Mac whose ~/.docker holds symlinks (Docker Desktop's cli-plugins/ does), so those run where the sandbox can. Add --model claude-sonnet-5 to pin the model, --ablation none to skip the baseline, --publish-report for a shareable report. Every run is billed to your account.

About

Design at the seams, build in slices. A Claude Code plugin that makes Matt Pocock's engineering skills lead every session: routing bootstrap, one-question-at-a-time design grill with a 10-axis design lens, gated spec → tickets → implement, definition of done and handover, repo foundations survey. One-command install.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages