Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,9 @@ You are a test architecture and coverage expert who evaluates whether the tests
- **Untested lifecycle branches** -- require coverage for every newly meaningful branch in lifecycle code, including "already loaded" guards and early-return branches after setup or global mutation. Do not accept only production-vs-non-production happy paths when the diff adds effect cleanup, script loading, event listener, timer, or DOM append/remove behavior.
- **Untested sentinel semantics** -- when a diff reuses an existing sentinel value (`null`, `undefined`, empty array/object, fallback enum) for a new meaning, require tests that prove consumers render, log, measure, or act on the new state truthfully. Tests that only prove the consumer does not crash are insufficient.
- **Mirror tests that miss the machine** -- for alignment, copy-list, or generated-shim tests, do not accept a test that compares one file to a hardcoded expected array or fixture unless the executable source of truth is checked too. Ask: "If the provisioner/source script changes but this expected array does not, does the test fail?" If no, report the missing source-of-truth assertion.
- **Tests that don't assert behavior (false confidence)** (violates Kent Beck's *behavior-sensitive* test desideratum) -- tests that call a function but only assert it doesn't throw, assert truthiness instead of specific values, or mock so heavily that the test verifies the mocks, not the code. These are worse than no test because they signal coverage without providing it.
- **Tests that don't assert behavior (false confidence)** (violates Kent Beck's *behavior-sensitive* test desideratum) -- tests that would still pass if the code under test were broken. Common shapes: asserting only that a call doesn't throw, or truthiness instead of specific values; an expected value computed by the same helper or renderer under test; a mock or fixture that supplies the result, ordering, or side effect the code under test is supposed to produce; a negative case rejected by a different guard than the one the test names. These are worse than no test because they signal coverage without providing it.
- **Production seams that exist only for tests** -- the diff adds an export, flag, wrapper, global, or injection hook that no production caller uses, so a test can reach an internal the real entry point could have exercised. Flag the seam and name that entry point. A seam that controls what no entry point can, such as time or randomness, is not this finding.
- **Duplicate coverage of one contract** -- a new test asserts a contract an existing test already owns, at another layer or as a near-copy, without a distinct risk the existing test cannot reach (such as a transport or lifecycle failure). Name the owning test and suggest extending it or a table-driven case instead.
- **Brittle implementation-coupled tests** (violates Kent Beck's *structure-insensitive* test desideratum) -- tests that break when you refactor implementation without changing behavior. Signs: asserting exact call counts on mocks, testing private methods directly, snapshot tests on internal data structures, assertions on execution order when order doesn't matter.
- **Nondeterministic or order-dependent tests** (violates Kent Beck's *deterministic* and *isolated* test desiderata) -- new tests that depend on real time (sleeps, `Date.now` without a fake clock), real network, shared mutable fixtures or module state another test also touches, or the order tests happen to run in. These pass today and flake later; flag the specific dependency, not "this might be flaky."
- **Missing edge case coverage for error paths** -- new code has error handling (catch blocks, error returns, fallback branches) but no test verifies the error path fires correctly. The happy path is tested; the sad path is not.
Expand Down
2 changes: 2 additions & 0 deletions skills/ce-work/references/implementation-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,8 @@ Guardrails for execution evidence:
- Do not skip verifying that a new or changed test fails for the expected reason before implementing the fix or feature
- Do not over-implement beyond the current behavior slice when working proof-first
- Do not add a duplicate regression test when an existing test is the right home; update or strengthen that test instead, then observe the failure before changing code
- A new or changed test must fail when the behavior it names breaks, and keep passing when only the implementation changes. It fails that bar when its expected value comes from the code under test, when a mock or fixture supplies the result the code should produce, or when it asserts calls between internal parts instead of what the code returns, stores, or sends across its boundary
- Do not add a production export, flag, wrapper, or hook that only tests use when the real entry point can drive the behavior; test through that entry point instead
- Skip proof-first discipline for trivial renames, pure configuration, pure styling, generated artifacts, and manual-only surfaces, but record the reason and replacement verification while continuing execution

**Test Discovery** — Before implementing changes to a file, find its existing test files (search for test/spec files that import, reference, or share naming patterns with the implementation file). When a plan specifies test scenarios or test files, start there, then check for additional test coverage the plan may not have enumerated. Changes to implementation files should be accompanied by corresponding test updates — new tests for new behavior, modified tests for changed behavior, removed or updated tests for deleted behavior.
Expand Down
Loading