diff --git a/.github/workflows/skill_evaluation.yaml b/.github/workflows/skill_evaluation.yaml new file mode 100644 index 0000000..e5766ec --- /dev/null +++ b/.github/workflows/skill_evaluation.yaml @@ -0,0 +1,55 @@ +name: Skill evaluation offline checks + +on: + pull_request: + paths: + - 'skills/cudaq-algorithms-dev/**' + - 'skills/cudaq-algorithms/**' + - '.github/workflows/skill_evaluation.yaml' + workflow_dispatch: + workflow_call: + +permissions: + contents: read + +jobs: + offline: + name: Suite, runner and reporting contracts + runs-on: ubuntu-latest + timeout-minutes: 15 + env: + PYTHONDONTWRITEBYTECODE: '1' + # The CI job never opts into namespace, scientific-runtime or model probes. + CUDAQ_RUNNER_TEST_ISOLATION: '0' + CUDAQ_RUNNER_TEST_PYTHON: '' + steps: + - name: Checkout repository + uses: actions/checkout@v4 + + - name: Set up Python + uses: actions/setup-python@v5 + with: + python-version: '3.12' + + - name: Install test dependencies + run: python -m pip install pytest jsonschema PyYAML numpy scipy pyscf openfermion + + - name: Validate canonical suite and private grading + if: ${{ !cancelled() }} + shell: bash + run: python -B -m unittest discover -s skills/cudaq-algorithms-dev/evals/tests + + - name: Test routing, coverage and reporting + if: ${{ !cancelled() }} + shell: bash + run: python -B -m pytest -q -p no:cacheprovider skills/cudaq-algorithms-dev/scripts/tests + + - name: Test runner with offline transports and trusted test commands + if: ${{ !cancelled() }} + shell: bash + run: python -B -m pytest -q -p no:cacheprovider skills/cudaq-algorithms-dev/evals/runner/tests + + - name: Check skill coverage and links + if: ${{ !cancelled() }} + shell: bash + run: python -B skills/cudaq-algorithms-dev/scripts/check_coverage.py diff --git a/.gitignore b/.gitignore index 406f1f2..d108076 100644 --- a/.gitignore +++ b/.gitignore @@ -8,3 +8,6 @@ __pycache__/ compile_commands.json timer.dat docs/sphinx/_build/ + +# Skill-development run evidence stays out of the repository +skills/cudaq-algorithms-dev/evals/results/ diff --git a/skills/cudaq-algorithms-dev/authoring/architecture.md b/skills/cudaq-algorithms-dev/authoring/architecture.md new file mode 100644 index 0000000..186c7c2 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/architecture.md @@ -0,0 +1,52 @@ +# Architecture + +## Design goal + +Represent CUDA-Q Algorithms as a durable library of independently selectable +scientific contracts. Applications consume and validate those contracts; they +do not define the taxonomy. + +This file is maintainer and reviewer policy. Scientific tasks normally begin at +[the catalog](../../cudaq-algorithms/references/catalog.md), not here. + +## Organization + +- `../../cudaq-algorithms/SKILL.md` is the application entry point. +- `../../cudaq-algorithms/references/catalog.md` selects nine scientific families; detailed + operation/object rows live in their family selectors. +- `../../cudaq-algorithms/references//.md` holds one independently selectable + primitive per focused file. A front door is navigation, not a multi-contract record. +- Shared representations and capabilities live in a focused support record + or family front door when multiple producers/consumers exchange them. +- `../../cudaq-algorithms/references/conventions.md` selects five convention records; + `validation.md` and `source-provenance.md` own validation and current-source lookup. +- `../../cudaq-algorithms/references/application-composition.md` describes application chains. +- `../../cudaq-algorithms/references/workflow.md` shares advisory and implementation guidance. +- This `authoring/` directory holds schemas, maintenance policy, and open design decisions. +- `../coverage/` holds the live contract inventory and per-feature evaluation mappings; + `../scripts/check_coverage.py` checks consistency. + +This development directory is not an application skill and has no `SKILL.md`. +The canonical delivery suite, evaluator configuration, required fixtures and +integrity checks live under `../evals/`; do not package them as application +guidance. Historical campaigns and superseded harnesses are archived outside +these delivery directories. Their results do not validate the current suite. + +Keep scientific family paths shallow, with no primitive subdirectories. Every +focused record must be linked from its family selector or directly from the root +catalog. Each record covers every applicable canonical schema field once; +short records may combine adjacent headings when boundaries remain explicit. + +## Authoring routes + +- [Record design](record-design.md): identity, granularity, metadata, composition, resources. +- [Capability design](capability-design.md): semantic boundaries and identifiers. +- [Source review](source-review.md): source authority and freshness. +- [Extension workflow](extension-workflow.md): lifecycle and coordinated changes. +- [Design decisions](design-decisions.md): open ownership and promotion questions. +- Templates: [primitive](templates/primitive-record-template.md), + [representation](templates/representation-record-template.md), + [capability](templates/capability-record-template.md), + [convention](templates/convention-record-template.md). +- [Coverage policy](../coverage/policy.md), [feature registry](../coverage/features.json), + and [delivery evaluation](../evals/EVAL.md). diff --git a/skills/cudaq-algorithms-dev/authoring/capability-design.md b/skills/cudaq-algorithms-dev/authoring/capability-design.md new file mode 100644 index 0000000..2086718 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/capability-design.md @@ -0,0 +1,23 @@ +# Capability Design + +## Capability composition + +Use identifiers of the form +`cudaq-algorithms..v`. Dotted capability names +are valid. The major version changes only for an incompatible semantic change. + +The currently adopted documentation identifiers are: + +- `cudaq-algorithms.state-preparation.unitary.v1`; +- `cudaq-algorithms.block-encoding.zero-flagged.v1`; +- `cudaq-algorithms.chemistry-integrals.v1`. + +Their status remains `provisional`; the identifier is resolved even though the +boundary has not been promoted to a stable taxonomy contract. + +Every capability record states its ID, status, owner, direction, boundary +representation, exact signature, semantic invariants, geometry, conventions, +execution boundary, providers, consumers, and unsupported conditions. +Composition requires the same ID and compatible major version, plus every +consumer invariant. Similar names and structural member presence are +insufficient. diff --git a/skills/cudaq-algorithms-dev/authoring/design-decisions.md b/skills/cudaq-algorithms-dev/authoring/design-decisions.md new file mode 100644 index 0000000..d04b6f9 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/design-decisions.md @@ -0,0 +1,46 @@ +# Design decisions and open questions + +## State-preparation ownership and capability promotion + +- Register or shape geometry and ownership: geometry is invariant 2 in the + [injection contract](../../cudaq-algorithms/references/state-preparation/injection-contract.md#capability-record--unitary-state-preparation). + At the + historical last review, ownership was **caller-owned in the reviewed source + only**: the consumer factory that called the preparation kernel allocated the + register, handed it over fresh in `|0...0>`, and expected exactly its own system + width. That is a `derived` historical description, **explicitly not a universal + or future policy.** Check current public source at use time. Whether the + caller or the primitive should own system and ancilla registers in general is + an **open** question in the + taxonomy design record outside this package (the + [kernel-boundary convention](../../cudaq-algorithms/references/conventions/kernel-boundaries.md) + repeats the scope limit); do not present the reviewed behavior as a library-wide guarantee, + nor the open question as settled. + +- Promotion criteria to a public protocol or compiler IR operation, and the + explicit decision still required: (a) a provider outside the two historically + reviewed providers that exercises the same boundary without widening it, (b) + a resolved register-ownership policy, (c) characterized behavior for the conditions + marked `unverified` in the + [injection boundaries](../../cudaq-algorithms/references/state-preparation/injection-contract.md#shared-unsupported-and-unverified-boundaries), + and (d) an explicit team decision recorded with an + owner. None of the four held at the historical last review; check current + public source and team records at use time. This record proposes no promotion. + +- **Scope limit.** This is a `derived` description of the injection seam found + during the last source review. Recheck the cited call sites in the current + checkout before relying on it. It is **not** a decided capability-level policy. + Whether the caller or the primitive should own system and ancilla registers + in general, and the exact input-state, width, ancilla, inverse, control, and + failure semantics of state preparation, remain **open** questions in the + taxonomy design record, which lives outside this skill package. Do not + present the current behavior as a library-wide guarantee for future + capabilities, and do not present the open questions as settled. + +## Capability status + +The adopted state-preparation, zero-flagged block-encoding, and chemistry-integral +documentation capability IDs remain provisional; they have not been promoted +to stable taxonomy contracts. The injection boundary has two packaged providers +and four independent consumer modules, but is not a public protocol, and stability +across future providers is not established. diff --git a/skills/cudaq-algorithms-dev/authoring/extension-workflow.md b/skills/cudaq-algorithms-dev/authoring/extension-workflow.md new file mode 100644 index 0000000..addc0b3 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/extension-workflow.md @@ -0,0 +1,62 @@ +# Extension Workflow + +Only add records for QROM, arithmetic, sparse-oracle, eigensolver, THC, or +additional resource-model concepts once current public source establishes a +contract. Roadmap names alone do not establish available primitives. + +## Lifecycle and evidence + +The following preserves the historical review vocabulary and promotion bar. +Store lifecycle history and executed evidence under +[coverage](../coverage/policy.md), not as status comments in operational +records. Current coverage uses scoped evidence, not blanket lifecycle stamps. + +`draft | verified | deprecated | removed` is the only record lifecycle +vocabulary. A capability separately uses +`candidate | provisional | stable taxonomy contract`. + +A record becomes `verified` only after its contract, runnable usage, relevant +scientific tests, and representative evals have actually passed on a recorded, +supported package/CUDA-Q version combination. Store exact source revisions, +dependencies, targets, and commands with that result. Source inspection alone +supports `source-checked` in the task that performs it, not `verified`. + +Deprecation records name the replacement, first deprecated version, and +behavioral differences. Incompatible contract changes prefer a new primitive +name and explicit migration over silent redefinition. + +## Growth rule + +Add knowledge in this order: + +1. shared representation or convention when demonstrated; +2. one independently selectable primitive contract; +3. a capability only when multiple producers or consumers justify it; +4. a composite protocol when its lower-level contracts are populated; +5. catalog routing, runnable usage, validation, and eval coverage in the same + change. + +Do not add records for roadmap concepts or speculative APIs. Split existing +records when real contracts have become independently selectable. + +## Incremental development rule + +Add or refine one independently selectable contract at a time. Update its +catalog entry, source provenance, complete contract, resource status, +independent validation method, runnable example or usage test, and evaluation +coverage together. Split a record whenever operation/object identity, return +type, execution layer, validation oracle, approximation behavior, resource +contract, or composition boundary can be selected independently. + +`verified` is a record lifecycle state, not shorthand for reading source or for +one successful run. Promotion requires a recorded, supported package/CUDA-Q +version combination; exact revisions, dependencies, targets, and commands +belong with the validation or evaluation result. + +## Skill evaluation versus scientific validation + +SkillEvaluator checks activation, routing, usefulness, safety, and answer +quality. It does not establish that a quantum circuit or numerical transform is +scientifically correct. Run baseline and with-skill eval arms for behavioral +uplift, and run repository tests or independent numerical oracles separately +for scientific claims. diff --git a/skills/cudaq-algorithms-dev/authoring/record-design.md b/skills/cudaq-algorithms-dev/authoring/record-design.md new file mode 100644 index 0000000..eca91ac --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/record-design.md @@ -0,0 +1,104 @@ +# Record Design + +## Primary identity and granularity + +The routing identity is: + +```text +scientific operation + mathematical object +``` + +Operations include prepare, load, encode, transform, evolve, measure, estimate, +synthesize, preprocess, analyze, and reconstruct. Objects include states, +fermionic operators, Pauli operators, integral tensors, block encodings, +polynomials, phase sequences, and resource descriptions. + +One primitive record corresponds to one contract a caller can select +independently. Split records when any of these differ materially: + +- operation or mathematical object; +- public entry point and input representation; +- return type or emitted kernel signature; +- execution layer or authorization implications; +- validation/rejection behavior or independent oracle; +- approximation/error behavior; +- resource contract; +- required/provided capability or composition boundary. + +Several symbols may remain in one record when they are inseparable parts of one +contract. One source module or class may require several records. File size is +evidence of a possible granularity problem, never the routing rule itself. + +## Record types + +1. **Primitive record:** one concrete, independently selectable operation. +2. **Representation record:** the meaning of an exchanged object; create only + after multiple producers and consumers interpret the same form. +3. **Capability record:** a reusable semantic composition boundary; create + only after multiple independent producers or consumers demonstrate it. + +These are documentation records, not automatic requests for a Python protocol, +ABC, compiler IR operation, or new public API. + +## Orthogonal metadata + +Classify, but do not route or organize directories, by: + +- **Kind:** quantum operation, classical transformation, + measurement/readout, simulation-only analysis, resource estimator. +- **Routine role:** driver, computational, auxiliary. Role describes problem + completeness, not where code executes. +- **Execution layer:** host preprocessing, kernel factory, device kernel, + observable/measurement, simulation-only host path, or mixed. +- **Abstraction:** leaf operation or composite protocol. +- **Parameterization:** none, construction-time, runtime. +- **Representation and capabilities.** +- **Exactness, uncertainty, and method.** +- **Domain, dependencies, error contract, resource contract, lifecycle.** + +Lifecycle is historical coverage metadata; the other scientific classifications +remain part of the live contract. Follow [coverage policy](../coverage/policy.md) +for that placement distinction. + +Host transforms such as chemistry loaders and factorizations are computational +routines when they solve an independently useful problem. A simulation-only +helper is not a hardware primitive merely because it consumes one. + +## Composite protocols + +A reusable driver may itself be a primitive. Its record must state required +lower-level capabilities, a source-grounded reference composition, applicability +conditions, alternatives, propagated conventions/errors/resources, and what an +alternative component must preserve. Never silently replace the reference +composition with a target-specific heuristic. + +## Resource claims + +Every executable primitive either gives a resource contract or explicitly says +that none exists. Every quantity identifies metric/unit, abstraction level, +architecture/execution assumptions, exact/bounded/estimated/measured status, +controlling parameters, confidence/limitations, and composition rule if known. + +Never compare logical operations, decomposition proxies, transpiled gates, +runtime, memory, or measured hardware cost as if they were one metric. Never +turn a benchmark or source comment into a fresh measurement. + +## HF/UCC record application + +The [HF/UCC record](../../cudaq-algorithms/references/state-preparation/state-preparation-hf-ucc.md) +is a **concrete primitive record**: it instantiates each applicable +canonical Primitive-Record field from `templates/primitive-record-template.md` +exactly once, for one contract — a Hartree-Fock reference occupation optionally +followed by a UCC product at amplitudes the caller already knows. The shared +seam it plugs into (kernel representation, unitary capability, consumer table, +common boundaries) is [injection contract](../../cudaq-algorithms/references/state-preparation/injection-contract.md); +cross-cutting layout, ownership, and validation conventions are +[the convention selector](../../cudaq-algorithms/references/conventions.md). Shared scientific +detail is linked rather than repeated; lifecycle and run bookkeeping follows +the template's separate coverage destination. + +## Provisional HF/UCC classification discussion + +The HF/UCC method is a fixed-parameter ansatz product. `ansatz` was a provisional +extension beyond the `direct | variational | heuristic` vocabulary. Exactness, +uncertainty, and method are recorded values, never routing identities. diff --git a/skills/cudaq-algorithms-dev/authoring/source-review.md b/skills/cudaq-algorithms-dev/authoring/source-review.md new file mode 100644 index 0000000..03096f1 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/source-review.md @@ -0,0 +1,44 @@ +# Source Review + +## Source ownership and freshness + +Current public code and authoritative tests in the checked-out repository +control API behavior. [Source lookup](../../cudaq-algorithms/references/source-provenance.md) +records common source/test/example locations; each record adds only +contract-specific stable symbols and paths. Follow the +[coverage policy](../coverage/policy.md) for evidence claims. Archived reviews +are audit material, not current contracts or compatibility promises. + +When a selected record differs from current source or tests: + +1. compare the relevant public symbol and tests; +2. treat current public source as authoritative for generated code; +3. report the drift and evidence level; +4. update a maintained record only when that update is in scope; +5. never retain a stale line-number claim merely because the prose is familiar. + +Do not copy historical repository hashes or dependency pins into primitive +records. Historical last-review values belong with the archived review. A validation or evaluation result instead records the exact revision and +dependencies actually used by that run. + +Use line numbers only for a non-obvious invariant that benefits from a precise +anchor. Prefer stable symbol and test names for ordinary provenance. + +## Freshness check + +Before implementing from a maintained record: + +```bash +git rev-parse HEAD +git status --short -- \ + python/cudaq_algorithms tests/python docs/sphinx \ + pyproject.toml .cudaq_version +``` + +Inspect only the selected symbols and their tests. If public signatures or +scientific assertions changed, follow current source for generated code, report +the drift, and update the maintained record only when that work is in scope. Do +not silently upgrade the lifecycle from `draft` or claim a newly verified +version range. If comparing against a previous review, obtain its actual +recorded revision from that review; do not assume a fixed historical hash is +the baseline for today's checkout. diff --git a/skills/cudaq-algorithms-dev/authoring/templates/capability-record-template.md b/skills/cudaq-algorithms-dev/authoring/templates/capability-record-template.md new file mode 100644 index 0000000..45c7d40 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/templates/capability-record-template.md @@ -0,0 +1,18 @@ +# Optional capability record + +Create only when multiple independent producers or consumers demonstrate a +reusable semantic boundary. + +- Stable ID: +- Status: candidate | provisional | stable taxonomy contract +- Contract type: documentation-only | source-level protocol (name it) +- Owner: +- Boundary representation and exact signature: +- Semantic invariants: +- Register or shape geometry and ownership: +- Convention requirements: +- Host/device/simulation boundary: +- Providers: +- Consumers: +- Unsupported and unverified conditions: +- Promotion criteria and decision owner: diff --git a/skills/cudaq-algorithms-dev/authoring/templates/convention-record-template.md b/skills/cudaq-algorithms-dev/authoring/templates/convention-record-template.md new file mode 100644 index 0000000..f0c20ca --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/templates/convention-record-template.md @@ -0,0 +1,14 @@ +# Convention record template + +## Convention record + +Each populated convention must state: + +- concept and canonical symbol; +- CUDA-Q Algorithms convention; +- common alternatives; +- exact translation at the boundary; +- invariants unaffected by representation; +- observable symptom of a mismatch; +- minimal independent verification; +- source paths, tests, applicable package requirements, and validation evidence. diff --git a/skills/cudaq-algorithms-dev/authoring/templates/primitive-record-template.md b/skills/cudaq-algorithms-dev/authoring/templates/primitive-record-template.md new file mode 100644 index 0000000..fa66ad2 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/templates/primitive-record-template.md @@ -0,0 +1,166 @@ +# [Primitive name] + +Operation + object: **[operation]** a **[mathematical object]**. + +This authoring checklist spans two destinations. Scientific contract fields +belong in the operational record. Lifecycle history, executed-version/run +records and authored evaluation mappings belong in +[coverage bookkeeping](../../coverage/policy.md) and the corresponding run evidence, +not in application references. The former `Status: draft.` line is retained +here as historical schema guidance, not text to copy into a new record. + +This file documents one independently selectable scientific contract. If two +operations differ in return type, execution layer, validation, approximation, +resources, or composition, create two records and link them from a family front +door. + +Use `Not applicable:` when a field cannot apply to this contract. Use +`Deferred:` only when the field is required but evidence is currently missing; +name the evidence or event that will resolve it. Never say to omit a heading and +then put a deferred statement under that omitted heading. + +## Identity and provenance + +- Owner: +- Public symbols and import paths: +- Contract-specific source paths: +- Authoritative tests: +- Authoritative documentation and runnable examples: +- Source provenance: [source-provenance.md](../../../cudaq-algorithms/references/source-provenance.md) +- Package/CUDA-Q versions executed (coverage run evidence): +- Historical lifecycle (archived review): draft | verified | deprecated | removed +- Implementation evidence (coverage, with scope): documented | implemented | compiled | executed | + numerically validated +- Replacement and migration notes: + +## Classification + +- Operation + mathematical object (primary identity): +- Kind: quantum operation | classical transformation | measurement/readout | + simulation-only analysis | resource estimator +- Routine role: driver | computational | auxiliary +- Abstraction level: leaf operation | composite protocol +- Parameterization: none | construction-time | runtime +- Execution layers: +- Input representations: +- Output representations: +- Domain: domain-independent | quantum-chemistry | other tag +- Required dependencies: +- Optional dependencies: +- Exactness: +- Uncertainty: +- Method: + +## Scientific contract + +- Purpose: +- Mathematical definition: +- Why and when to use: +- When not to use: +- Approximation controls: + +## Inputs + +- Arguments: +- Shapes/ranks: +- Dtypes/domains: +- Units: +- Ordering/layout: +- Normalization: +- Required mathematical properties: +- Validation and rejection behavior: + +## Outputs + +- Return type or emitted kernel signature: +- Mathematical meaning: +- Shape/register geometry: +- Normalization, sign, and phase: +- Observable or measurement interpretation: +- Error/status information: + +## Capabilities and composition + +For every provided or required capability: + +- Stable ID: `cudaq-algorithms..v` +- Capability status: candidate | provisional | stable taxonomy contract +- Direction: provides | requires +- Owning record: +- Boundary representation and exact signature: +- Semantic invariants: +- Shape/register geometry: +- Normalization, sign, phase, and ordering: +- Convention requirements: +- Host/device/simulation boundary: +- Unsupported conditions: + +If no capability applies, say `Not applicable:` and explain the direct concrete +composition boundary. A capability ID is a documentation identifier unless the +record names a source-level protocol or public symbol. + +## Composite protocol + +For `composite protocol` records: + +- Required lower-level capabilities: +- Canonical reference composition: +- Default recipe and applicability conditions: +- Materially different alternatives: +- Propagated conventions: +- Propagated errors: +- Propagated resources: +- Component-substitution requirements: + +For leaf records, write `Not applicable: leaf operation.` + +## Accuracy and limitations + +- Error behavior or bounds: +- Precision sensitivity: +- Unsupported inputs: +- Known implementation limitations: +- Unsupported, absent, and unverified behavior: + +## Resources + +For every quantity, state metric/unit, abstraction level, +architecture/execution assumptions, exact/bounded/estimated/measured status, +controlling parameters, confidence/limitations, and composition rule. If no +estimator exists, say so and document only exact structural facts. + +## Validation + +- Independent oracle: +- Invariants: +- Representative cases: +- Predeclared tolerances: +- Expected failure/adversarial case: +- Runnable example or usage test: +- Execution record (coverage evidence): command, date, package/CUDA-Q version, target, precision, + result +- Evidence status per claim: derived | source-checked | compiled | executed | + numerically validated | measured | assumed | unverified. Add `unexecuted` as + an execution-state qualifier when no successful run occurred in the current + task. Reserve `source-checked` for current-task inspection; describe durable + historical evidence with its recorded scope and retain its review context + under the [coverage policy](../../coverage/policy.md). Source provenance provides + current-source lookup, not historical run authority. + +## Evaluation coverage + +Store this section's mappings and results in the coverage registry, not in the +operational record. + +- Positive selection/application: +- Convention or misconception: +- Capability composition: +- Invalid/unsupported boundary: +- Negative activation: +- Eval status: authored | baseline run | with-skill run | compared + +## External alignment + +- Literature conventions: +- External package translations: +- Known semantic differences: diff --git a/skills/cudaq-algorithms-dev/authoring/templates/representation-record-template.md b/skills/cudaq-algorithms-dev/authoring/templates/representation-record-template.md new file mode 100644 index 0000000..c7a8ad9 --- /dev/null +++ b/skills/cudaq-algorithms-dev/authoring/templates/representation-record-template.md @@ -0,0 +1,16 @@ +# Optional representation record + +Create only when multiple producers and consumers exchange the same object. + +- Object name and canonical symbol: +- Public type or structural form: +- Mathematical meaning: +- Shape, layout, ordering, dtype, and units: +- Normalization, sign, and phase convention: +- Applicability preconditions: +- Producers: +- Consumers: +- Invariants: +- Observable symptom of misinterpretation: +- Unsupported or ambiguous forms: +- Source/tests/docs/example evidence: diff --git a/skills/cudaq-algorithms-dev/coverage/features.json b/skills/cudaq-algorithms-dev/coverage/features.json new file mode 100644 index 0000000..8c7f969 --- /dev/null +++ b/skills/cudaq-algorithms-dev/coverage/features.json @@ -0,0 +1,634 @@ +{ + "schema_version": 1, + "behavioral_eval_suites": [ + "evals/evals.json" + ], + "features": [ + { + "id": "state-preparation-givens-schedule", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-givens-schedule.md", + "symbols": [ + "cudaq_algorithms.stateprep.make_givens_rotation_schedule" + ], + "behavioral_evals": [ + "state-preparation-provider-selection" + ] + }, + { + "id": "state-preparation-slater-determinant-kernel", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-slater-determinant-kernel.md", + "symbols": [ + "cudaq_algorithms.stateprep.slater_determinant_kernel" + ], + "behavioral_evals": [ + "state-preparation-provider-selection" + ] + }, + { + "id": "state-preparation-hf-ucc", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-hf-ucc.md", + "symbols": [ + "cudaq_algorithms.stateprep.hartree_fock_ucc_kernel" + ], + "behavioral_evals": [ + "state-preparation-provider-selection", + "state-preparation-ucc-parameterization-boundary", + "state-preparation-uccsd-open-shell-parity", + "state-preparation-width-mismatch-unknown" + ] + }, + { + "id": "state-preparation-kernel-hartree-fock", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-hartree-fock.md", + "symbols": [ + "cudaq_algorithms.stateprep.hartree_fock" + ], + "behavioral_evals": [ + "state-preparation-reference-excitation-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-hartree-fock-occupation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-hartree-fock-occupation.md", + "symbols": [ + "cudaq_algorithms.stateprep.hartree_fock_occupation" + ], + "behavioral_evals": [ + "state-preparation-reference-excitation-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-single-excitation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-single-excitation.md", + "symbols": [ + "cudaq_algorithms.stateprep.single_excitation" + ], + "behavioral_evals": [ + "state-preparation-reference-excitation-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-double-excitation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-double-excitation.md", + "symbols": [ + "cudaq_algorithms.stateprep.double_excitation" + ], + "behavioral_evals": [ + "state-preparation-reference-excitation-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-uccsd", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-uccsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.uccsd" + ], + "behavioral_evals": [ + "state-preparation-ucc-parameterization-boundary", + "state-preparation-uccsd-open-shell-parity" + ] + }, + { + "id": "state-preparation-kernel-uccgsd", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-uccgsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.uccgsd" + ], + "behavioral_evals": [ + "state-preparation-grouped-ucc-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-upccgsd", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-upccgsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.upccgsd" + ], + "behavioral_evals": [ + "state-preparation-grouped-ucc-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-ceo", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-ceo.md", + "symbols": [ + "cudaq_algorithms.stateprep.ceo" + ], + "behavioral_evals": [ + "state-preparation-grouped-ucc-device-boundary", + "operator-pool-ceo-units" + ] + }, + { + "id": "state-preparation-kernel-fixed-parameter-ucc", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-fixed-parameter-ucc.md", + "symbols": [ + "cudaq_algorithms.stateprep.fixed_parameter_ucc" + ], + "behavioral_evals": [ + "state-preparation-ucc-parameterization-boundary", + "state-preparation-grouped-ucc-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-givens-rotation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-givens-rotation.md", + "symbols": [ + "cudaq_algorithms.stateprep.givens_rotation" + ], + "behavioral_evals": [ + "state-preparation-givens-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-phase-givens-rotation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-phase-givens-rotation.md", + "symbols": [ + "cudaq_algorithms.stateprep.phase_givens_rotation" + ], + "behavioral_evals": [ + "state-preparation-givens-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-slater-determinant", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-slater-determinant.md", + "symbols": [ + "cudaq_algorithms.stateprep.slater_determinant" + ], + "behavioral_evals": [ + "state-preparation-givens-device-boundary" + ] + }, + { + "id": "state-preparation-kernel-complex-slater-determinant", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-kernel-complex-slater-determinant.md", + "symbols": [ + "cudaq_algorithms.stateprep.complex_slater_determinant" + ], + "behavioral_evals": [ + "state-preparation-givens-device-boundary" + ] + }, + { + "id": "operator-pool-uccsd", + "family": "state-preparation", + "record": "references/state-preparation/operator-pool-uccsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.make_uccsd_operator_pool" + ], + "behavioral_evals": [ + "state-preparation-ucc-parameterization-boundary", + "state-preparation-uccsd-open-shell-parity", + "operator-pool-selection-boundary" + ] + }, + { + "id": "operator-pool-uccgsd", + "family": "state-preparation", + "record": "references/state-preparation/operator-pool-uccgsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.make_uccgsd_operator_pool" + ], + "behavioral_evals": [ + "operator-pool-selection-boundary" + ] + }, + { + "id": "operator-pool-upccgsd", + "family": "state-preparation", + "record": "references/state-preparation/operator-pool-upccgsd.md", + "symbols": [ + "cudaq_algorithms.stateprep.make_upccgsd_operator_pool" + ], + "behavioral_evals": [ + "operator-pool-selection-boundary" + ] + }, + { + "id": "operator-pool-ceo", + "family": "state-preparation", + "record": "references/state-preparation/operator-pool-ceo.md", + "symbols": [ + "cudaq_algorithms.stateprep.make_ceo_operator_pool" + ], + "behavioral_evals": [ + "operator-pool-ceo-units" + ] + }, + { + "id": "state-preparation-resources-givens", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-resources-givens.md", + "symbols": [ + "cudaq_algorithms.stateprep.estimate_givens_resources" + ], + "behavioral_evals": [ + "state-preparation-provider-selection" + ] + }, + { + "id": "state-preparation-resources-hartree-fock", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-resources-hartree-fock.md", + "symbols": [ + "cudaq_algorithms.stateprep.estimate_hartree_fock_resources" + ], + "behavioral_evals": [ + "state-preparation-provider-selection" + ] + }, + { + "id": "state-preparation-resources-hartree-fock-occupation", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-resources-hartree-fock-occupation.md", + "symbols": [ + "cudaq_algorithms.stateprep.estimate_hartree_fock_occupation_resources" + ], + "behavioral_evals": [ + "state-preparation-provider-selection" + ] + }, + { + "id": "state-preparation-resources-fixed-parameter-ucc", + "family": "state-preparation", + "record": "references/state-preparation/state-preparation-resources-fixed-parameter-ucc.md", + "symbols": [ + "cudaq_algorithms.stateprep.estimate_fixed_parameter_ucc_resources" + ], + "behavioral_evals": [ + "state-preparation-provider-selection", + "state-preparation-ucc-parameterization-boundary" + ] + }, + { + "id": "pauli-lcu", + "family": "block-encoding", + "record": "references/block-encoding/pauli-lcu.md", + "symbols": [ + "cudaq_algorithms.PauliLCU" + ], + "behavioral_evals": [ + "block-encoding-capability-boundary" + ] + }, + { + "id": "qubitization-walk", + "family": "qubitization", + "record": "references/qubitization/qubitization-walk.md", + "symbols": [ + "cudaq_algorithms.Walk.kernel" + ], + "behavioral_evals": [ + "qubitization-walk-moment-boundary" + ] + }, + { + "id": "qubitization-moments", + "family": "qubitization", + "record": "references/qubitization/qubitization-moments.md", + "symbols": [ + "cudaq_algorithms.Walk.moment", + "cudaq_algorithms.Walk.moments" + ], + "behavioral_evals": [ + "state-preparation-injection-composition", + "qubitization-walk-moment-boundary", + "repository-implementation-third-moment" + ] + }, + { + "id": "qsvt-sequence", + "family": "qsvt", + "record": "references/qsvt/qsvt-sequence.md", + "symbols": [ + "cudaq_algorithms.PhaseSequence", + "cudaq_algorithms.QSVT.kernel", + "cudaq_algorithms.QSVT.controlled_kernel" + ], + "behavioral_evals": [ + "qsvt-phase-and-recovery-boundary" + ] + }, + { + "id": "qsvt-recovery", + "family": "qsvt", + "record": "references/qsvt/qsvt-recovery.md", + "symbols": [ + "cudaq_algorithms.recover_real_time_evolution" + ], + "behavioral_evals": [ + "qsvt-phase-and-recovery-boundary", + "qsvt-paraphrase-convention" + ] + }, + { + "id": "trotter-planning", + "family": "trotter", + "record": "references/trotter/trotter-planning.md", + "symbols": [ + "cudaq_algorithms.trotter.make_trotter_terms", + "cudaq_algorithms.TrotterOrdering" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "trotter-kernel-factory", + "family": "trotter", + "record": "references/trotter/trotter-kernel-factory.md", + "symbols": [ + "cudaq_algorithms.Trotter.kernel", + "cudaq_algorithms.trotter.Trotter.kernel" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "trotter-state-kernel-factory", + "family": "trotter", + "record": "references/trotter/trotter-state-kernel-factory.md", + "symbols": [ + "cudaq_algorithms.Trotter.state_kernel", + "cudaq_algorithms.trotter.Trotter.state_kernel" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "trotter-apply-kernel", + "family": "trotter", + "record": "references/trotter/trotter-apply-kernel.md", + "symbols": [ + "cudaq_algorithms.trotter.apply_trotter" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "trotter-resources-planned", + "family": "trotter", + "record": "references/trotter/trotter-resources-planned.md", + "symbols": [ + "cudaq_algorithms.Trotter.resources", + "cudaq_algorithms.trotter.Trotter.resources" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "trotter-resources-raw", + "family": "trotter", + "record": "references/trotter/trotter-resources-raw.md", + "symbols": [ + "cudaq_algorithms.trotter.estimate_trotter_resources" + ], + "behavioral_evals": [ + "trotter-evolution-resource-boundary" + ] + }, + { + "id": "jordan-wigner", + "family": "fermion-transforms", + "record": "references/fermion-transforms/jordan-wigner.md", + "symbols": [ + "cudaq_algorithms.fermion.jordan_wigner" + ], + "behavioral_evals": [ + "fermion-transform-selection-boundary", + "chemistry-end-to-end-composition" + ] + }, + { + "id": "bravyi-kitaev", + "family": "fermion-transforms", + "record": "references/fermion-transforms/bravyi-kitaev.md", + "symbols": [ + "cudaq_algorithms.fermion.bravyi_kitaev" + ], + "behavioral_evals": [ + "fermion-transform-selection-boundary" + ] + }, + { + "id": "chemistry-from-fcidump", + "family": "chemistry", + "record": "references/chemistry/chemistry-from-fcidump.md", + "symbols": [ + "cudaq_algorithms.chemistry.from_fcidump" + ], + "behavioral_evals": [ + "chemistry-end-to-end-composition" + ] + }, + { + "id": "chemistry-from-pyscf", + "family": "chemistry", + "record": "references/chemistry/chemistry-from-pyscf.md", + "symbols": [ + "cudaq_algorithms.chemistry.from_pyscf" + ], + "behavioral_evals": [ + "chemistry-bridge-dependency-boundary" + ] + }, + { + "id": "chemistry-from-psi4", + "family": "chemistry", + "record": "references/chemistry/chemistry-from-psi4.md", + "symbols": [ + "cudaq_algorithms.chemistry.from_psi4" + ], + "behavioral_evals": [ + "chemistry-bridge-dependency-boundary" + ] + }, + { + "id": "chemistry-spin-orbital-tensors", + "family": "chemistry", + "record": "references/chemistry/chemistry-spin-orbital-tensors.md", + "symbols": [ + "cudaq_algorithms.chemistry.spin_orbital_tensors" + ], + "behavioral_evals": [ + "chemistry-bridge-dependency-boundary", + "chemistry-end-to-end-composition" + ] + }, + { + "id": "chemistry-qubit-hamiltonian", + "family": "chemistry", + "record": "references/chemistry/chemistry-qubit-hamiltonian.md", + "symbols": [ + "cudaq_algorithms.chemistry.qubit_hamiltonian" + ], + "behavioral_evals": [ + "chemistry-bridge-dependency-boundary" + ] + }, + { + "id": "double-factorization-explicit", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-explicit.md", + "symbols": [ + "cudaq_algorithms.double_factorization.explicit_double_factorization" + ], + "behavioral_evals": [ + "double-factorization-host-contracts" + ] + }, + { + "id": "double-factorization-compressed", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-compressed.md", + "symbols": [ + "cudaq_algorithms.double_factorization.compressed_double_factorization" + ], + "behavioral_evals": [ + "double-factorization-encoding-boundary", + "chemistry-end-to-end-composition" + ] + }, + { + "id": "double-factorization-reconstruction", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-reconstruction.md", + "symbols": [ + "cudaq_algorithms.double_factorization.reconstruct_eri" + ], + "behavioral_evals": [ + "double-factorization-encoding-boundary", + "chemistry-end-to-end-composition" + ] + }, + { + "id": "double-factorization-error", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-error.md", + "symbols": [ + "cudaq_algorithms.double_factorization.factorization_error" + ], + "behavioral_evals": [ + "double-factorization-encoding-boundary", + "chemistry-end-to-end-composition" + ] + }, + { + "id": "double-factorization-modified-one-body", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-modified-one-body.md", + "symbols": [ + "cudaq_algorithms.double_factorization.modified_one_body_integrals" + ], + "behavioral_evals": [ + "double-factorization-host-contracts" + ] + }, + { + "id": "double-factorization-one-norm", + "family": "double-factorization", + "record": "references/double-factorization/double-factorization-one-norm.md", + "symbols": [ + "cudaq_algorithms.double_factorization.double_factorization_one_norm" + ], + "behavioral_evals": [ + "double-factorization-host-contracts" + ] + }, + { + "id": "simulation-good-subspace", + "family": "simulation", + "record": "references/simulation/simulation-good-subspace.md", + "symbols": [ + "cudaq_algorithms.sim_utils.good_subspace" + ], + "behavioral_evals": [ + "simulation-analysis-hardware-boundary", + "implicit-ftqc-application-advisory" + ] + }, + { + "id": "simulation-action", + "family": "simulation", + "record": "references/simulation/simulation-action.md", + "symbols": [ + "cudaq_algorithms.sim_utils.action" + ], + "behavioral_evals": [ + "block-encoding-capability-boundary", + "simulation-analysis-hardware-boundary" + ] + }, + { + "id": "simulation-transform", + "family": "simulation", + "record": "references/simulation/simulation-transform.md", + "symbols": [ + "cudaq_algorithms.sim_utils.transform" + ], + "behavioral_evals": [ + "qsvt-phase-and-recovery-boundary", + "simulation-analysis-hardware-boundary", + "qsvt-paraphrase-convention" + ] + }, + { + "id": "simulation-evolve", + "family": "simulation", + "record": "references/simulation/simulation-evolve.md", + "symbols": [ + "cudaq_algorithms.sim_utils.evolve" + ], + "behavioral_evals": [ + "simulation-analysis-hardware-boundary" + ] + } + ], + "support_records": [ + "references/application-composition.md", + "references/block-encoding/block-encoding.md", + "references/catalog.md", + "references/chemistry/chemistry-bridges.md", + "references/conventions.md", + "references/conventions/kernel-boundaries.md", + "references/conventions/numerical-comparison.md", + "references/conventions/ordering.md", + "references/conventions/resource-levels.md", + "references/conventions/spectral-processing.md", + "references/double-factorization/double-factorization.md", + "references/fermion-transforms/fermion-transforms.md", + "references/qsvt/qsvt.md", + "references/qubitization/qubitization.md", + "references/simulation/simulation-analysis.md", + "references/source-provenance.md", + "references/state-preparation/hf-ucc-validation.md", + "references/state-preparation/injection-contract.md", + "references/state-preparation/operator-pools.md", + "references/state-preparation/state-preparation.md", + "references/state-preparation/ucc-parameterization.md", + "references/trotter/trotter.md", + "references/validation.md", + "references/workflow.md" + ] +} diff --git a/skills/cudaq-algorithms-dev/coverage/policy.md b/skills/cudaq-algorithms-dev/coverage/policy.md new file mode 100644 index 0000000..90f5e27 --- /dev/null +++ b/skills/cudaq-algorithms-dev/coverage/policy.md @@ -0,0 +1,66 @@ +# Feature coverage and evidence policy + +This is maintainer bookkeeping. Application agents use +[scientific validation](../../cudaq-algorithms/references/validation.md); +maintainers start at [authoring architecture](../authoring/architecture.md). + +## Live registry + +[`features.json`](features.json) inventories 52 independently selectable +operation/object contracts. Shared selectors, representations, conventions and +workflows are `support_records`, not extra primitives. Register additions and +renames together with their records and routing links. + +- `record`, `family` and `symbols` identify the public contract. +- `behavioral_evals` lists authored cases in the canonical + [`evals/evals.json`](../evals/evals.json). A mapping is not evidence that a + case ran or passed; an empty list means no explicit mapping. +- Historical campaign pointers, execution claims and old metadata indices + were archived with the pre-delivery tree. The live registry has no bundled + run evidence. Zero evidence counts mean none is attached here, not that + scientific checks failed or that historical runs never occurred. + +The checker retains optional campaign/evidence validation for future work. +Only attach a campaign when its referenced manifests and results are present +under the declared support root. Each execution must identify the required +public API, case, arm, repetition count and narrow scientific scope. Keep +campaign artifacts outside the delivered folders and attach them only in a +separate review workspace when using this capability. + +## Evidence claims + +Source inspection, compilation, execution and numerical validation are distinct. +Record the actual source/skill versions, dependencies, target, precision, +commands, independent oracle and tolerances with each result. Matching one +record's bytes does not validate the whole current skill or a newer runtime. +Historical evidence keeps its original scope and does not certify the 62-case +delivery suite. Do not restore stale success claims merely to increase coverage +counts after cleanup. + +For optional historical campaign checks, `scientific_pass` requires the +recorded public and held-out numerical variants and host/device evidence for +every listed repetition. `executed` and `blocked` do not count as numerical +passes. Report scientific correctness separately from agent task success. + +Behavioral comparisons require matched baseline and with-skill arms; scientific +claims require independent oracles or appropriate repository tests. Follow the +[delivery evaluation protocol](../evals/EVAL.md). No bundled model results are +implied by a successful static check. + +## Deterministic checks + +From the repository root: + +```bash +python3 -B skills/cudaq-algorithms-dev/scripts/check_coverage.py +python3 -B -m unittest discover -s skills/cudaq-algorithms-dev/scripts/tests +``` + +Use `--json` for diagnostics or `--feature FEATURE_ID` for one contract. Default +roots are sibling `cudaq-algorithms` and `cudaq-algorithms-dev` directories; +pass `--root` and `--support-root` for another split layout. An explicit +`--root` alone retains combined-layout compatibility. + +Checks cover inventory, local file links, runtime reachability, selectable +family-table entries, authored IDs and any explicitly attached run evidence. +They do not rerun science, grade model answers or validate Markdown anchors. diff --git a/skills/cudaq-algorithms-dev/evals/EVAL.md b/skills/cudaq-algorithms-dev/evals/EVAL.md new file mode 100644 index 0000000..06ae7be --- /dev/null +++ b/skills/cudaq-algorithms-dev/evals/EVAL.md @@ -0,0 +1,270 @@ +# Delivery evaluation + +[`evals.json`](evals.json) is the single maintained suite: **62 unchanged cases**, +comprising 42 regression and 20 science cases. It has 59 positive skill cases +and 3 negative controls. [`manifest.json`](manifest.json) pins the original +case order, complete case records, source provenance and fixture hashes. +Intentional future changes require explicit manifest review and rebaselining. +Old 42- or 7-case measurements are not results for this delivery suite. + +The runtime skill also ships a standalone [catalog evaluation package](../../cudaq-algorithms/evals/EVAL.md) +for NVIDIA SkillEvaluator. Its 62 cases and three negative controls decode to +exactly this suite, with byte-identical local fixtures; `test_catalog_evals.py` +guards against drift. The catalog config contains only supported SkillEvaluator +keys. Keep this directory's campaign, grading, and reporting infrastructure here. + +## Paired campaign protocol + +The maintained [multi-model runner](runner/README.md) provides endpoint preflight, +frozen campaigns, isolated execution, private grading and canonical export through +`scripts/run_eval.py`. Its first backend supports OpenAI-compatible Chat +Completions endpoints. Offline CI validates the infrastructure and all 24 +registered executable checks. Live endpoint health and actual model outcomes +must still be established for each campaign; absent verification is explicitly +reported as unassessed. + +Use the canonical suite, [`config.yml`](config.yml), declared `files/` fixtures, +and runtime skill [`../../cudaq-algorithms`](../../cudaq-algorithms). The +runner must execute every case with and without the skill, for **exactly five +independent paired repetitions**: 62 × 5 × 2 = 620 records per model. Include +records for failures, unanswered runs and ungraded runs. Never stop on pass, +retry until a pass, or select the best attempt. A seed identifies a pair; it +does not promise deterministic model-provider behavior. + +The harness configuration uses `n_attempts: 1`, `stop_on_pass: false`, and the +unchanged `pass_threshold: 0.50`. The separate `evaluation_protocol` section is +reporting metadata, **not an implicit Harbor seed-loop feature**. The runner +must loop `[0, 1, 2, 3, 4]` itself, select budgets before running, and copy the +actual protocol into the results. Do not pass custom reporting metadata to a +harness unless its configuration adapter supports it. + +A full campaign includes a frontier model and a mid-tier model. A single-model +bundle is valid and visibly reports missing tier coverage. Keep each model's +results separate. Use isolated throwaway worktrees for implementation tasks. +Both arms receive the same prompts, scientific inputs, repository/library +access, dependencies, tool limits and time budgets; only skill availability +changes. Keep assertions, expected outputs and other answer material out of +worker inputs. Record whether the skill was merely `listed` or explicitly +`injected`; observed opening after injection is not natural trigger recall. + +Before execution, register one `case_contracts` entry for each canonical ID: +`executable_check`, `implementation`, and a nonempty `rationale` describing the +check and applicability. Use identical contracts across arms and models. +Implementation cases must have an executable check. Other checks apply only +where the authored task warrants them; advice and predictions do not acquire +an extra universal simulator or hardware requirement. In science cases, grade +substantive CUDA-Q Algorithms use against the original rubric; an unused +import is insufficient, and permitted independent classical references remain +valid. + +The maintained runner appends identical public artifact-format instructions to +both arms' system messages, preserving each user question verbatim. It freezes +all numerical checker specifications in the campaign and fingerprints their +source. The `after_final_answer_v1` policy performs one private check after the +final answer, within the remaining task budget, on qpp-cpu/fp64. Completion costs +include that observed check; preceding independent failed checks are zero when +this first check succeeds. Worker self-tests remain transcript evidence, not +trusted completion events. Numerical verification and original rubric assessment +of library use, method and interpretation are separate requirements. + +For applicable scientific checks, preregister the oracle, inputs, tolerance, +target and precision. Check intermediate stages and pin the full statevector +when needed to distinguish register order, phase, normalization or partial +state errors. Follow the requested scientific quantity and the source-grounded +[validation guidance](../../cudaq-algorithms/references/validation.md). Preserve +full transcripts, executable commands, artifacts, numerical outputs and check +logs, including failed attempts. Classify the five tripwires from transcript +evidence: controlled measurement, sample feedback, kernel definitions in a +heredoc, unnecessary kernel reminting in loops, and misuse of partial +statevectors. Assess their meaning against the task and inspected source; +do not count an intentional demonstration or a valid requested operation as +a failure merely because a keyword appears. + +The three existing negative cases are the control subgroup. They are not +necessarily unrelated maintenance controls from a broader evaluation guide; +do not change the 62 cases or claim coverage they do not provide. + +## Result bundle and evidence + +[`results.schema.json`](results.schema.json) defines the required JSON shape +using JSON Schema draft 2020-12. All record fields are required and fixed +objects reject extra fields. `null` records missing evidence; it is not a +passing grade, a zero count, or free resource usage. + +Top-level fields are `schema_version: 1`, `suite_sha256` (SHA-256 of the raw +`evals.json` bytes), `source_revision`, `skill_revision`, `skill_label`, +`environment`, `protocol`, `case_contracts`, and `models`. Store immutable +source/skill revisions and record CUDA-Q, CUDA-Q Algorithms, Python, +dependencies, simulator and precision as string-valued environment entries. +The protocol requires five distinct integer `seeds`, positive `budget_seconds`, +a positive integer or null `tool_call_limit`, `skill_exposure`, and nullable +`time_regression_limit`/`token_regression_limit` ratios. Each stated tolerance +must be at least 1; null means the regression target was not assessed. + +Each model has a unique `id`, a `tier` of `frontier` or `mid`, independently +measured `wall_seconds` for each arm, `notes` for regression/science/controls, +and `runs`. Notes should identify task IDs and evidence showing where the +skill changed the work or misled the model. Each run contains: + +- `case_id`, `seed`, `arm` (`baseline` or `skill`), and `outcome`: `answered`, + `budget_timeout`, `tool_limit`, `backend_error`, `error`, or `no_answer`. +- `skill_opened` (baseline must be null), one boolean/null per original + `assertions` entry in order, and boolean/null `expected_output` and + `claimed_success`. Unanswered outcomes cannot have true assertions, + expected output or claimed success; answered runs may still be ungraded. +- `verification`: `passed`, `failed`, `not_run`, or `not_applicable`. + `not_applicable` is mandatory exactly when the registered case has no + executable check. `verification_evidence` is a relative path to an existing, + nonempty check-log file inside the bundle directory for passed/failed checks + and null otherwise. +- `transcript`: a relative path to an existing, nonempty file inside the bundle + directory. `grading_evidence` follows the same rule and is required when + any assertion, expected output, critical failure, tripwire or executable + check has been assessed. These paths must still resolve inside the bundle + after resolving symlinks. Save full transcripts and grading/check logs; + references may share a file when it contains the relevant evidence. +- Nullable nonnegative `critical_failures` and the five `tripwires` counts: + `controlled_measurement`, `sample_feedback`, `heredoc_kernel`, + `remint_in_loop`, `partial_statevector`. Zero requires an actual transcript + assessment; null means unknown. +- `resources`: nonnegative `task_seconds`, `wall_seconds`, + `backend_wait_seconds`, nullable nonnegative integer `tokens`/`tool_calls`, + and nullable nonnegative measured `cost_usd`. Run wall time equals task + time plus backend waiting within floating tolerance. Record total token + usage consistently, including skill reading and unsuccessful work. +- `verified_completion`: present as an object exactly when verification + passed, otherwise null. Its `task_seconds`, `wall_seconds`, nullable + `tokens`/`tool_calls`, and integer `failed_attempts` describe cumulative + work through the first successful executable check, including earlier + failed attempts within that run. Completion costs cannot exceed known + corresponding run totals; completion wall time cannot be below task time. + Waiting accumulated before completion cannot exceed the entire run's wait. + +The reporter rejects duplicate JSON keys, nonfinite numbers, wrong suite +hashes, missing/extra/duplicate case-seed-arm records, wrong assertion counts, +unsafe, absent or empty evidence files, and inconsistent evidence/resource fields. +Each arm's measured campaign wall time must cover its longest run. Campaign +wall time is not the sum of task times; concurrent tasks and backend waits can +overlap. Schema and reporter checks enforce recording and reporting contracts; +they do not establish scientific truth or execute the campaign. + +## Generate the report + +Run from the repository root, after collecting real evidence: + +```bash +python3 -m pip install jsonschema +python3 skills/cudaq-algorithms-dev/scripts/report_eval.py /path/to/campaign/results.json --validate-only +python3 skills/cudaq-algorithms-dev/scripts/report_eval.py /path/to/campaign/results.json --output /tmp/cudaq-eval-report +``` + +Invalid input exits with status 2 before creating output. The output directory +must be absent or empty. Keep both the bundle and generated reports outside +the runtime and development skill trees. The reporter does not run models +or invent missing measurements. Its tests use +synthetic records only. + +An evaluation handoff is incomplete until this command succeeds and its +generated reports, details and source evidence are supplied. A valid report +may contain failed or unassessed targets; validation is not a claim that +delivery targets were met. Regenerate the Markdown from the JSON instead of +editing calculated cells by hand. + +The output `index.md` links one `model-01/`, `model-02/`, … directory per model. +Each `report.md` contains **exactly two Markdown tables** for the full suite: + +1. Quality and target results: `Metric | Baseline (no skill) | With skill + (