Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 55 additions & 0 deletions .github/workflows/skill_evaluation.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
name: Skill evaluation offline checks

on:
pull_request:
paths:
- 'skills/cudaq-algorithms-dev/**'
- 'skills/cudaq-algorithms/**'
- '.github/workflows/skill_evaluation.yaml'
workflow_dispatch:
workflow_call:

permissions:
contents: read

jobs:
offline:
name: Suite, runner and reporting contracts
runs-on: ubuntu-latest

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this should use the Nvidia CPU runners.

timeout-minutes: 15
env:
PYTHONDONTWRITEBYTECODE: '1'
# The CI job never opts into namespace, scientific-runtime or model probes.
CUDAQ_RUNNER_TEST_ISOLATION: '0'
CUDAQ_RUNNER_TEST_PYTHON: ''
steps:
- name: Checkout repository
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.12'

- name: Install test dependencies
run: python -m pip install pytest jsonschema PyYAML numpy scipy pyscf openfermion

- name: Validate canonical suite and private grading
if: ${{ !cancelled() }}
shell: bash
run: python -B -m unittest discover -s skills/cudaq-algorithms-dev/evals/tests

- name: Test routing, coverage and reporting
if: ${{ !cancelled() }}
shell: bash
run: python -B -m pytest -q -p no:cacheprovider skills/cudaq-algorithms-dev/scripts/tests

- name: Test runner with offline transports and trusted test commands
if: ${{ !cancelled() }}
shell: bash
run: python -B -m pytest -q -p no:cacheprovider skills/cudaq-algorithms-dev/evals/runner/tests

- name: Check skill coverage and links
if: ${{ !cancelled() }}
shell: bash
run: python -B skills/cudaq-algorithms-dev/scripts/check_coverage.py
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,6 @@ __pycache__/
compile_commands.json
timer.dat
docs/sphinx/_build/

# Skill-development run evidence stays out of the repository
skills/cudaq-algorithms-dev/evals/results/
52 changes: 52 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/architecture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# Architecture

## Design goal

Represent CUDA-Q Algorithms as a durable library of independently selectable
scientific contracts. Applications consume and validate those contracts; they
do not define the taxonomy.

This file is maintainer and reviewer policy. Scientific tasks normally begin at
[the catalog](../../cudaq-algorithms/references/catalog.md), not here.

## Organization

- `../../cudaq-algorithms/SKILL.md` is the application entry point.
- `../../cudaq-algorithms/references/catalog.md` selects nine scientific families; detailed
operation/object rows live in their family selectors.
- `../../cudaq-algorithms/references/<family>/<record>.md` holds one independently selectable
primitive per focused file. A front door is navigation, not a multi-contract record.
- Shared representations and capabilities live in a focused support record
or family front door when multiple producers/consumers exchange them.
- `../../cudaq-algorithms/references/conventions.md` selects five convention records;
`validation.md` and `source-provenance.md` own validation and current-source lookup.
- `../../cudaq-algorithms/references/application-composition.md` describes application chains.
- `../../cudaq-algorithms/references/workflow.md` shares advisory and implementation guidance.
- This `authoring/` directory holds schemas, maintenance policy, and open design decisions.
- `../coverage/` holds the live contract inventory and per-feature evaluation mappings;
`../scripts/check_coverage.py` checks consistency.

This development directory is not an application skill and has no `SKILL.md`.
The canonical delivery suite, evaluator configuration, required fixtures and
integrity checks live under `../evals/`; do not package them as application
guidance. Historical campaigns and superseded harnesses are archived outside
these delivery directories. Their results do not validate the current suite.

Keep scientific family paths shallow, with no primitive subdirectories. Every
focused record must be linked from its family selector or directly from the root
catalog. Each record covers every applicable canonical schema field once;
short records may combine adjacent headings when boundaries remain explicit.

## Authoring routes

- [Record design](record-design.md): identity, granularity, metadata, composition, resources.
- [Capability design](capability-design.md): semantic boundaries and identifiers.
- [Source review](source-review.md): source authority and freshness.
- [Extension workflow](extension-workflow.md): lifecycle and coordinated changes.
- [Design decisions](design-decisions.md): open ownership and promotion questions.
- Templates: [primitive](templates/primitive-record-template.md),
[representation](templates/representation-record-template.md),
[capability](templates/capability-record-template.md),
[convention](templates/convention-record-template.md).
- [Coverage policy](../coverage/policy.md), [feature registry](../coverage/features.json),
and [delivery evaluation](../evals/EVAL.md).
23 changes: 23 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/capability-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
# Capability Design

## Capability composition

Use identifiers of the form
`cudaq-algorithms.<dotted-capability-name>.v<major>`. Dotted capability names
are valid. The major version changes only for an incompatible semantic change.

The currently adopted documentation identifiers are:

- `cudaq-algorithms.state-preparation.unitary.v1`;
- `cudaq-algorithms.block-encoding.zero-flagged.v1`;
- `cudaq-algorithms.chemistry-integrals.v1`.

Their status remains `provisional`; the identifier is resolved even though the
boundary has not been promoted to a stable taxonomy contract.

Every capability record states its ID, status, owner, direction, boundary
representation, exact signature, semantic invariants, geometry, conventions,
execution boundary, providers, consumers, and unsupported conditions.
Composition requires the same ID and compatible major version, plus every
consumer invariant. Similar names and structural member presence are
insufficient.
46 changes: 46 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/design-decisions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Design decisions and open questions

## State-preparation ownership and capability promotion

- Register or shape geometry and ownership: geometry is invariant 2 in the
[injection contract](../../cudaq-algorithms/references/state-preparation/injection-contract.md#capability-record--unitary-state-preparation).
At the
historical last review, ownership was **caller-owned in the reviewed source
only**: the consumer factory that called the preparation kernel allocated the
register, handed it over fresh in `|0...0>`, and expected exactly its own system
width. That is a `derived` historical description, **explicitly not a universal
or future policy.** Check current public source at use time. Whether the
caller or the primitive should own system and ancilla registers in general is
an **open** question in the
taxonomy design record outside this package (the
[kernel-boundary convention](../../cudaq-algorithms/references/conventions/kernel-boundaries.md)
repeats the scope limit); do not present the reviewed behavior as a library-wide guarantee,
nor the open question as settled.

- Promotion criteria to a public protocol or compiler IR operation, and the
explicit decision still required: (a) a provider outside the two historically
reviewed providers that exercises the same boundary without widening it, (b)
a resolved register-ownership policy, (c) characterized behavior for the conditions
marked `unverified` in the
[injection boundaries](../../cudaq-algorithms/references/state-preparation/injection-contract.md#shared-unsupported-and-unverified-boundaries),
and (d) an explicit team decision recorded with an
owner. None of the four held at the historical last review; check current
public source and team records at use time. This record proposes no promotion.

- **Scope limit.** This is a `derived` description of the injection seam found
during the last source review. Recheck the cited call sites in the current
checkout before relying on it. It is **not** a decided capability-level policy.
Whether the caller or the primitive should own system and ancilla registers
in general, and the exact input-state, width, ancilla, inverse, control, and
failure semantics of state preparation, remain **open** questions in the
taxonomy design record, which lives outside this skill package. Do not
present the current behavior as a library-wide guarantee for future
capabilities, and do not present the open questions as settled.

## Capability status

The adopted state-preparation, zero-flagged block-encoding, and chemistry-integral
documentation capability IDs remain provisional; they have not been promoted
to stable taxonomy contracts. The injection boundary has two packaged providers
and four independent consumer modules, but is not a public protocol, and stability
across future providers is not established.
62 changes: 62 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/extension-workflow.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# Extension Workflow

Only add records for QROM, arithmetic, sparse-oracle, eigensolver, THC, or
additional resource-model concepts once current public source establishes a
contract. Roadmap names alone do not establish available primitives.

## Lifecycle and evidence

The following preserves the historical review vocabulary and promotion bar.
Store lifecycle history and executed evidence under
[coverage](../coverage/policy.md), not as status comments in operational
records. Current coverage uses scoped evidence, not blanket lifecycle stamps.

`draft | verified | deprecated | removed` is the only record lifecycle
vocabulary. A capability separately uses
`candidate | provisional | stable taxonomy contract`.

A record becomes `verified` only after its contract, runnable usage, relevant
scientific tests, and representative evals have actually passed on a recorded,
supported package/CUDA-Q version combination. Store exact source revisions,
dependencies, targets, and commands with that result. Source inspection alone
supports `source-checked` in the task that performs it, not `verified`.

Deprecation records name the replacement, first deprecated version, and
behavioral differences. Incompatible contract changes prefer a new primitive
name and explicit migration over silent redefinition.

## Growth rule

Add knowledge in this order:

1. shared representation or convention when demonstrated;
2. one independently selectable primitive contract;
3. a capability only when multiple producers or consumers justify it;
4. a composite protocol when its lower-level contracts are populated;
5. catalog routing, runnable usage, validation, and eval coverage in the same
change.

Do not add records for roadmap concepts or speculative APIs. Split existing
records when real contracts have become independently selectable.

## Incremental development rule

Add or refine one independently selectable contract at a time. Update its
catalog entry, source provenance, complete contract, resource status,
independent validation method, runnable example or usage test, and evaluation
coverage together. Split a record whenever operation/object identity, return
type, execution layer, validation oracle, approximation behavior, resource
contract, or composition boundary can be selected independently.

`verified` is a record lifecycle state, not shorthand for reading source or for
one successful run. Promotion requires a recorded, supported package/CUDA-Q
version combination; exact revisions, dependencies, targets, and commands
belong with the validation or evaluation result.

## Skill evaluation versus scientific validation

SkillEvaluator checks activation, routing, usefulness, safety, and answer
quality. It does not establish that a quantum circuit or numerical transform is
scientifically correct. Run baseline and with-skill eval arms for behavioral
uplift, and run repository tests or independent numerical oracles separately
for scientific claims.
104 changes: 104 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/record-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
# Record Design

## Primary identity and granularity

The routing identity is:

```text
scientific operation + mathematical object
```

Operations include prepare, load, encode, transform, evolve, measure, estimate,
synthesize, preprocess, analyze, and reconstruct. Objects include states,
fermionic operators, Pauli operators, integral tensors, block encodings,
polynomials, phase sequences, and resource descriptions.

One primitive record corresponds to one contract a caller can select
independently. Split records when any of these differ materially:

- operation or mathematical object;
- public entry point and input representation;
- return type or emitted kernel signature;
- execution layer or authorization implications;
- validation/rejection behavior or independent oracle;
- approximation/error behavior;
- resource contract;
- required/provided capability or composition boundary.

Several symbols may remain in one record when they are inseparable parts of one
contract. One source module or class may require several records. File size is
evidence of a possible granularity problem, never the routing rule itself.

## Record types

1. **Primitive record:** one concrete, independently selectable operation.
2. **Representation record:** the meaning of an exchanged object; create only
after multiple producers and consumers interpret the same form.
3. **Capability record:** a reusable semantic composition boundary; create
only after multiple independent producers or consumers demonstrate it.

These are documentation records, not automatic requests for a Python protocol,
ABC, compiler IR operation, or new public API.

## Orthogonal metadata

Classify, but do not route or organize directories, by:

- **Kind:** quantum operation, classical transformation,
measurement/readout, simulation-only analysis, resource estimator.
- **Routine role:** driver, computational, auxiliary. Role describes problem
completeness, not where code executes.
- **Execution layer:** host preprocessing, kernel factory, device kernel,
observable/measurement, simulation-only host path, or mixed.
- **Abstraction:** leaf operation or composite protocol.
- **Parameterization:** none, construction-time, runtime.
- **Representation and capabilities.**
- **Exactness, uncertainty, and method.**
- **Domain, dependencies, error contract, resource contract, lifecycle.**

Lifecycle is historical coverage metadata; the other scientific classifications
remain part of the live contract. Follow [coverage policy](../coverage/policy.md)
for that placement distinction.

Host transforms such as chemistry loaders and factorizations are computational
routines when they solve an independently useful problem. A simulation-only
helper is not a hardware primitive merely because it consumes one.

## Composite protocols

A reusable driver may itself be a primitive. Its record must state required
lower-level capabilities, a source-grounded reference composition, applicability
conditions, alternatives, propagated conventions/errors/resources, and what an
alternative component must preserve. Never silently replace the reference
composition with a target-specific heuristic.

## Resource claims

Every executable primitive either gives a resource contract or explicitly says
that none exists. Every quantity identifies metric/unit, abstraction level,
architecture/execution assumptions, exact/bounded/estimated/measured status,
controlling parameters, confidence/limitations, and composition rule if known.

Never compare logical operations, decomposition proxies, transpiled gates,
runtime, memory, or measured hardware cost as if they were one metric. Never
turn a benchmark or source comment into a fresh measurement.

## HF/UCC record application

The [HF/UCC record](../../cudaq-algorithms/references/state-preparation/state-preparation-hf-ucc.md)
is a **concrete primitive record**: it instantiates each applicable
canonical Primitive-Record field from `templates/primitive-record-template.md`
exactly once, for one contract — a Hartree-Fock reference occupation optionally
followed by a UCC product at amplitudes the caller already knows. The shared
seam it plugs into (kernel representation, unitary capability, consumer table,
common boundaries) is [injection contract](../../cudaq-algorithms/references/state-preparation/injection-contract.md);
cross-cutting layout, ownership, and validation conventions are
[the convention selector](../../cudaq-algorithms/references/conventions.md). Shared scientific
detail is linked rather than repeated; lifecycle and run bookkeeping follows
the template's separate coverage destination.

## Provisional HF/UCC classification discussion

The HF/UCC method is a fixed-parameter ansatz product. `ansatz` was a provisional
extension beyond the `direct | variational | heuristic` vocabulary. Exactness,
uncertainty, and method are recorded values, never routing identities.
44 changes: 44 additions & 0 deletions skills/cudaq-algorithms-dev/authoring/source-review.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Source Review

## Source ownership and freshness

Current public code and authoritative tests in the checked-out repository
control API behavior. [Source lookup](../../cudaq-algorithms/references/source-provenance.md)
records common source/test/example locations; each record adds only
contract-specific stable symbols and paths. Follow the
[coverage policy](../coverage/policy.md) for evidence claims. Archived reviews
are audit material, not current contracts or compatibility promises.

When a selected record differs from current source or tests:

1. compare the relevant public symbol and tests;
2. treat current public source as authoritative for generated code;
3. report the drift and evidence level;
4. update a maintained record only when that update is in scope;
5. never retain a stale line-number claim merely because the prose is familiar.

Do not copy historical repository hashes or dependency pins into primitive
records. Historical last-review values belong with the archived review. A validation or evaluation result instead records the exact revision and
dependencies actually used by that run.

Use line numbers only for a non-obvious invariant that benefits from a precise
anchor. Prefer stable symbol and test names for ordinary provenance.

## Freshness check

Before implementing from a maintained record:

```bash
git rev-parse HEAD
git status --short -- \
python/cudaq_algorithms tests/python docs/sphinx \
pyproject.toml .cudaq_version
```

Inspect only the selected symbols and their tests. If public signatures or
scientific assertions changed, follow current source for generated code, report
the drift, and update the maintained record only when that work is in scope. Do
not silently upgrade the lifecycle from `draft` or claim a newly verified
version range. If comparing against a previous review, obtain its actual
recorded revision from that review; do not assume a fixed historical hash is
the baseline for today's checkout.
Loading
Loading