Skip to content

Add CUDA-Q Algorithms agent skill - #46

Open
kvmto wants to merge 7 commits into
NVIDIA:mainfrom
kvmto:algo-skill
Open

kvmto wants to merge 7 commits into
NVIDIA:mainfrom
kvmto:algo-skill

Conversation

@kvmto

@kvmto kvmto commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a CUDA-Q Algorithms agent skill that translates researchers’ inputs and requested scientific outputs into library operations, composed workflows, and numerical verification. For example, a molecular-energy or lattice-dynamics request routes to the relevant scientific guide, API contracts, and current source without requiring the researcher to name repository files.

The runtime skill lives in skills/cudaq-algorithms; authoring, coverage, evaluation, and reporting tools live in skills/cudaq-algorithms-dev.

Runtime skill

  • Adds focused routing, a reference catalog, API contracts, and source/test pointers for state preparation, block encoding and Pauli LCU, qubitization, QSP/QSVT, Suzuki–Trotter evolution, chemistry, fermion transforms, double factorization, and simulation analysis.
  • Adds application-composition guidance and domain-specific Workflow and Verification sections, including molecular and condensed-matter use cases.
  • Documents register ordering, normalization, signs and phases, energy offsets, symmetry sectors, and host/kernel boundaries. Verification guidance distinguishes independent numerical checks, executed circuits, and resource estimates.

Evaluation infrastructure

  • Maintains one manifest-pinned 62-case suite: 42 regression cases and 20 science cases, including three negative controls. The manifest validates complete case records, ordering, provenance, and fixtures.
  • Defines five paired baseline/skill repetitions: 620 attempts per model. Both arms receive the same questions, inputs, library access, and budgets; skill exposure is recorded explicitly.
  • Adds an OpenAI-compatible tool-loop runner and ten NVIDIA Build model configurations, with completion/tool-call preflight, frozen campaign inputs, isolated workspaces, resumable pending work, and bounded transport retries. Provider errors and measured backend waiting are retained.
  • Adds private, evidence-cited rubric assessment and 24 executable checks: all 20 science cases plus four implementation regression cases. Scientific checks compare runnable artifacts against independent numerical references.
  • Adds monitoring, a strict result schema, evidence/hash validation, and two generated tables per model: quality against targets, and effort/completion measurements. Reports retain failures and unknown measurements, with paired and subgroup details in JSON.

The committed runner supports OpenAI-compatible endpoints; native Codex/Claude adapters are outside this PR. Rubric scores, executable correctness, and resource measurements are reported separately. Full multi-model performance results are still being collected.

Validation

Verified on PR head ee5710be:

The evaluation workflow runs on relevant pull requests and supports manual and reusable invocation. It requires no model credentials; live model campaigns use the documented runner and external evidence directories.

See the evaluation protocol and runner usage.

Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
@wsttiger wsttiger added the enhancement New feature or request label Sep 17, 2026
kvmto and others added 2 commits September 23, 2026 18:54
Operational skill: SKILL.md, 76 reference records and the router, as measured on the 42-evaluation boundary suite with Codex and Claude Code. Development companion: authoring method, coverage registry, evaluation definitions and tooling, without run evidence (ignored via .gitignore), machine-bound experiment controllers or code depending on packages that are not public. Machine paths replaced by environment variables, SPDX headers added, yapf applied.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
@kvmto
kvmto marked this pull request as ready for review October 5, 2026 08:58
kvmto added 2 commits October 5, 2026 17:00
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
jobs:
offline:
name: Suite, runner and reporting contracts
runs-on: ubuntu-latest

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this should use the Nvidia CPU runners.

@anjbur anjbur left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A couple things from the structural side:

  • The cudaq-algorithms skill will need its own evals, rather than using ones inside cudaq-algorithms-dev. This should also include negative checks for when the skill won't be triggered.
  • Any non-ASCII characters should be removed. The description line and one table row in cudaq-algorithms/SKILL.md use em and en dashes, and there are also many reference files with non-ASCII characters
  • The cudaq-algorithms skill needs a skill-card.md file. The evaluation details will come later, from the nvskills-ci run, but the broader details can be added. Example

In general we'll want SkillEvaluator to pass locally before merge, as that'll be a requirement for this to enter the broader skills catalog.

Comment thread skills/cudaq-algorithms/SKILL.md Outdated
@@ -0,0 +1,164 @@
---
name: cudaq-algorithms
description: Use when designing, implementing, debugging, reviewing, validating, or composing cudaq_algorithms / CUDA-Q Algorithms APIs and fault-tolerant primitives, including scientific requests expressed as molecular geometry/integrals to energies, lattice Hamiltonians to dynamics, or prepared states to spectral and conditional observables. Covers state preparation, Pauli/block encoding, qubitization, QSP/QSVT, Suzuki–Trotter evolution, fermion transforms, chemistry, double factorization, and statevector analysis. Not for CUDA-Q installation, backend setup, basic standalone kernels such as Bell states, or unrelated quantum-computing questions.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This description line should be shortened to 50-150 characters. The full trigger list can go in the body of the skill

version: "0.2.0"
---

# CUDA-Q Algorithms

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It should be explicitly stated somewhere in this file that the skill relies on a checkout of the algorithms repo.

---

# CUDA-Q Algorithms

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'd be great to add a scripts table somewhere in this file with details about route.py. Something like this.

---
name: cudaq-algorithms
description: Use when designing, implementing, debugging, reviewing, validating, or composing cudaq_algorithms / CUDA-Q Algorithms APIs and fault-tolerant primitives, including scientific requests expressed as molecular geometry/integrals to energies, lattice Hamiltonians to dynamics, or prepared states to spectral and conditional observables. Covers state preparation, Pauli/block encoding, qubitization, QSP/QSVT, Suzuki–Trotter evolution, fermion transforms, chemistry, double factorization, and statevector analysis. Not for CUDA-Q installation, backend setup, basic standalone kernels such as Bell states, or unrelated quantum-computing questions.
license: Apache-2.0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'd be great to add tags and metadata.tags lists to this header as well

Comment thread .github/workflows/skill_evaluation.yaml Outdated
set -o pipefail
python -B skills/cudaq-algorithms-dev/scripts/check_coverage.py 2>&1 | tee "$RUNNER_TEMP/cudaq-skill-evaluation/coverage.log"

- name: Upload failure diagnostics

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This step can be removed, along with the directory prep step and the | tee "$RUNNER_TEMP/..." pipe on each run step. All of that output is being written to the CI logs with or without tee, so the job log already contains everything that this artifact would.

@kvmto

kvmto commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Hi Angela, could you take a look at the CI and evaluation side of this PR when you have a chance?

I’d start with the workflow, EVAL.md, and the runner README. I’d especially appreciate your thoughts on running this reliably in CI, handling endpoint failures and retries, keeping credentials and execution isolated, and making sure the reports reflect what actually happened.

Here are some files to help you find your way around:

.github/workflows/skill_evaluation.yaml

skills/cudaq-algorithms-dev/evals/
  EVAL.md
  config.yml
  manifest.json
  models.json
  results.schema.json
  runner/README.md
  runner/campaign.py
  runner/cli.py
  runner/providers.py
  runner/workspace.py
  runner/verification.py
  runner/assessment.py

skills/cudaq-algorithms-dev/scripts/
  run_eval.py
  monitor_eval.py
  report_eval.py
  check_coverage.py

skills/cudaq-algorithms-dev/coverage/
  features.json
  policy.md

And the related tests, if useful:

skills/cudaq-algorithms-dev/evals/tests/
  test_active_evals.py
  test_assessment.py

skills/cudaq-algorithms-dev/evals/runner/tests/
  test_campaign.py
  test_cli.py
  test_providers.py
  test_transport_budget.py
  test_transport_recovery.py
  test_workspace.py
  test_verification.py

skills/cudaq-algorithms-dev/scripts/tests/
  test_monitor_eval.py
  test_report_eval.py
  test_check_coverage.py

Please feel free to follow anything else that catches your attention—the list is just a starting point. Whatever time you can spare would be a big help, and if you have time for a deeper look, all the better.

Thanks a lot for helping with this. I know reviewing takes time, and I really appreciate it.

@kvmto

kvmto commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Hi Scott, could you take a look at the skill and the scientific workflows when you have a chance?

I’d really like your take on whether this matches what you had in mind for researchers: choosing the right library pieces, getting the physics and conventions right, and checking the results properly.

I’d start with SKILL.md and the application-composition guide, then try following the block-encoding, chemistry, and QSVT examples through to their evaluation checks.

Here are the relevant files:

skills/cudaq-algorithms/
  SKILL.md
  scripts/route.py

  references/catalog.md
  references/source-provenance.md
  references/workflow.md
  references/application-composition.md
  references/validation.md

  references/conventions/kernel-boundaries.md
  references/conventions/numerical-comparison.md
  references/conventions/ordering.md
  references/conventions/resource-levels.md
  references/conventions/spectral-processing.md

  references/block-encoding/block-encoding.md
  references/block-encoding/pauli-lcu.md

  references/chemistry/chemistry-bridges.md
  references/chemistry/chemistry-from-pyscf.md
  references/chemistry/chemistry-qubit-hamiltonian.md
  references/chemistry/chemistry-spin-orbital-tensors.md

  references/qsvt/qsvt.md
  references/qsvt/qsvt-sequence.md
  references/qsvt/qsvt-recovery.md

The questions, assertions, and numerical checks are here:

skills/cudaq-algorithms-dev/evals/
  evals.json
  case_contracts.json
  runner/checkers/__init__.py
  runner/checkers/common.py
  runner/checkers/states_walk.py
  runner/checkers/chemistry.py

These six cases would be a useful first sample:

  • science-block-encoding-heralded-observables
  • science-block-encoding-truncation-tradeoff
  • science-chemistry-reference-quality
  • science-chemistry-water-bending
  • science-qsvt-low-energy-filter-choice
  • science-qsvt-calibration-robustness

And the related tests, if useful:

skills/cudaq-algorithms-dev/evals/runner/tests/
  test_checker_common.py
  test_checker_registry.py
  test_checkers_states_walk.py
  test_checkers_chemistry.py

skills/cudaq-algorithms-dev/scripts/tests/
  test_route.py

Please feel free to explore other workflows or question the approach more broadly. These are just suggestions for where to start, and I’d welcome your thoughts wherever you think something could work better.

Thanks again for the ideas you shared and for any time you can spend on this. I’m sure your experience and ideas will help make it better. Thanks a lot!

Signed-off-by: Kevin Mato <kmato@nvidia.com>
@wsttiger

wsttiger commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

@kvmto — nice work on the restructure (the smaller -dev/ and shipping evals with the skill are good changes). One heads-up for a future PR, since the library moved under the skill:

QROM + unary iteration are now merged on main. python/cudaq_algorithms/primitives/ ships QROM and unary_iteration_kernels (UnaryIterationKernels). The skill predates that merge and still treats QROM as absent, which creates one real problem and a couple of gaps:

  • The QROM eval is now wrong. The case asking for the "QROM API, resource count, and benchmark" requires the agent to answer that no QROM API exists / is "absent from the current public cudaq_algorithms source." That's now false — QROM is public source — so the eval would penalize a correct answer (an agent that checks source and reports QROM at primitives/_qrom.py gets graded as inventing an API).
  • No QROM / primitives record, and QROM/unary iteration are absent from catalog.md and source-provenance.md — a coverage gap now that they're shipped primitives.

Suggested follow-up (not blocking this PR):

  1. Add a primitives reference record for QROM / unary iteration (signature, the select vs select_swap/QROAM variants, the strictly-unitary coherent resource counts the module documents).
  2. List it in catalog.md and source-provenance.md (source primitives/_qrom.py, _unary_iteration.py; tests test_primitives_qrom.py, test_primitives_unary_iteration.py).
  3. Re-aim the QROM eval at the benchmark axis (there's still no measured QROM benchmark to report) rather than the existence axis.
  4. Leave arithmetic, alias sampling, and sparse as not-installed for now — they're correct as-is; add them when they merge (arithmetic is in flight on Add reversible integer arithmetic primitives #47, the others are still on feature branches).

Root cause is skill↔library drift: availability claims go stale silently as things merge. A small CI check that cross-validates the skill against the live package surface — for anything the skill (or an eval) marks absent, assert it's still absent in python/cudaq_algorithms; for each documented family, assert its named public symbols still import — would have caught this automatically.

@kvmto

kvmto commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

The latest changes enable automatic skill preloading for Codex and Claude projects and make injected exposure the default for new evaluations.

The initial results are encouraging. Average rubric scores improve with the skill across every tested configuration in this snapshot. With injected guidance, requested-output passes increase from 35/55 to 48/55 for Codex and 12/33 to 30/33 for Claude. Both also show lower median task times, although token usage increases for Codex and decreases for Claude.

How to interpret injection: listed mode makes the public skill available for discovery; injected mode supplies its full instructions upfront. Private grading criteria and expected answers remain withheld. Preloading makes the guidance consistently available to users. The new project setup implements that through project instructions, whose priority differs from the evaluation’s system-prompt injection.

Some rubric items explicitly reward consulting the skill, so rubric improvements partly reflect workflow adherence. Independent executable checks provide complementary evidence and show a more varied picture. The tables therefore report quality, completion, execution checks, time, and tokens separately.

Evaluation of the open-source models through the tested APIs has been extremely slow in this run. Median attempt wall times in this snapshot range from roughly 5–11 minutes, compared with about 35–65 seconds for the native Codex and Claude runs. Substantial measured backend waiting has also slowed collection, so these API results remain partial. These are end-to-end timings for the tested tasks, providers, and configurations, not an isolated measure of model capability.

Snapshot: data frozen 7 October 2026, 15:35:43 UTC, validated at 15:36:04 UTC; one paired repetition across 62 questions. 812/1,240 attempts finished. Codex and Claude execution is complete for this repetition, with grading still partial; the other models have partial execution and grading.

All arrows mean baseline → skill, compared within the same exposure mode.

Codex GPT-5.6 Sol — execution complete; grading partial

Metric Listed Injected
Finished attempts / 62 per arm 62/62 → 62/62 62/62 → 62/62
Answers delivered 61 → 60 61 → 61
Fully assessed attempts 54 → 56 57 → 60
Matched fully assessed pairs 50 55
Mean rubric score / 10 6.77 → 7.60 6.69 → 8.45
Strict rubric passes 13/50 → 22/50 15/55 → 28/55
Requested-output passes 31/50 → 36/50 35/55 → 48/55
Known critical failures 0 → 0 0 → 1
Independent checks: pass/fail/not run 22/1/1 → 21/2/1 22/1/1 → 23/0/1
Matched finished pairs for timing 62 62
Median task time (s) 55.2 → 64.0 55.0 → 50.2
Mean task time (s) 82.5 → 99.4 80.2 → 68.1
Cumulative task time (min) 85.3 → 102.8 82.9 → 70.4
Median wall time (s) 55.2 → 64.0 55.0 → 50.2
Cumulative wall time (min) 85.3 → 102.8 82.9 → 70.4
Cumulative backend waiting (min) 0.0 → 0.0 0.0 → 0.0
Matched pairs with complete token usage 60 61
Median tokens (thousands) 58.2 → 68.6 65.3 → 90.5
Total tokens (millions, complete pairs) 4.49 → 4.80 5.02 → 6.12

Claude Sonnet 5 — execution complete; grading partial

Metric Listed Injected
Finished attempts / 62 per arm 62/62 → 62/62 62/62 → 62/62
Answers delivered 60 → 60 60 → 61
Fully assessed attempts 39 → 46 48 → 44
Matched fully assessed pairs 31 33
Mean rubric score / 10 6.12 → 8.03 5.48 → 8.75
Strict rubric passes 4/31 → 11/31 3/33 → 19/33
Requested-output passes 13/31 → 19/31 12/33 → 30/33
Known critical failures 2 → 1 3 → 1
Independent checks: pass/fail/not run 15/8/1 → 17/6/1 17/6/1 → 20/3/1
Matched finished pairs for timing 62 62
Median task time (s) 64.9 → 47.0 59.4 → 34.9
Mean task time (s) 122.2 → 87.1 125.2 → 58.2
Cumulative task time (min) 126.3 → 90.0 129.4 → 60.2
Median wall time (s) 64.9 → 47.0 59.4 → 34.9
Cumulative wall time (min) 126.3 → 90.0 129.4 → 60.2
Cumulative backend waiting (min) 0.0 → 0.0 0.0 → 0.0
Matched pairs with complete token usage 58 59
Median tokens (thousands) 92.0 → 84.0 67.7 → 63.6
Total tokens (millions, complete pairs) 9.16 → 6.30 8.32 → 4.87

Kimi, low reasoning — partial results

Metric Listed Injected
Finished attempts / 62 per arm 25/62 → 25/62 30/62 → 30/62
Answers delivered 12 → 17 17 → 19
Fully assessed attempts 19 → 16 19 → 20
Matched fully assessed pairs 13 13
Mean rubric score / 10 1.88 → 4.96 2.21 → 4.08
Strict rubric passes 1/13 → 5/13 1/13 → 4/13
Requested-output passes 1/13 → 7/13 2/13 → 5/13
Known critical failures 1 → 0 0 → 0
Independent checks: pass/fail/not run 0/0/0 → 0/0/0 0/0/1 → 0/0/1
Matched finished pairs for timing 25 30
Median task time (s) 557.9 → 425.5 400.8 → 269.1
Mean task time (s) 517.7 → 454.6 448.7 → 408.9
Cumulative task time (min) 215.7 → 189.4 224.4 → 204.4
Median wall time (s) 561.8 → 609.5 460.7 → 294.2
Cumulative wall time (min) 268.5 → 247.1 261.8 → 256.7
Cumulative backend waiting (min) 52.8 → 57.7 37.4 → 52.3
Matched pairs with complete token usage 2 4
Median tokens (thousands) 10.5 → 18.3 5.4 → 7.7
Total tokens (millions, complete pairs) 0.02 → 0.04 0.02 → 0.03

Muse — partial results

Metric Listed Injected
Finished attempts / 62 per arm 25/62 → 26/62 27/62 → 26/62
Answers delivered 23 → 22 19 → 25
Fully assessed attempts 15 → 14 16 → 19
Matched fully assessed pairs 10 12
Mean rubric score / 10 3.89 → 5.72 3.00 → 7.68
Strict rubric passes 0/10 → 1/10 0/12 → 4/12
Requested-output passes 3/10 → 4/10 3/12 → 10/12
Known critical failures 2 → 0 0 → 0
Independent checks: pass/fail/not run 0/0/0 → 0/0/0 0/0/0 → 0/0/0
Matched finished pairs for timing 25 26
Median task time (s) 558.6 → 424.8 478.2 → 310.1
Mean task time (s) 530.4 → 406.8 557.3 → 320.0
Cumulative task time (min) 221.0 → 169.5 241.5 → 138.7
Median wall time (s) 643.5 → 518.2 658.4 → 368.6
Cumulative wall time (min) 292.5 → 237.5 298.0 → 224.4
Cumulative backend waiting (min) 71.5 → 68.1 56.5 → 85.7
Matched pairs with complete token usage 6 9
Median tokens (thousands) 454.9 → 256.7 333.2 → 245.4
Total tokens (millions, complete pairs) 3.82 → 1.70 4.02 → 2.09

Nemotron Super — partial results

Metric Listed Injected
Finished attempts / 62 per arm 25/62 → 25/62 26/62 → 26/62
Answers delivered 22 → 20 23 → 24
Fully assessed attempts 14 → 16 15 → 18
Matched fully assessed pairs 9 10
Mean rubric score / 10 0.60 → 3.57 3.43 → 6.91
Strict rubric passes 0/9 → 1/9 0/10 → 4/10
Requested-output passes 0/9 → 1/9 2/10 → 3/10
Known critical failures 4 → 1 1 → 1
Independent checks: pass/fail/not run 0/0/0 → 0/0/0 0/0/0 → 0/0/0
Matched finished pairs for timing 25 26
Median task time (s) 125.4 → 157.8 138.2 → 86.9
Mean task time (s) 159.4 → 251.9 201.7 → 168.7
Cumulative task time (min) 66.4 → 105.0 87.4 → 73.1
Median wall time (s) 460.0 → 360.1 643.5 → 328.1
Cumulative wall time (min) 241.6 → 294.4 285.8 → 237.1
Cumulative backend waiting (min) 175.2 → 189.4 198.4 → 164.0
Matched pairs with complete token usage 3 4
Median tokens (thousands) 201.9 → 923.6 561.9 → 468.4
Total tokens (millions, complete pairs) 1.16 → 2.08 4.19 → 1.88

Reading the tables: Quality metrics use matched fully assessed pairs. Timing uses matched finished pairs, including unsuccessful attempts; wall time includes measured backend waiting. Cumulative time sums attempt durations rather than measuring elapsed campaign time. Token comparisons use only pairs with complete usage records; missing usage remains unknown. Independent-check counts cover applicable finished attempts; 0/0/0 means no applicable finished attempts, not a correctness pass. Different coverage and incomplete grading prevent reliable cross-model rankings or a clean listed-versus-injected comparison at this stage.

Performance should be interpreted within the tested operating conditions. Model capability, reasoning configuration, task complexity, available tools and execution budget all define those conditions. This study did not isolate the causal effect of reasoning effort. Earlier results also used different configurations and should be treated as separate experiments.

Scientific-agent evaluation is demanding because a satisfactory result must connect several stages: understand the research question, choose an appropriate method, use the requested library, produce an executable artifact and validate its output. The present evaluation makes those stages visible and provides a useful basis for improving them.

These findings support a focused product-validation phase using Codex, Claude, Muse, Nemotron Super and Kimi. The immediate priorities are reliable artifact delivery, precise convention handling and bounded execution for expensive workflows. The 62 questions should remain fixed while targeted paired checks establish whether those changes improve outcomes.

The evidence supports further investment in a scoped pilot. Broader readiness should be established for each supported model and workflow through reproducible correctness and reliability measurements.

@kvmto
kvmto requested review from anjbur and wsttiger October 7, 2026 16:49

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants