Repository navigation
Conversation
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Operational skill: SKILL.md, 76 reference records and the router, as measured on the 42-evaluation boundary suite with Codex and Claude Code. Development companion: authoring method, coverage registry, evaluation definitions and tooling, without run evidence (ignored via .gitignore), machine-bound experiment controllers or code depending on packages that are not public. Machine paths replaced by environment variables, SPDX headers added, yapf applied. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
Signed-off-by: Kevin Mato <kmato@nvidia.com>
| jobs: | ||
| offline: | ||
| name: Suite, runner and reporting contracts | ||
| runs-on: ubuntu-latest |
There was a problem hiding this comment.
I think this should use the Nvidia CPU runners.
anjbur
left a comment
There was a problem hiding this comment.
A couple things from the structural side:
- The cudaq-algorithms skill will need its own evals, rather than using ones inside cudaq-algorithms-dev. This should also include negative checks for when the skill won't be triggered.
- Any non-ASCII characters should be removed. The description line and one table row in cudaq-algorithms/SKILL.md use em and en dashes, and there are also many reference files with non-ASCII characters
- The cudaq-algorithms skill needs a
skill-card.mdfile. The evaluation details will come later, from the nvskills-ci run, but the broader details can be added. Example
In general we'll want SkillEvaluator to pass locally before merge, as that'll be a requirement for this to enter the broader skills catalog.
| @@ -0,0 +1,164 @@ | |||
| --- | |||
| name: cudaq-algorithms | |||
| description: Use when designing, implementing, debugging, reviewing, validating, or composing cudaq_algorithms / CUDA-Q Algorithms APIs and fault-tolerant primitives, including scientific requests expressed as molecular geometry/integrals to energies, lattice Hamiltonians to dynamics, or prepared states to spectral and conditional observables. Covers state preparation, Pauli/block encoding, qubitization, QSP/QSVT, Suzuki–Trotter evolution, fermion transforms, chemistry, double factorization, and statevector analysis. Not for CUDA-Q installation, backend setup, basic standalone kernels such as Bell states, or unrelated quantum-computing questions. | |||
There was a problem hiding this comment.
This description line should be shortened to 50-150 characters. The full trigger list can go in the body of the skill
| version: "0.2.0" | ||
| --- | ||
|
|
||
| # CUDA-Q Algorithms |
There was a problem hiding this comment.
It should be explicitly stated somewhere in this file that the skill relies on a checkout of the algorithms repo.
| --- | ||
|
|
||
| # CUDA-Q Algorithms | ||
|
|
There was a problem hiding this comment.
It'd be great to add a scripts table somewhere in this file with details about route.py. Something like this.
| --- | ||
| name: cudaq-algorithms | ||
| description: Use when designing, implementing, debugging, reviewing, validating, or composing cudaq_algorithms / CUDA-Q Algorithms APIs and fault-tolerant primitives, including scientific requests expressed as molecular geometry/integrals to energies, lattice Hamiltonians to dynamics, or prepared states to spectral and conditional observables. Covers state preparation, Pauli/block encoding, qubitization, QSP/QSVT, Suzuki–Trotter evolution, fermion transforms, chemistry, double factorization, and statevector analysis. Not for CUDA-Q installation, backend setup, basic standalone kernels such as Bell states, or unrelated quantum-computing questions. | ||
| license: Apache-2.0 |
There was a problem hiding this comment.
It'd be great to add tags and metadata.tags lists to this header as well
| set -o pipefail | ||
| python -B skills/cudaq-algorithms-dev/scripts/check_coverage.py 2>&1 | tee "$RUNNER_TEMP/cudaq-skill-evaluation/coverage.log" | ||
|
|
||
| - name: Upload failure diagnostics |
There was a problem hiding this comment.
This step can be removed, along with the directory prep step and the | tee "$RUNNER_TEMP/..." pipe on each run step. All of that output is being written to the CI logs with or without tee, so the job log already contains everything that this artifact would.
|
Hi Angela, could you take a look at the CI and evaluation side of this PR when you have a chance? I’d start with the workflow, Here are some files to help you find your way around: And the related tests, if useful: Please feel free to follow anything else that catches your attention—the list is just a starting point. Whatever time you can spare would be a big help, and if you have time for a deeper look, all the better. Thanks a lot for helping with this. I know reviewing takes time, and I really appreciate it. |
|
Hi Scott, could you take a look at the skill and the scientific workflows when you have a chance? I’d really like your take on whether this matches what you had in mind for researchers: choosing the right library pieces, getting the physics and conventions right, and checking the results properly. I’d start with Here are the relevant files: The questions, assertions, and numerical checks are here: These six cases would be a useful first sample:
And the related tests, if useful: Please feel free to explore other workflows or question the approach more broadly. These are just suggestions for where to start, and I’d welcome your thoughts wherever you think something could work better. Thanks again for the ideas you shared and for any time you can spend on this. I’m sure your experience and ideas will help make it better. Thanks a lot! |
Signed-off-by: Kevin Mato <kmato@nvidia.com>
|
@kvmto — nice work on the restructure (the smaller QROM + unary iteration are now merged on
Suggested follow-up (not blocking this PR):
Root cause is skill↔library drift: availability claims go stale silently as things merge. A small CI check that cross-validates the skill against the live package surface — for anything the skill (or an eval) marks absent, assert it's still absent in |
|
The latest changes enable automatic skill preloading for Codex and Claude projects and make injected exposure the default for new evaluations. The initial results are encouraging. Average rubric scores improve with the skill across every tested configuration in this snapshot. With injected guidance, requested-output passes increase from 35/55 to 48/55 for Codex and 12/33 to 30/33 for Claude. Both also show lower median task times, although token usage increases for Codex and decreases for Claude. How to interpret injection: listed mode makes the public skill available for discovery; injected mode supplies its full instructions upfront. Private grading criteria and expected answers remain withheld. Preloading makes the guidance consistently available to users. The new project setup implements that through project instructions, whose priority differs from the evaluation’s system-prompt injection. Some rubric items explicitly reward consulting the skill, so rubric improvements partly reflect workflow adherence. Independent executable checks provide complementary evidence and show a more varied picture. The tables therefore report quality, completion, execution checks, time, and tokens separately. Evaluation of the open-source models through the tested APIs has been extremely slow in this run. Median attempt wall times in this snapshot range from roughly 5–11 minutes, compared with about 35–65 seconds for the native Codex and Claude runs. Substantial measured backend waiting has also slowed collection, so these API results remain partial. These are end-to-end timings for the tested tasks, providers, and configurations, not an isolated measure of model capability. Snapshot: data frozen 7 October 2026, 15:35:43 UTC, validated at 15:36:04 UTC; one paired repetition across 62 questions. 812/1,240 attempts finished. Codex and Claude execution is complete for this repetition, with grading still partial; the other models have partial execution and grading. All arrows mean baseline → skill, compared within the same exposure mode. Codex GPT-5.6 Sol — execution complete; grading partial
Claude Sonnet 5 — execution complete; grading partial
Kimi, low reasoning — partial results
Muse — partial results
Nemotron Super — partial results
Reading the tables: Quality metrics use matched fully assessed pairs. Timing uses matched finished pairs, including unsuccessful attempts; wall time includes measured backend waiting. Cumulative time sums attempt durations rather than measuring elapsed campaign time. Token comparisons use only pairs with complete usage records; missing usage remains unknown. Independent-check counts cover applicable finished attempts; Performance should be interpreted within the tested operating conditions. Model capability, reasoning configuration, task complexity, available tools and execution budget all define those conditions. This study did not isolate the causal effect of reasoning effort. Earlier results also used different configurations and should be treated as separate experiments. Scientific-agent evaluation is demanding because a satisfactory result must connect several stages: understand the research question, choose an appropriate method, use the requested library, produce an executable artifact and validate its output. The present evaluation makes those stages visible and provides a useful basis for improving them. These findings support a focused product-validation phase using Codex, Claude, Muse, Nemotron Super and Kimi. The immediate priorities are reliable artifact delivery, precise convention handling and bounded execution for expensive workflows. The 62 questions should remain fixed while targeted paired checks establish whether those changes improve outcomes. The evidence supports further investment in a scoped pilot. Broader readiness should be established for each supported model and workflow through reproducible correctness and reliability measurements. |
Summary
Adds a CUDA-Q Algorithms agent skill that translates researchers’ inputs and requested scientific outputs into library operations, composed workflows, and numerical verification. For example, a molecular-energy or lattice-dynamics request routes to the relevant scientific guide, API contracts, and current source without requiring the researcher to name repository files.
The runtime skill lives in
skills/cudaq-algorithms; authoring, coverage, evaluation, and reporting tools live inskills/cudaq-algorithms-dev.Runtime skill
Evaluation infrastructure
The committed runner supports OpenAI-compatible endpoints; native Codex/Claude adapters are outside this PR. Rubric scores, executable correctness, and resource measurements are reported separately. Full multi-model performance results are still being collected.
Validation
Verified on PR head
ee5710be:The evaluation workflow runs on relevant pull requests and supports manual and reusable invocation. It requires no model credentials; live model campaigns use the documented runner and external evidence directories.
See the evaluation protocol and runner usage.