Skip to content

[WIP] NPU GPU co-execution in QNN EP - #794

Draft
qti-ashwshan wants to merge 1 commit into
mainfrom
dev/ashwshan/gpu-npu-co-execution
Draft

qti-ashwshan wants to merge 1 commit into
mainfrom
dev/ashwshan/gpu-npu-co-execution

Conversation

@qti-ashwshan

Copy link
Copy Markdown
Collaborator

Description

  • Segmented execution: HTP and GPU segments compiled from a single fused node; intermediate buffers plumbed between segments with DX12 MEMHANDLE shadow tensors for GPU-external I/O.
  • Skip GPU validation for nodes HTP already accepted (skip_nodes on GetSupportedNodes overload) so only HTP-rejected ops are validated against GPU.
  • Debug instrumentation: STANDALONE-TENSOR-DUMP, MEM-BIND traces, CREATESTATE-DEBUG for compute-info wiring.

Motivation and Context

- Segmented execution: HTP and GPU segments compiled from a single
  fused node; intermediate buffers plumbed between segments with
  DX12 MEMHANDLE shadow tensors for GPU-external I/O.
- Skip GPU validation for nodes HTP already accepted (skip_nodes on
  GetSupportedNodes overload) so only HTP-rejected ops are validated
  against GPU.
- Debug instrumentation: STANDALONE-TENSOR-DUMP, MEM-BIND traces,
  CREATESTATE-DEBUG for compute-info wiring.
@qti-ashwshan qti-ashwshan self-assigned this Sep 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant