diff --git a/docs/execution_providers/QNN-ExecutionProvider.md b/docs/execution_providers/QNN-ExecutionProvider.md index 6338ff8caa..cd0602a0f7 100644 --- a/docs/execution_providers/QNN-ExecutionProvider.md +++ b/docs/execution_providers/QNN-ExecutionProvider.md @@ -1728,6 +1728,9 @@ To enable new operator support in EP, areas to visit: A **User-Defined Operation (UDO)** allows developers to extend the Qualcomm® Neural Network (QNN) runtimes with custom operators. UDO enables execution of operations that are not natively supported in the default QNN op set, while maintaining compatibility with model conversion, compilation, and runtime execution. +For an end-to-end MyAdd UDO reference, including CPU, HTP, and on-device +commands, see the [QNN UDO sample](../../qcom/samples/qnn_udo_myadd/README.md). + ### Overview A UDO lets you define and register custom operations—describing their inputs, outputs, parameters, data types, and backend behavior—so they can run on: diff --git a/qcom/samples/qnn_udo_myadd/.gitignore b/qcom/samples/qnn_udo_myadd/.gitignore new file mode 100644 index 0000000000..b75f930091 --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/.gitignore @@ -0,0 +1,9 @@ +# Artifacts generated while building or running the sample. +/artifacts/ +/build/ +/run_udo_sample +/test_data_set_0/ +/__pycache__/ +QNNExecutionProvider_*.json +/*.onnx +/*.bin diff --git a/qcom/samples/qnn_udo_myadd/README.md b/qcom/samples/qnn_udo_myadd/README.md new file mode 100644 index 0000000000..6de2d71de7 --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/README.md @@ -0,0 +1,351 @@ +# QNN EP UDO Sample: MyAdd + +End-to-end reference showing how to run an ONNX model with a QNN +User-Defined Operation (UDO) through the ORT QNN Execution Provider. + +**Op**: `MyAdd` (`domain="example"`) — computes `output = input + constant`. +**Backends demonstrated**: QNN CPU (float32), QNN HTP x86 simulator (QDQ uint8), and on-device HTP (arm64). + +--- + +## Directory contents + +| File | Purpose | +|------|---------| +| `gen_myadd_model.py` | Generate `myadd_fp32.onnx` (CPU) and `myadd_qdq.onnx` (HTP) | +| `gen_myadd_test_data.py` | Generate `onnxruntime_plugin_ep_onnx_test` protobuf inputs and references for on-device HTP | +| `run_udo_sample.cc` | C++ standalone sample — CPU and HTP modes | +| `run_udo_sample.py` | Python sample — CPU and HTP modes | +| `build_op_package.sh` | Build `libMyAddOpPackage_cpu.so` / `libMyAddOpPackage_htp.so` | + +The sample uses the canonical MyAdd op-package source shared with the UDO unit +tests: `onnxruntime/test/providers/qnn/udo/` (`MyAddOpPackageCpu.xml`, +`MyAddOpPackageHtp.xml`, `MyAddCPU.cpp`, `MyAddHTP.cpp`, and `HTP_Makefile`). +Run the commands from a repository checkout so `build_op_package.sh` can locate +those inputs. + +--- + +## Prerequisites + +| Requirement | Version | Notes | +|-------------|---------|-------| +| QAIRT SDK | any version with `qnn-op-package-generator` | Set `QNN_SDK_ROOT=/qairt/` | +| LLVM | 21.1.8 | Set `LLVM_TOOL_DIR=/LLVM-21.1.8-Linux-X64` | +| Hexagon SDK | 6.5.0.0 | HTP only; set `HEXAGON_SDK_ROOT=/6.5.0.0` | +| Python | **3.12** | Use 3.12 throughout. The op-package generator supports 3.10/3.12, but the Python sample needs a QNN-EP-compatible host ORT (≥ 1.24), and public PyPI `onnxruntime` has no cp310 wheels past 1.23.2 — too old for the plugin. On 3.12, `pip install onnxruntime` gets a compatible release. | +| onnx, numpy | any recent | `pip install onnx numpy` | +| ONNX Runtime C/C++ package | version ABI-compatible with the QNN EP build | Extract the released core ORT package and set `ORT_PREBUILT_ROOT`; it supplies public headers and `libonnxruntime.so`. | +| QNN EP build | this repo | Set `ORT_BUILD` to its Release directory. It supplies `libonnxruntime_providers_qnn.so` and the test binaries. | +| onnxruntime (Python) | ≥ 1.24 | In a Python 3.12 venv: `pip install onnxruntime`. QNN EP libraries come from the built `onnxruntime_qnn` wheel (`build/linux-x86_64/Release/dist/`). | + +--- + +## Step 1 — Generate ONNX models + +```bash +cd qcom/samples/qnn_udo_myadd/ +python3 gen_myadd_model.py --constant 2.0 --outdir . +# Produces: myadd_fp32.onnx, myadd_qdq.onnx +``` + +--- + +## Step 2 — Build the QNN op packages + +```bash +export QNN_SDK_ROOT=/path/to/qairt/ +export LLVM_TOOL_DIR=/path/to/LLVM-21.1.8-Linux-X64 +export HEXAGON_SDK_ROOT=/path/to/Hexagon_SDK/6.5.0.0 # HTP only + +./build_op_package.sh all +# Produces: +# artifacts/libMyAddOpPackage_cpu.so (QNN CPU op package) +# artifacts/libMyAddOpPackage_htp.so (QNN HTP op package) +``` + +Individual targets: `./build_op_package.sh cpu` or `htp`. + +--- + +## Step 3a — Run C++ sample + +`run_udo_sample.cc` and the following commands target Linux only. They use +Linux shared-library names and `LD_LIBRARY_PATH`. + +```bash +# Use a compatible released core ORT package for headers and libonnxruntime.so. +# The QNN EP build supplies the plugin. +ORT_BUILD=/path/to/qnn-ep/build/linux-x86_64/Release +ORT_PREBUILT_ROOT=/path/to/onnxruntime-linux-x64- +ORT_HEADERS=${ORT_PREBUILT_ROOT}/include +ORT_LIB=${ORT_PREBUILT_ROOT}/lib + +# This sample registers the QNN plugin by filename. Place a symlink beside the +# prebuilt libonnxruntime.so so ORT can resolve the plugin. Do not replace an +# existing plugin in a packaged deployment. +ln -s ${ORT_BUILD}/libonnxruntime_providers_qnn.so \ + ${ORT_LIB}/libonnxruntime_providers_qnn.so + +g++ -std=c++17 run_udo_sample.cc \ + -I${ORT_HEADERS} \ + -L${ORT_LIB} -lonnxruntime \ + -Wl,-rpath,${ORT_LIB} \ + -o run_udo_sample + +# Set env var so the QNN EP factory registers the custom-op domain automatically +export ORT_QNN_CUSTOM_OP_DOMAINS="example:MyAdd" + +# CPU backend (libQnnCpu.so must be on LD_LIBRARY_PATH) +LD_LIBRARY_PATH=${QNN_SDK_ROOT}/lib/x86_64-linux-clang:${ORT_LIB}:${LD_LIBRARY_PATH} \ + ./run_udo_sample cpu myadd_fp32.onnx artifacts/libMyAddOpPackage_cpu.so + +# HTP backend (libQnnHtp.so must be on LD_LIBRARY_PATH) +LD_LIBRARY_PATH=${QNN_SDK_ROOT}/lib/x86_64-linux-clang:${ORT_LIB}:${LD_LIBRARY_PATH} \ + ./run_udo_sample htp myadd_qdq.onnx artifacts/libMyAddOpPackage_htp.so +``` + +Expected output (CPU): +``` +=== QNN CPU backend === +Max absolute error vs (input + 2.0): 0.00e+00 +PASS +``` + +Expected output (HTP): +``` +=== QNN HTP backend === +Max absolute error vs (input + 2.0): ... (QDQ tol: 0.0314) +PASS +``` + +The HTP sample permits two output-quantization steps over `[0, 4]`, so its +tolerance is `2 × (4/255) ≈ 0.0314`. + +### Host verification cross-checks + +Confirm that the CPU model is assigned to QNN rather than falling back to the +CPU EP: + +```bash +LD_LIBRARY_PATH=${QNN_SDK_ROOT}/lib/x86_64-linux-clang:${ORT_LIB}:${LD_LIBRARY_PATH} \ +ORT_LOG_LEVEL=1 ./run_udo_sample cpu myadd_fp32.onnx artifacts/libMyAddOpPackage_cpu.so 2>&1 \ + | grep -i "node.*assign\|partition\|MyAdd" +``` + +The corresponding gtests in the same QNN EP build output validate the CPU and +HTP paths: + +```bash +${ORT_BUILD}/onnxruntime_provider_test \ + --gtest_filter="QnnCPUBackendTests.UDO_Op_MyAdd_AutoDomainOnly" +${ORT_BUILD}/onnxruntime_provider_test \ + --gtest_filter="QnnHTPBackendTests.UDO_Op_MyAdd" +``` + +--- + +## Step 3b — Run Python sample + +```bash +# Path to the onnxruntime-qnn package directory (ships libonnxruntime_providers_qnn.so +# and all QNN backend libs: libQnnCpu.so, libQnnHtp.so, etc.) +QNN_PKG=$(python3 -c "import onnxruntime_qnn, os; print(os.path.dirname(onnxruntime_qnn.__file__))") +ORT_LIB=$(python3 -c "import onnxruntime, os; print(os.path.join(os.path.dirname(onnxruntime.__file__), 'capi'))") + +# Set env var so the QNN EP factory registers the custom-op domain automatically +export ORT_QNN_CUSTOM_OP_DOMAINS="example:MyAdd" + +# CPU backend +LD_LIBRARY_PATH=${QNN_PKG}:${ORT_LIB}:${LD_LIBRARY_PATH} \ +python3 run_udo_sample.py cpu myadd_fp32.onnx \ + --op-package artifacts/libMyAddOpPackage_cpu.so \ + --qnn-ep-lib ${QNN_PKG}/libonnxruntime_providers_qnn.so + +# HTP backend (libQnnHtp.so is bundled in QNN_PKG) +LD_LIBRARY_PATH=${QNN_PKG}:${ORT_LIB}:${LD_LIBRARY_PATH} \ +python3 run_udo_sample.py htp myadd_qdq.onnx \ + --op-package artifacts/libMyAddOpPackage_htp.so \ + --qnn-ep-lib ${QNN_PKG}/libonnxruntime_providers_qnn.so +``` + +--- + +## Step 4 — On-device HTP (arm64) + +`build_op_package.sh` targets the x86 simulator only. On-device HTP requires +**two** separately-built op-package halves (the aarch64-android registration lib +and the hexagon-vNN DSP skel), plus the correct test runner. + +### 4a — Build both op-package halves for the device arch + +Set `HEXAGON_VER` to an HTP target compatible with the device and QAIRT runtime +(obtain a supported target from the verbose-log line `Setting libnative architecture +to vNN`, or from the SoC datasheet). The QAIRT SDK must ship +`lib/hexagon-v${HEXAGON_VER}`. A runtime may select a higher native architecture +while executing a compatible lower-target package, so the two values need not be +identical. + +```bash +REPO_ROOT=$(git rev-parse --show-toplevel) +OP_PACKAGE_DIR=${REPO_ROOT}/onnxruntime/test/providers/qnn/udo +export DEVICE_SERIAL= # from `adb devices` +export HEXAGON_VER=75 # e.g. 75, 79, 81 — must be device/runtime-compatible +export QNN_SDK_ROOT= +export HEXAGON_SDK_ROOT=/6.5.0.0 +export LLVM_TOOL_DIR=/LLVM-21.1.8-Linux-X64 + +BUILD=/tmp/udo_arm64 +rm -rf "${BUILD}" +PYTHONPATH=${QNN_SDK_ROOT}/lib/python \ +python3 ${QNN_SDK_ROOT}/bin/x86_64-linux-clang/qnn-op-package-generator \ + -p ${OP_PACKAGE_DIR}/MyAddOpPackageHtp.xml -o "${BUILD}" +cp ${OP_PACKAGE_DIR}/MyAddHTP.cpp "${BUILD}/MyAddOpPackage/src/ops/MyAdd.cpp" +cp ${OP_PACKAGE_DIR}/HTP_Makefile "${BUILD}/MyAddOpPackage/Makefile" + +env QNN_SDK_ROOT="${QNN_SDK_ROOT}" HEXAGON_SDK_ROOT="${HEXAGON_SDK_ROOT}" \ + PATH="${LLVM_TOOL_DIR}/bin:${PATH}" \ + make -C "${BUILD}/MyAddOpPackage" \ + "X86_CXX=${LLVM_TOOL_DIR}/bin/clang++ -stdlib=libc++" \ + htp_aarch64 htp_v${HEXAGON_VER} +# Outputs: +# ARM lib : ${BUILD}/MyAddOpPackage/libs/aarch64-android/libQnnMyAddOpPackage.so +# DSP skel: ${BUILD}/MyAddOpPackage/build/hexagon-v${HEXAGON_VER}/libQnnMyAddOpPackage.so +``` + +### 4b — Sign the DSP skel (if required) + +If the device's process domain requires skel signing, sign the DSP skel before +deployment. Refer to the QAIRT SDK signing documentation. + +### 4c — Deploy artifacts + +Deploy the ARM lib at the top level of `${DEVICE_DIR}` and the DSP skel into a +separate `${DSP_DIR}` under the **exact same filename**. The test runner consumes +an `onnxruntime_plugin_ep_onnx_test` test-case directory (`model.onnx` + `test_data_set_0/`) and treats +subdirectories under that root as test cases, so `${DSP_DIR}` must be separate. + +```bash +DEVICE_DIR=/data/local/tmp/udo_test # model/test-data root consumed by the runner +DSP_DIR=/data/local/tmp/udo_test_dsp # keep outside DEVICE_DIR +ORT_ARM64_BUILD=/path/to/build/android-aarch64/Release + +adb -s ${DEVICE_SERIAL} shell "rm -rf ${DEVICE_DIR} ${DSP_DIR}; mkdir -p ${DEVICE_DIR}/test_data_set_0 ${DSP_DIR}" + +# ORT arm64 binaries + libs +adb -s ${DEVICE_SERIAL} push ${ORT_ARM64_BUILD}/onnxruntime_plugin_ep_onnx_test ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${ORT_ARM64_BUILD}/libonnxruntime.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${ORT_ARM64_BUILD}/libonnxruntime_providers_qnn.so ${DEVICE_DIR}/ + +# QNN aarch64-android backend libs + device-arch skel/stub +adb -s ${DEVICE_SERIAL} push ${QNN_SDK_ROOT}/lib/aarch64-android/libQnnHtp.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${QNN_SDK_ROOT}/lib/aarch64-android/libQnnHtpPrepare.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${QNN_SDK_ROOT}/lib/aarch64-android/libQnnSystem.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${QNN_SDK_ROOT}/lib/aarch64-android/libQnnHtpV${HEXAGON_VER}Stub.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${QNN_SDK_ROOT}/lib/hexagon-v${HEXAGON_VER}/unsigned/libQnnHtpV${HEXAGON_VER}Skel.so ${DEVICE_DIR}/ + +# MyAdd op package: ARM lib at top level, DSP skel outside the runner root. +# The runner treats subdirectories under DEVICE_DIR as test cases, so do not use DEVICE_DIR/dsp. +adb -s ${DEVICE_SERIAL} push ${BUILD}/MyAddOpPackage/libs/aarch64-android/libQnnMyAddOpPackage.so ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push ${BUILD}/MyAddOpPackage/build/hexagon-v${HEXAGON_VER}/libQnnMyAddOpPackage.so ${DSP_DIR}/ + +# Model + deterministic test data (float reference = input + 2.0) +TC=/tmp/udo_qdq_testcase +rm -rf "${TC}" +mkdir -p "${TC}" +cp myadd_qdq.onnx "${TC}/model.onnx" +python3 gen_myadd_test_data.py --constant 2.0 --outdir "${TC}" +adb -s ${DEVICE_SERIAL} push "${TC}/model.onnx" ${DEVICE_DIR}/ +adb -s ${DEVICE_SERIAL} push "${TC}/test_data_set_0/." ${DEVICE_DIR}/test_data_set_0/ +``` + +### 4d — Run on device (dual CPU+HTP op-package registration) + +On-device HTP requires the op package registered for **both** processors in one +comma-separated `op_packages` string: + +- **CPU entry** — absolute path to the aarch64-android lib, target `:CPU` + (ARM-side graph prepare) +- **HTP entry** — **bare filename** of the DSP skel, target `:HTP` + (NSP kernel execution; resolved via `ADSP_LIBRARY_PATH`) + +> Supplying only one half fails: `:CPU` alone → `INVALID_HANDLE (6001)` at +> execute (no DSP kernel registered); `:HTP` alone → `Could not find an +> implementation for MyAdd` at model load. The skel target must be compatible +> with the device/runtime; an incompatible skel can finalize at the host API but +> fail on the DSP. + +```bash +DEVICE_DIR=/data/local/tmp/udo_test +DSP_DIR=/data/local/tmp/udo_test_dsp +adb -s ${DEVICE_SERIAL} shell "cd ${DEVICE_DIR} && \ + LD_LIBRARY_PATH=${DEVICE_DIR} \ + ADSP_LIBRARY_PATH='${DSP_DIR};${DEVICE_DIR};/dsp/cdsp;/vendor/lib/rfsa/adsp;/system/lib/rfsa/adsp' \ + ORT_QNN_CUSTOM_OP_DOMAINS=example:MyAdd \ + ./onnxruntime_plugin_ep_onnx_test \ + --plugin_ep_libs 'QNNExecutionProvider|${DEVICE_DIR}/libonnxruntime_providers_qnn.so' \ + --plugin_eps 'QNNExecutionProvider' \ + --plugin_ep_options 'backend_type|htp offload_graph_io_quantization|0 op_packages|MyAdd:${DEVICE_DIR}/libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:CPU,MyAdd:libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:HTP' \ + -a 0.04 -t 0.02 -j 1 \ + ${DEVICE_DIR}" +``` + +`ORT_QNN_CUSTOM_OP_DOMAINS=example:MyAdd` registers the ONNX custom-op schema at +factory-load time (same mechanism as Steps 3a/3b). + +Expected output: `Succeeded: 1`. The QDQ output is checked against the float +reference with `-a 0.04 -t 0.02`; two quantization steps over `[0, 4]` are +about `0.0314`. + +--- + +## Step 5 — EPContext binary on-device (arm64) + +This step reuses the deployment from Step 4, including both op-package halves. +Context generation requires the custom-op schema; inference from the generated +context model intentionally does not. Keep context artifacts outside +`${DEVICE_DIR}`: the runner discovers every `.onnx` and subdirectory in that +directory as a model/test case. + +```bash +# Generate an external context ONNX + .bin. Keep ORT_QNN_CUSTOM_OP_DOMAINS for this source-model run. +CONTEXT_DIR=/data/local/tmp/udo_context +DSP_DIR=/data/local/tmp/udo_test_dsp +adb -s ${DEVICE_SERIAL} shell "rm -rf ${CONTEXT_DIR}; mkdir -p ${CONTEXT_DIR}" +adb -s ${DEVICE_SERIAL} shell "cd ${DEVICE_DIR} && \ + LD_LIBRARY_PATH=${DEVICE_DIR} \ + ADSP_LIBRARY_PATH='${DSP_DIR};${DEVICE_DIR};/dsp/cdsp;/vendor/lib/rfsa/adsp;/system/lib/rfsa/adsp' \ + ORT_QNN_CUSTOM_OP_DOMAINS=example:MyAdd \ + ./onnxruntime_plugin_ep_onnx_test -b \ + -C 'ep.context_enable|1 ep.context_file_path|${CONTEXT_DIR}/myadd_ctx.onnx' \ + --plugin_ep_libs 'QNNExecutionProvider|${DEVICE_DIR}/libonnxruntime_providers_qnn.so' \ + --plugin_eps 'QNNExecutionProvider' \ + --plugin_ep_options 'backend_type|htp offload_graph_io_quantization|0 op_packages|MyAdd:${DEVICE_DIR}/libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:CPU,MyAdd:libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:HTP' \ + -a 0.04 -t 0.02 -j 1 ${DEVICE_DIR}" + +adb -s ${DEVICE_SERIAL} shell "ls -lh ${CONTEXT_DIR}/myadd_ctx.onnx ${CONTEXT_DIR}/*.bin" + +# Run the context model. Copy both artifacts so its relative .bin reference resolves. +# Do NOT set ORT_QNN_CUSTOM_OP_DOMAINS here. +adb -s ${DEVICE_SERIAL} shell "cp ${CONTEXT_DIR}/myadd_ctx.onnx ${DEVICE_DIR}/model.onnx && cp ${CONTEXT_DIR}/*.bin ${DEVICE_DIR}/" +adb -s ${DEVICE_SERIAL} shell "cd ${DEVICE_DIR} && \ + LD_LIBRARY_PATH=${DEVICE_DIR} \ + ADSP_LIBRARY_PATH='${DSP_DIR};${DEVICE_DIR};/dsp/cdsp;/vendor/lib/rfsa/adsp;/system/lib/rfsa/adsp' \ + ./onnxruntime_plugin_ep_onnx_test \ + --plugin_ep_libs 'QNNExecutionProvider|${DEVICE_DIR}/libonnxruntime_providers_qnn.so' \ + --plugin_eps 'QNNExecutionProvider' \ + --plugin_ep_options 'backend_type|htp offload_graph_io_quantization|0 op_packages|MyAdd:${DEVICE_DIR}/libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:CPU,MyAdd:libQnnMyAddOpPackage.so:MyAddOpPackageInterfaceProvider:HTP' \ + -a 0.04 -t 0.02 -j 1 ${DEVICE_DIR}" +``` + +Expected output: `Succeeded: 1` for both runs. The generated `.onnx` and `.bin` +must be non-empty. The context run still needs both `op_packages` entries, but +must succeed without `ORT_QNN_CUSTOM_OP_DOMAINS`. + +--- + +## Cross-reference + +- Unit test (C++ gtest): `onnxruntime/test/providers/qnn/udo_op_test.cc` +- Build automation: `cmake/onnxruntime_unittests_udo.cmake` +- ORT QNN EP documentation: `docs/execution_providers/QNN-ExecutionProvider.md` §"QNN User-Defined Operation" diff --git a/qcom/samples/qnn_udo_myadd/build_op_package.sh b/qcom/samples/qnn_udo_myadd/build_op_package.sh new file mode 100755 index 0000000000..d4861b77e6 --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/build_op_package.sh @@ -0,0 +1,105 @@ +#!/usr/bin/env bash +# Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries. +# SPDX-License-Identifier: MIT +# +# build_op_package.sh -- Build MyAdd QNN op package(s). +# +# Usage: +# ./build_op_package.sh cpu # build CPU x86 op package +# ./build_op_package.sh htp # build HTP x86 op package +# ./build_op_package.sh all # build both +# +# Required environment variables: +# QNN_SDK_ROOT – path to QAIRT SDK root (e.g. .../qairt/) +# LLVM_TOOL_DIR – path to LLVM bin dir (e.g. .../LLVM-21.1.8-Linux-X64) +# HEXAGON_SDK_ROOT – (HTP only) path to Hexagon SDK version dir (e.g. .../6.5.0.0) +# +# Outputs (under this sample's artifacts/ directory): +# artifacts/libMyAddOpPackage_cpu.so (CPU target) +# artifacts/libMyAddOpPackage_htp.so (HTP target) + +set -euo pipefail + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +REPO_ROOT="$(cd "${SCRIPT_DIR}/../../.." && pwd)" +OP_PACKAGE_DIR="${REPO_ROOT}/onnxruntime/test/providers/qnn/udo" +ARTIFACT_DIR="${SCRIPT_DIR}/artifacts" +BUILD_DIR="${ARTIFACT_DIR}/build" + +build_cpu() { + QNN_SDK_ROOT="${QNN_SDK_ROOT:?QNN_SDK_ROOT must be set}" + LLVM_TOOL_DIR="${LLVM_TOOL_DIR:?LLVM_TOOL_DIR must be set}" + echo ">>> Building CPU x86 op package..." + mkdir -p "${ARTIFACT_DIR}" + local cpu_build="${BUILD_DIR}/cpu" + rm -rf "${cpu_build}" + + # Step 1: generate skeleton + PYTHONPATH="${QNN_SDK_ROOT}/lib/python" \ + python3 "${QNN_SDK_ROOT}/bin/x86_64-linux-clang/qnn-op-package-generator" \ + -p "${OP_PACKAGE_DIR}/MyAddOpPackageCpu.xml" \ + -o "${cpu_build}" + + # Step 2: copy pre-implemented kernel + /bin/cp "${OP_PACKAGE_DIR}/MyAddCPU.cpp" \ + "${cpu_build}/MyAddOpPackage/src/ops/MyAdd.cpp" + + # Step 3: build + QNN_SDK_ROOT="${QNN_SDK_ROOT}" \ + PATH="${LLVM_TOOL_DIR}/bin:${PATH}" \ + make -C "${cpu_build}/MyAddOpPackage" \ + "CXX=${LLVM_TOOL_DIR}/bin/clang++ -stdlib=libc++ -static-libstdc++ -Wl,--exclude-libs,ALL" \ + all_x86 + + # Step 4: copy output + /bin/cp "${cpu_build}/MyAddOpPackage/libs/x86_64-linux-clang/libMyAddOpPackage.so" \ + "${ARTIFACT_DIR}/libMyAddOpPackage_cpu.so" + echo ">>> CPU package: ${ARTIFACT_DIR}/libMyAddOpPackage_cpu.so" +} + +build_htp() { + QNN_SDK_ROOT="${QNN_SDK_ROOT:?QNN_SDK_ROOT must be set}" + LLVM_TOOL_DIR="${LLVM_TOOL_DIR:?LLVM_TOOL_DIR must be set}" + HEXAGON_SDK_ROOT="${HEXAGON_SDK_ROOT:?HEXAGON_SDK_ROOT must be set for HTP build}" + mkdir -p "${ARTIFACT_DIR}" + local htp_build="${BUILD_DIR}/htp" + rm -rf "${htp_build}" + + # Step 1: generate skeleton + PYTHONPATH="${QNN_SDK_ROOT}/lib/python" \ + python3 "${QNN_SDK_ROOT}/bin/x86_64-linux-clang/qnn-op-package-generator" \ + -p "${OP_PACKAGE_DIR}/MyAddOpPackageHtp.xml" \ + -o "${htp_build}" + + # Step 2: copy pre-implemented kernel + custom HTP Makefile + /bin/cp "${OP_PACKAGE_DIR}/MyAddHTP.cpp" \ + "${htp_build}/MyAddOpPackage/src/ops/MyAdd.cpp" + /bin/cp "${OP_PACKAGE_DIR}/HTP_Makefile" \ + "${htp_build}/MyAddOpPackage/Makefile" + + # Step 3: build + QNN_SDK_ROOT="${QNN_SDK_ROOT}" \ + HEXAGON_SDK_ROOT="${HEXAGON_SDK_ROOT}" \ + PATH="${LLVM_TOOL_DIR}/bin:${PATH}" \ + make -C "${htp_build}/MyAddOpPackage" \ + "X86_CXX=${LLVM_TOOL_DIR}/bin/clang++ -stdlib=libc++" \ + htp_x86 + + # Step 4: copy output + /bin/cp "${htp_build}/MyAddOpPackage/build/x86_64-linux-clang/libQnnMyAddOpPackage.so" \ + "${ARTIFACT_DIR}/libMyAddOpPackage_htp.so" + echo ">>> HTP package: ${ARTIFACT_DIR}/libMyAddOpPackage_htp.so" +} + +TARGET="${1:-all}" +case "${TARGET}" in + cpu) build_cpu ;; + htp) build_htp ;; + all) build_cpu; build_htp ;; + *) + echo "Usage: $0 [cpu|htp|all]" + exit 1 + ;; +esac + +echo ">>> Done." diff --git a/qcom/samples/qnn_udo_myadd/gen_myadd_model.py b/qcom/samples/qnn_udo_myadd/gen_myadd_model.py new file mode 100644 index 0000000000..fd5c6471aa --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/gen_myadd_model.py @@ -0,0 +1,135 @@ +# Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries. +# SPDX-License-Identifier: MIT +""" +Generate two ONNX models containing a single MyAdd UDO node: + - myadd_fp32.onnx : float32 model (for QNN CPU backend) + - myadd_qdq.onnx : uint8 QDQ model (DQ -> MyAdd -> Q, for QNN HTP backend) + +MyAdd computes: output = input + constant + input : shape [1, 32], float32 + output : shape [1, 32], float32 + constant attr : float, default 2.0 + +Usage: + python gen_myadd_model.py [--constant 2.0] [--outdir .] +""" + +import argparse +from pathlib import Path + +import numpy as np +import onnx +from onnx import TensorProto, helper, numpy_helper + +DOMAIN = "example" +OP_TYPE = "MyAdd" +INPUT_SHAPE = [1, 32] + + +def make_fp32_model(constant: float) -> onnx.ModelProto: + """Float32 model: input -> MyAdd -> output.""" + input_vi = helper.make_tensor_value_info("input", TensorProto.FLOAT, INPUT_SHAPE) + output_vi = helper.make_tensor_value_info("output", TensorProto.FLOAT, INPUT_SHAPE) + + constant_attr = helper.make_attribute("constant", constant) + node = helper.make_node(OP_TYPE, inputs=["input"], outputs=["output"], domain=DOMAIN) + node.attribute.append(constant_attr) + + graph = helper.make_graph([node], "myadd_fp32", [input_vi], [output_vi]) + opset = helper.make_opsetid(DOMAIN, 1) + model = helper.make_model(graph, opset_imports=[opset]) + model.ir_version = 8 + onnx.checker.check_model(model, full_check=False) + return model + + +def make_qdq_model(constant: float) -> onnx.ModelProto: + """ + QDQ model for HTP backend: + input (f32) -> QuantizeLinear -> input_q (u8) -> DequantizeLinear -> input_dq (f32) + -> MyAdd -> output_dq (f32) -> QuantizeLinear -> output_q (u8) -> DequantizeLinear -> output (f32) + + Scale/zero-point are computed from [-1, 1] range matching udo_op_test.cc test input. + """ + # Input quantization: [-1, 1] range -> uint8 (scale=2/255, zp=128 so 0.0 maps to 128) + scale_in = np.float32(2.0 / 255.0) + zp_in = np.uint8(128) + + # Output quantization: [0, 4] range (conservatively covers input[-1,1]+constant[2.0]) + # zp=0 (asymmetric, all values positive) + scale_out = np.float32(4.0 / 255.0) + zp_out = np.uint8(0) + + def quant_init(name: str, val) -> onnx.TensorProto: + return numpy_helper.from_array(np.array(val), name=name) + + # scale/zero_point initializers + inits = [ + quant_init("scale_in", scale_in), + quant_init("zp_in", zp_in), + quant_init("scale_out", scale_out), + quant_init("zp_out", zp_out), + ] + + # Value infos + def make_f32_value_info(name: str) -> onnx.ValueInfoProto: + return helper.make_tensor_value_info(name, TensorProto.FLOAT, INPUT_SHAPE) + + input_vi = make_f32_value_info("input") + output_vi = make_f32_value_info("output") + + # Declared type/shape for the MyAdd output (intermediate tensor feeding the + # output QuantizeLinear). The QNN EP auto-registration path registers a + # placeholder op with no shape/type inference, so ORT resolves the custom-op + # output type from this value_info at model-load time. Without it, load fails + # with "type inference failed" for the custom-domain node. + output_dq_vi = make_f32_value_info("output_dq") + + # Nodes: Q -> DQ -> MyAdd -> Q -> DQ + q_in = helper.make_node("QuantizeLinear", ["input", "scale_in", "zp_in"], ["input_q"], axis=None) + dq_in = helper.make_node("DequantizeLinear", ["input_q", "scale_in", "zp_in"], ["input_dq"], axis=None) + + constant_attr = helper.make_attribute("constant", constant) + myadd = helper.make_node(OP_TYPE, ["input_dq"], ["output_dq"], domain=DOMAIN) + myadd.attribute.append(constant_attr) + + q_out = helper.make_node("QuantizeLinear", ["output_dq", "scale_out", "zp_out"], ["output_q"], axis=None) + dq_out = helper.make_node("DequantizeLinear", ["output_q", "scale_out", "zp_out"], ["output"], axis=None) + + graph = helper.make_graph( + [q_in, dq_in, myadd, q_out, dq_out], + "myadd_qdq", + [input_vi], + [output_vi], + initializer=inits, + value_info=[output_dq_vi], + ) + onnx_opset = helper.make_opsetid("", 21) + custom_opset = helper.make_opsetid(DOMAIN, 1) + model = helper.make_model(graph, opset_imports=[onnx_opset, custom_opset]) + model.ir_version = 8 + onnx.checker.check_model(model, full_check=False) + return model + + +def main(): + parser = argparse.ArgumentParser(description="Generate MyAdd UDO ONNX models") + parser.add_argument("--constant", type=float, default=2.0, help="Value added to each input element (default: 2.0)") + parser.add_argument("--outdir", default=".", help="Output directory") + args = parser.parse_args() + + outdir = Path(args.outdir) + outdir.mkdir(parents=True, exist_ok=True) + + fp32_path = outdir / "myadd_fp32.onnx" + qdq_path = outdir / "myadd_qdq.onnx" + + onnx.save(make_fp32_model(args.constant), fp32_path) + print(f"Saved {fp32_path}") + + onnx.save(make_qdq_model(args.constant), qdq_path) + print(f"Saved {qdq_path}") + + +if __name__ == "__main__": + main() diff --git a/qcom/samples/qnn_udo_myadd/gen_myadd_test_data.py b/qcom/samples/qnn_udo_myadd/gen_myadd_test_data.py new file mode 100644 index 0000000000..186acd4bdb --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/gen_myadd_test_data.py @@ -0,0 +1,37 @@ +# Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries. +# SPDX-License-Identifier: MIT +"""Generate protobuf test data for onnxruntime_plugin_ep_onnx_test. + +The input matches the standalone samples: 32 float32 values evenly spaced in +[-1, 1]. The reference is the unquantized mathematical result (input + +constant); the on-device runner applies the documented QDQ tolerance. +""" + +import argparse +from pathlib import Path + +import numpy as np +from onnx import numpy_helper + + +def main(): + parser = argparse.ArgumentParser(description="Generate MyAdd ONNX test data") + parser.add_argument("--constant", type=float, default=2.0, help="Value added by MyAdd (default: 2.0)") + parser.add_argument("--outdir", required=True, help="Test-case directory that will contain test_data_set_0") + args = parser.parse_args() + + data_dir = Path(args.outdir) / "test_data_set_0" + data_dir.mkdir(parents=True, exist_ok=True) + input_data = np.linspace(-1.0, 1.0, 32, dtype=np.float32).reshape(1, 32) + output_data = input_data + np.float32(args.constant) + + with (data_dir / "input_0.pb").open("wb") as f: + f.write(numpy_helper.from_array(input_data, name="input").SerializeToString()) + with (data_dir / "output_0.pb").open("wb") as f: + f.write(numpy_helper.from_array(output_data, name="output").SerializeToString()) + + print(f"Saved {data_dir}/input_0.pb and output_0.pb") + + +if __name__ == "__main__": + main() diff --git a/qcom/samples/qnn_udo_myadd/run_udo_sample.cc b/qcom/samples/qnn_udo_myadd/run_udo_sample.cc new file mode 100644 index 0000000000..c0fbe0f853 --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/run_udo_sample.cc @@ -0,0 +1,180 @@ +// Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries. +// SPDX-License-Identifier: MIT + +// Standalone C++ reference sample for running a QNN UDO (User-Defined Op). +// +// Demonstrates the MyAdd UDO (output = input + constant) on: +// - QNN CPU backend (fp32 model, no QDQ wrapping) +// - QNN HTP backend (QDQ uint8 model, DQ -> MyAdd -> Q fusion) +// +// Before running, export the custom-op domain so the QNN EP factory registers +// the schema automatically: +// +// export ORT_QNN_CUSTOM_OP_DOMAINS="example:MyAdd" +// +// Build (after building ORT and the MyAdd op package): +// g++ -std=c++17 run_udo_sample.cc \ +// -I \ +// -L -lonnxruntime \ +// -Wl,-rpath, \ +// -o run_udo_sample +// +// Usage: +// # CPU backend +// LD_LIBRARY_PATH=/lib/x86_64-linux-clang: \ +// ./run_udo_sample cpu myadd_fp32.onnx /path/to/libMyAddOpPackage_cpu.so +// +// # HTP backend (x86 simulator or on-device) +// LD_LIBRARY_PATH=/lib/x86_64-linux-clang: \ +// ./run_udo_sample htp myadd_qdq.onnx /path/to/libMyAddOpPackage_htp.so + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "onnxruntime_cxx_api.h" + +static constexpr const char* kQnnEpName = "QNNExecutionProvider"; +static constexpr float kConstant = 2.0f; +static constexpr int kElements = 32; + +// Register the QNN EP plugin library and append it to session options using +// the v2 plugin API (RegisterExecutionProviderLibrary + AppendExecutionProvider_V2). +static void AppendQnnEp(Ort::Env& env, + Ort::SessionOptions& so, + const std::unordered_map& ep_options) { + // Register the QNN EP shared library with the environment. + env.RegisterExecutionProviderLibrary(kQnnEpName, "libonnxruntime_providers_qnn.so"); + + // Query all registered EP devices and filter to those from the QNN EP. + std::vector all_devices = env.GetEpDevices(); + std::vector qnn_devices; + for (const auto& dev : all_devices) { + if (std::string(dev.EpName()) == kQnnEpName) + qnn_devices.push_back(dev); + } + if (qnn_devices.empty()) + throw std::runtime_error("No QNN EP device found after registration."); + + so.AppendExecutionProvider_V2(env, qnn_devices, ep_options); +} + +static std::vector make_input() { + std::vector v(kElements); + for (int i = 0; i < kElements; ++i) + v[i] = -1.0f + 2.0f * i / (kElements - 1); + return v; +} + +static void run_cpu(const std::string& model_path, const std::string& pkg_path) { + printf("\n=== QNN CPU backend ===\n"); + + Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "UDOSampleCPU"); + Ort::SessionOptions so; + + // Register QNN EP and configure with op_packages. + // ORT_QNN_CUSTOM_OP_DOMAINS (set in the environment before this process started) + // handles the custom-domain schema registration automatically. + std::string op_packages = "MyAdd:" + pkg_path + ":MyAddOpPackageInterfaceProvider"; + AppendQnnEp(env, so, {{"backend_type", "cpu"}, {"op_packages", op_packages}}); + + Ort::Session session(env, model_path.c_str(), so); + + // Run inference with float32 input in [-1, 1]. + auto input = make_input(); + std::vector shape = {1, kElements}; + auto mem = Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault); + auto in_tensor = Ort::Value::CreateTensor(mem, input.data(), input.size(), + shape.data(), shape.size()); + const char* in_names[] = {"input"}; + const char* out_names[] = {"output"}; + auto outputs = session.Run(Ort::RunOptions{}, in_names, &in_tensor, 1, out_names, 1); + + // Verify output ≈ input + constant. + const float* out = outputs[0].GetTensorData(); + float max_err = 0.0f; + for (int i = 0; i < kElements; ++i) + max_err = std::max(max_err, std::fabs(out[i] - (input[i] + kConstant))); + + printf("Max absolute error vs (input + %.1f): %.2e\n", kConstant, max_err); + fputs(max_err <= 1e-4f ? "PASS\n" : "FAIL: error exceeds threshold\n", stdout); +} + +static void run_htp(const std::string& model_path, const std::string& pkg_path) { + printf("\n=== QNN HTP backend ===\n"); + + Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "UDOSampleHTP"); + Ort::SessionOptions so; + + // Register QNN HTP EP. + // ORT_QNN_CUSTOM_OP_DOMAINS handles schema registration automatically. + // The x86 HTP op package is the host-side half, so target it as CPU for + // host graph preparation. On-device execution registers the DSP skel + // separately with the HTP target (see the README's on-device flow). + std::string op_packages = "MyAdd:" + pkg_path + ":MyAddOpPackageInterfaceProvider:CPU"; + AppendQnnEp(env, so, {{"backend_type", "htp"}, {"offload_graph_io_quantization", "0"}, {"op_packages", op_packages}}); + + Ort::Session session(env, model_path.c_str(), so); + + // Run inference. + auto input = make_input(); + std::vector shape = {1, kElements}; + auto mem = Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault); + auto in_tensor = Ort::Value::CreateTensor(mem, input.data(), input.size(), + shape.data(), shape.size()); + const char* in_names[] = {"input"}; + const char* out_names[] = {"output"}; + auto outputs = session.Run(Ort::RunOptions{}, in_names, &in_tensor, 1, out_names, 1); + + // Verify within QDQ tolerance (~2 quantization steps for uint8 over [-1,1]). + const float* out = outputs[0].GetTensorData(); + float max_err = 0.0f; + for (int i = 0; i < kElements; ++i) + max_err = std::max(max_err, std::fabs(out[i] - (input[i] + kConstant))); + + // QDQ tolerance: 2x the output quantization scale (4/255, covering [0,4] output range). + const float qdq_tol = 4.0f / 255.0f * 2; + printf("Max absolute error vs (input + %.1f): %.4f (QDQ tol: %.4f)\n", + kConstant, max_err, qdq_tol); + fputs(max_err <= qdq_tol ? "PASS\n" : "FAIL: error exceeds QDQ tolerance\n", stdout); +} + +int main(int argc, char* argv[]) { + if (argc < 4) { + fprintf(stderr, + "Usage:\n" + " export ORT_QNN_CUSTOM_OP_DOMAINS=\"example:MyAdd\"\n" + " %s cpu \n" + " %s htp \n", + argv[0], argv[0]); + return 1; + } + + std::string backend = argv[1]; + std::string model = argv[2]; + std::string pkg = argv[3]; + + try { + if (backend == "cpu") + run_cpu(model, pkg); + else if (backend == "htp") + run_htp(model, pkg); + else { + fprintf(stderr, "Unknown backend '%s'. Use 'cpu' or 'htp'.\n", backend.c_str()); + return 1; + } + } catch (const Ort::Exception& e) { + fprintf(stderr, "ORT error: %s\n", e.what()); + return 1; + } catch (const std::exception& e) { + fprintf(stderr, "Error: %s\n", e.what()); + return 1; + } + + return 0; +} diff --git a/qcom/samples/qnn_udo_myadd/run_udo_sample.py b/qcom/samples/qnn_udo_myadd/run_udo_sample.py new file mode 100644 index 0000000000..63405de76f --- /dev/null +++ b/qcom/samples/qnn_udo_myadd/run_udo_sample.py @@ -0,0 +1,137 @@ +# Copyright (c) Qualcomm Technologies, Inc. and/or its subsidiaries. +# SPDX-License-Identifier: MIT +""" +Python reference sample: run MyAdd UDO on QNN EP. + +MyAdd computes: output = input + constant + input : shape [1, 32], float32 values in [-1, 1] + output : shape [1, 32], float32 + +Before running, export the custom-op domain so the QNN EP factory can register +the schema automatically (no separate schema library needed): + + export ORT_QNN_CUSTOM_OP_DOMAINS="example:MyAdd" + +Usage: + # QNN CPU backend (fp32 model) + LD_LIBRARY_PATH=/lib/x86_64-linux-clang: \\ + python run_udo_sample.py cpu myadd_fp32.onnx \\ + --op-package artifacts/libMyAddOpPackage_cpu.so \\ + --qnn-ep-lib + + # QNN HTP backend (QDQ model, x86 simulator or on-device) + LD_LIBRARY_PATH=/lib/x86_64-linux-clang: \\ + python run_udo_sample.py htp myadd_qdq.onnx \\ + --op-package artifacts/libMyAddOpPackage_htp.so \\ + --qnn-ep-lib + +Domain registration: + ORT_QNN_CUSTOM_OP_DOMAINS is read by the QNN EP factory at library-load time. + It registers a placeholder schema for the custom domain so ORT can load the + model. The actual QNN hardware kernel comes from the --op-package library. + +QNN EP registration (v2 plugin API): + The QNN EP is registered as a plugin library via + ort.register_execution_provider_library() + ort.get_ep_devices() + + so.add_provider_for_devices(), matching the C++ AppendExecutionProvider_V2 path. +""" + +import argparse +import sys + +import numpy as np + +import onnxruntime as ort + +CONSTANT = 2.0 +INPUT_SHAPE = (1, 32) +QNN_EP_NAME = "QNNExecutionProvider" + + +def build_input() -> np.ndarray: + n = INPUT_SHAPE[1] + return np.linspace(-1.0, 1.0, n, dtype=np.float32).reshape(INPUT_SHAPE) + + +def append_qnn_ep(so: ort.SessionOptions, qnn_ep_lib: str, ep_options: dict) -> None: + """Register the QNN EP plugin and append it to session options (v2 API). + + qnn_ep_lib must be an absolute path to libonnxruntime_providers_qnn.so. + """ + ort.register_execution_provider_library(QNN_EP_NAME, qnn_ep_lib) + devices = [d for d in ort.get_ep_devices() if d.ep_name == QNN_EP_NAME] + if not devices: + raise RuntimeError("No QNN EP device found after registration.") + so.add_provider_for_devices(devices, ep_options) + + +def run_cpu(model_path: str, op_package: str, qnn_ep_lib: str) -> None: + print("\n=== QNN CPU backend ===") + so = ort.SessionOptions() + + op_packages_str = f"MyAdd:{op_package}:MyAddOpPackageInterfaceProvider" + append_qnn_ep(so, qnn_ep_lib, {"backend_type": "cpu", "op_packages": op_packages_str}) + + sess = ort.InferenceSession(model_path, sess_options=so) + x = build_input() + [output] = sess.run(["output"], {"input": x}) + + expected = x + CONSTANT + max_err = float(np.max(np.abs(output - expected))) + print(f"Max absolute error vs (input + {CONSTANT}): {max_err:.2e}") + if max_err > 1e-4: + raise RuntimeError("error exceeds threshold") + print("PASS") + + +def run_htp(model_path: str, op_package: str, qnn_ep_lib: str) -> None: + print("\n=== QNN HTP backend ===") + so = ort.SessionOptions() + + op_packages_str = f"MyAdd:{op_package}:MyAddOpPackageInterfaceProvider:CPU" + append_qnn_ep( + so, + qnn_ep_lib, + { + "backend_type": "htp", + "offload_graph_io_quantization": "0", + "op_packages": op_packages_str, + }, + ) + + sess = ort.InferenceSession(model_path, sess_options=so) + x = build_input() + [output] = sess.run(["output"], {"input": x}) + + expected = x + CONSTANT + max_err = float(np.max(np.abs(output - expected))) + # QDQ tolerance: 2x the output quantization scale (4/255, covering [0,4] output range). + qdq_tol = 4.0 / 255.0 * 2 + print(f"Max absolute error vs (input + {CONSTANT}): {max_err:.4f} (QDQ tol: {qdq_tol:.4f})") + if max_err > qdq_tol: + raise RuntimeError("error exceeds QDQ tolerance") + print("PASS") + + +def main() -> int: + parser = argparse.ArgumentParser(description="Run MyAdd UDO on QNN EP") + parser.add_argument("backend", choices=["cpu", "htp"], help="QNN backend to use") + parser.add_argument("model", help="Path to ONNX model (myadd_fp32.onnx or myadd_qdq.onnx)") + parser.add_argument("--op-package", required=True, help="Path to libMyAddOpPackage_.so") + parser.add_argument("--qnn-ep-lib", required=True, help="Absolute path to libonnxruntime_providers_qnn.so") + args = parser.parse_args() + + try: + if args.backend == "cpu": + run_cpu(args.model, args.op_package, args.qnn_ep_lib) + else: + run_htp(args.model, args.op_package, args.qnn_ep_lib) + except RuntimeError as error: + print(f"FAIL: {error}", file=sys.stderr) + return 1 + + return 0 + + +if __name__ == "__main__": + sys.exit(main())