Skip to content

riscv: add SpacemiT IME2 matrix-extension path for Gemm (fp16) - #7032

Open
HougeLangley wants to merge 2 commits into
Tencent:masterfrom
HougeLangley:spacemit-ime2-gemm
Open

HougeLangley wants to merge 2 commits into
Tencent:masterfrom
HougeLangley:spacemit-ime2-gemm

Conversation

@HougeLangley

Copy link
Copy Markdown
Contributor

Summary

Add an optional fp16 GEMM path for the RISC-V Gemm layer that uses the
SpacemiT IME2 matrix extension (smt.vfwmadot, xsmtvdotii) found on the
A100 cluster of SpacemiT K3 SoCs (VLEN=1024).

smt.vfwmadot computes C[8x8]fp32 += A[8x8]fp16 x B^T[8x8]fp16, which is
exactly the transB=1 GEMM shape used by LLM decoders. On a SpacemiT K3
Pico-ITX board this gives a large, bit-exact speedup over the existing
RVV path.

What

  • New file src/layer/riscv/gemm_riscv_ime2.h: runtime probe (SIGILL-safe),
    A/B tile packing, and the vfwmadot micro-kernel (8x8 and 2x2 register-blocked).
  • gemm_riscv.{h,cpp} / gemm_riscv_zfh.cpp: an opt-in branch in
    create_pipeline() / forward() that selects the IME2 path only when the
    layer is a plain alpha * A * B^T (constantB, transB, no transposed output)
    with no C term, and the current CPU can execute smt.vfwmadot.
    Otherwise it falls back to the existing fp16 path unchanged.
  • New CMake option NCNN_RISCV_SPACEMIT_IME2 (default OFF), with a
    check_cxx_source_compiles probe for the xsmtvdotii extension. When the
    compiler does not support it, the option is force-disabled and the build is
    identical to before. When OFF, the new code is compiled out entirely.
  • cmake/ncnn_add_layer.cmake: the zfh+rvv variant march gains _xsmtvdotii
    only when the option is enabled.

Safety / fallback behaviour

  • The extension is gated both at compile time (CMake option + compiler
    probe) and at runtime (ime2_probe() executes the instruction under a
    SIGILL handler; on cores without it — e.g. the K3's X100 cluster — it returns
    false and the normal path is used).
  • The IME2 path handles alpha * A * B^T with an optional empty C. Any C term
    (runtime C blob or non-empty constant C with beta != 0) falls back to the
    standard path, which is still built for such layers.
  • Verified bit-exact against the existing path (see below).

Testing (SpacemiT K3, riscv64, gcc 17 + binutils 2.47)

Correctness:

  • tests/test_gemm* suite: identical pass/fail results to pristine master,
    on both the A100 and X100 clusters. (A few gemm tests already fail on A100 on
    pristine master — a pre-existing VLEN=1024 RVV issue unrelated to this PR.)
  • Byte-identical output vs the existing RVV path on a real LLM
    (Qwen3-0.6B, greedy decoding, 64 tokens).
  • Standalone kernel verified against a scalar reference: max relative error 0.

Performance (Qwen3-0.6B fp16, 8 A100 threads):

path prefill tok/s decode tok/s
RVV (A100) 12.9 5.1
RVV (X100, best non-IME cluster) 30.6 4.9
IME2 (A100) 83.0 10.1

Notes / limitations

  • The matrix unit only exists on the A100 cluster (VLEN=1024); the kernel is a
    no-op elsewhere. ncnn's weight packing is VLEN-dependent (packn = vlenb/4),
    so a process must stay pinned to one cluster — this is a pre-existing ncnn
    constraint, not introduced here.
  • Only the fp16 vfwmadot path is implemented. The int8 vmadot path (works on
    both X100 IME1 and A100 IME2) is a possible follow-up for quantized models.

Add an optional fp16 GEMM path for the RISC-V Gemm layer using the
SpacemiT IME2 matrix extension (smt.vfwmadot, xsmtvdotii) found on the
A100 cluster of SpacemiT K3 SoCs (VLEN=1024).

smt.vfwmadot computes C[8x8]fp32 += A[8x8]fp16 x B^T[8x8]fp16, exactly
the transB=1 shape used by LLM decoders.

The path is selected only for a plain alpha*A*B^T with no C term and
only when the current CPU can execute the instruction (a runtime
SIGILL-safe probe); otherwise it falls back to the existing fp16 path,
which is still built for layers that may carry a C term.

Gated by new CMake option NCNN_RISCV_SPACEMIT_IME2 (default OFF) with a
compiler support probe; when OFF or unsupported the code is compiled out
and the build is identical to before.

Tested on SpacemiT K3: identical pass/fail to pristine master across the
gemm test suite, and byte-identical output on Qwen3-0.6B. Measured
~6.4x prefill / ~2x decode vs the RVV path on the same A100 cluster.
@tencent-adm

Copy link
Copy Markdown
Member

CLA assistant check
Thank you for your submission, we really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants