riscv: add SpacemiT IME2 matrix-extension path for Gemm (fp16) - #7032
Open
HougeLangley wants to merge 2 commits into
Open
HougeLangley wants to merge 2 commits into
HougeLangley wants to merge 2 commits into
Conversation
Add an optional fp16 GEMM path for the RISC-V Gemm layer using the SpacemiT IME2 matrix extension (smt.vfwmadot, xsmtvdotii) found on the A100 cluster of SpacemiT K3 SoCs (VLEN=1024). smt.vfwmadot computes C[8x8]fp32 += A[8x8]fp16 x B^T[8x8]fp16, exactly the transB=1 shape used by LLM decoders. The path is selected only for a plain alpha*A*B^T with no C term and only when the current CPU can execute the instruction (a runtime SIGILL-safe probe); otherwise it falls back to the existing fp16 path, which is still built for layers that may carry a C term. Gated by new CMake option NCNN_RISCV_SPACEMIT_IME2 (default OFF) with a compiler support probe; when OFF or unsupported the code is compiled out and the build is identical to before. Tested on SpacemiT K3: identical pass/fail to pristine master across the gemm test suite, and byte-identical output on Qwen3-0.6B. Measured ~6.4x prefill / ~2x decode vs the RVV path on the same A100 cluster.
Member
|
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add an optional fp16 GEMM path for the RISC-V
Gemmlayer that uses theSpacemiT IME2 matrix extension (
smt.vfwmadot,xsmtvdotii) found on theA100 cluster of SpacemiT K3 SoCs (VLEN=1024).
smt.vfwmadotcomputesC[8x8]fp32 += A[8x8]fp16 x B^T[8x8]fp16, which isexactly the
transB=1GEMM shape used by LLM decoders. On a SpacemiT K3Pico-ITX board this gives a large, bit-exact speedup over the existing
RVV path.
What
src/layer/riscv/gemm_riscv_ime2.h: runtime probe (SIGILL-safe),A/B tile packing, and the
vfwmadotmicro-kernel (8x8 and 2x2 register-blocked).gemm_riscv.{h,cpp}/gemm_riscv_zfh.cpp: an opt-in branch increate_pipeline()/forward()that selects the IME2 path only when thelayer is a plain
alpha * A * B^T(constantB, transB, no transposed output)with no C term, and the current CPU can execute
smt.vfwmadot.Otherwise it falls back to the existing fp16 path unchanged.
NCNN_RISCV_SPACEMIT_IME2(default OFF), with acheck_cxx_source_compilesprobe for thexsmtvdotiiextension. When thecompiler does not support it, the option is force-disabled and the build is
identical to before. When OFF, the new code is compiled out entirely.
cmake/ncnn_add_layer.cmake: the zfh+rvv variant march gains_xsmtvdotiionly when the option is enabled.
Safety / fallback behaviour
probe) and at runtime (
ime2_probe()executes the instruction under aSIGILL handler; on cores without it — e.g. the K3's X100 cluster — it returns
false and the normal path is used).
alpha * A * B^Twith an optional empty C. Any C term(runtime C blob or non-empty constant C with beta != 0) falls back to the
standard path, which is still built for such layers.
Testing (SpacemiT K3, riscv64, gcc 17 + binutils 2.47)
Correctness:
tests/test_gemm*suite: identical pass/fail results to pristine master,on both the A100 and X100 clusters. (A few gemm tests already fail on A100 on
pristine master — a pre-existing VLEN=1024 RVV issue unrelated to this PR.)
(Qwen3-0.6B, greedy decoding, 64 tokens).
Performance (Qwen3-0.6B fp16, 8 A100 threads):
Notes / limitations
no-op elsewhere. ncnn's weight packing is VLEN-dependent (
packn = vlenb/4),so a process must stay pinned to one cluster — this is a pre-existing ncnn
constraint, not introduced here.
vfwmadotpath is implemented. The int8vmadotpath (works onboth X100 IME1 and A100 IME2) is a possible follow-up for quantized models.