Skip to content

QuadratureTreeSHAP GPU: Remove the O(D²) partner scan in interaction values - #12653

Merged
RAMitchell merged 2 commits into
dmlc:masterfrom
ron-wettenstein:feature/qts-gpu-shared-path-features
Oct 6, 2026
Merged

RAMitchell merged 2 commits into
dmlc:masterfrom
ron-wettenstein:feature/qts-gpu-shared-path-features

Conversation

@ron-wettenstein

@ron-wettenstein ron-wettenstein commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

This PR is split out of #12652. See that PR for more details, including benchmark results against master for this change and the three related ones.

This PR include contribution 1 of #12652 :

Keep the features of the current path in shared memory (interaction values).
To compute interactions, the kernel repeatedly needs the distinct features on the current tree path. It used to rebuild this list at every step by rescanning the path in global memory, which costs O(D²) per step for a tree of depth D. It now keeps the list up to date in shared memory as it walks the tree, so the lookup costs O(D). On a T4, interaction values are about 1.1–1.2× faster for depth-6 trees and about 2× faster for depth-16 trees.

…ctions

The interaction kernel needs the distinct features on the current tree path at every return edge. Keep
each depth's split feature, and whether a deeper split on the same feature shadows it, in per-warp
shared memory, updated as the walk enters and leaves nodes. This replaces the O(D^2) rescan of the
path in global memory with an O(D) pass over shared memory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a3441f03-e04c-44ee-87ea-394e67a3a0ec
📥 Commits

Reviewing files that changed from the base of the PR and between 6958e39 and 231155f.

📒 Files selected for processing (2)
  • src/predictor/interpretability/shap.cu
  • tests/cpp/predictor/test_shap.cu

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


Walkthrough

GPU SHAP interaction traversal now tracks split features and whether earlier path entries are shadowed by repeated splits. Partner enumeration walks the active path and uses only unshadowed entries. The interaction kernel uses the expanded shared-state type and adds synchronization around traversal state reuse and changes. A new test compares CPU and GPU SHAP outputs for partial row tiles.

Suggested reviewers: ramitchell

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 23115

The investigated interaction traversal preserves the expected partner selection. No identified issue remains to resolve before normal merge checks.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Finish shared path reads before reusing node and stage slots, and synchronize stage reads before the warp leader advances the traversal. Apply the ordering to SHAP values and interactions, and cover full and partial row tiles with a CPU/GPU regression test suitable for Racecheck.
@RAMitchell

Copy link
Copy Markdown
Member

I fixed a shared memory race condition that was there before this PR.

@RAMitchell
RAMitchell merged commit 1b89f9e into dmlc:master Oct 6, 2026
84 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants