Skip to content

Reuse a runtime-descriptor LTO driver for multi-output AST JIT - #24407

Draft
vyasr wants to merge 5 commits into
NVIDIA:mainfrom
vyasr:codex/ast-runtime-lto
Draft

vyasr wants to merge 5 commits into
NVIDIA:mainfrom
vyasr:codex/ast-runtime-lto

Conversation

@vyasr

@vyasr vyasr commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Description

AST (expression-tree) JIT evaluation currently compiles generated expression code together with the transform driver that traverses rows and handles inputs and outputs. Changing the expression or its types can therefore require recompiling driver code as well, even though the row traversal is largely reusable.

This PR separates compatible multi-output evaluations into two compiled fragments: the generated expression operation and a schema-independent driver. LTO (link-time optimization) links those fragments into an executable GPU kernel. The driver receives runtime descriptors containing column pointers, null masks, offsets, and scalar information rather than specializing its interface for each input/output schema. This allows its compiled fragment to be reused across different expressions and supported schemas.

  • Use the new path for BOOL8, integer, and floating-point inputs/outputs. Keep source JIT for single-output evaluation and unsupported representations, including decimals and temporal types, whose logical semantics need more than the supported storage-type interface.
  • Pack aligned input/output descriptors into one asynchronous transfer, using the current memory resource for scratch storage.
  • Cache linked kernels by generated source, embedded bundle, CUDA runtime/driver, target SM, and LTO architecture. Resolve fragments only on a linked-library miss, avoiding that work when the executable kernel is already cached.
  • Preserve expression lowering through Row IR, result ordering, null/error semantics, output allocation, and RTCX cache behavior without changing public APIs.
  • Add focused coverage for per-type overflow errors, nullable grid-stride execution and partial warps, linked-kernel cache identity, sliced inputs, scalars, and empty inputs. Existing arithmetic and decimal tests continue to cover ordinary/NULLIFY results and source-JIT fallback.

Checklist

  • I am familiar with the Contributing Guidelines.
  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

@vyasr vyasr added tests Unit testing for project libcudf Affects libcudf (C++/CUDA) code. Performance Performance related issue improvement Improvement / enhancement to an existing function non-breaking Non-breaking change labels Oct 1, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@vyasr

vyasr commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

@vyasr vyasr added the 2 - In Progress Currently a work in progress label Oct 1, 2026
@vyasr

vyasr commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

1 similar comment
@vyasr

vyasr commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

vyasr added 5 commits October 10, 2026 17:17
Keep source JIT for single-output and representation-incompatible types. Cache the linked AST kernel by source identity, resolve fragments only on misses, and transfer runtime descriptors together. Add self-contained integer, nullable, throwing, scalar, sliced, and decimal-fallback coverage without depending on the test-batching PR.
@vyasr
vyasr force-pushed the codex/ast-runtime-lto branch from 571fbe2 to c6fa2f1 Compare October 11, 2026 01:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2 - In Progress Currently a work in progress improvement Improvement / enhancement to an existing function libcudf Affects libcudf (C++/CUDA) code. non-breaking Non-breaking change Performance Performance related issue tests Unit testing for project

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant