Skip to content

Read Iceberg v2 position delete files in DataFusion scans - #412

Merged
JanKaul merged 3 commits into
JanKaul:mainfrom
Embucket:upstream-position-delete-read
Sep 24, 2026
Merged

JanKaul merged 3 commits into
JanKaul:mainfrom
Embucket:upstream-position-delete-read

Conversation

@osipovartem

@osipovartem osipovartem commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • read Iceberg v2 Parquet position-delete files in datafusion_iceberg scans
  • dispatch Content::PositionDeletes by content_offset: v3 Puffin deletion vectors keep the Read Iceberg v3 deletion vectors in DataFusion scans #414 loader, while v2 Parquet delete files use a new streaming loader
  • merge v2 positions into the shared path-keyed roaring-bitmap index and reuse IcebergDvExec
  • apply the Iceberg sequence rule while building the index (delete sequence >= data sequence, including equality)
  • keep internal data-path and Parquet row-number columns collision-free and out of user results

This implements the v2 read side described by the Iceberg delete files specification and converges it with the v3 deletion-vector path introduced by #414. Equality and positional deletes in the same partition remain an explicit follow-up, matching #414's current behavior.

Performance

  • data files without positional deletes keep the existing scan path
  • delete files are read with bounded concurrency and streamed in record batches
  • only the reserved file_path and pos fields are projected by Iceberg field ID
  • rows are filtered by the shared vectorized Arrow bitmap operator; there is no sort, repartition, join, or per-row ScalarValue conversion
  • query predicates and row-group pruning remain enabled through the true Parquet RowNumber virtual column

Validation

  • cargo +1.95.0 test -p iceberg-rust table::position_delete::tests --lib
  • cargo +1.95.0 test -p iceberg-rust table::deletion_vector --lib
  • cargo +1.95.0 test -p datafusion_iceberg --test position_delete
  • cargo +1.95.0 test -p datafusion_iceberg table::tests::row_number_virtual_column_drives_dv_filter_with_pushdown --lib
  • cargo +1.95.0 fmt --all -- --check
  • git diff --check

The integration test covers multiple v2 delete files, duplicate positions, field-ID-based projection with noncanonical column names, filtered reads, later appends, and collisions with internal metadata column names. Focused tests cover inclusive sequence semantics and the existing v3 DV path.

@JanKaul

JanKaul commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Convergence with #414 (v3 deletion vectors)

#414 adds read support for v3 deletion vectors and, in doing so, introduces infrastructure this PR can reuse. Both features apply positional deletes keyed by (data_file_path, absolute_row_position), so once #414 merges there's a clean path to land v2 and v3 on a single operator. Sharing the plan here so we can coordinate — nothing to change in this PR right now.

What #414 provides that this PR can reuse

  • The Parquet RowNumber virtual column (this PR uses it too) — the shared way to get a row's true absolute file position, correct under row-group pruning and repartitioning.
  • IcebergDvExec: a streaming physical node that filters rows by probing a per-file roaring bitmap at each row's RowNumber. It consumes HashMap<data_file_path, DeletionVector> and preserves scan ordering; files with no delete are untouched.
  • Eager per-file delete-index loading.

Proposed integration (after #414 merges)

  1. Dispatch Content::PositionDeletes by content_offset: Some → v3 DV (Puffin), None → v2 position-delete file. This is exactly the guard this PR already has (the not_impl_err! on content_offset), turned from an error into a router.
  2. Add load_position_deletes: read each v2 delete file's (file_path, pos) rows (spec field-ids 2147483546 / 2147483545, already encoded here) into a HashMap<data_file_path, RoaringTreemap>, applying the sequence-number rule (delete.seq >= data.seq, equal applies — the same semantics as position_delete_sequence_filter) at build time.
  3. Merge into the shared index and let IcebergDvExec filter both v2 and v3.

What that reuses from this PR: the content_offset dispatch, the sequence-number applicability rule, and the v2 delete-file schema / reserved field-ids. unique_internal_column_name is also worth upstreaming into the shared path (the DV branch currently hardcodes its internal column names).

What it would replace: the SortMergeJoin + Repartition + Sort pipeline, in favor of the streaming bitmap probe. Trade-off: the index materializes deleted positions in memory (compact via roaring, and it's the small side), while the join spills the data side; the index also preserves scan ordering and per-file skip. It further removes the need for the sequence-number scan columns, since the rule is applied at index-build time rather than per-row in a JoinFilter.

Scope note: coexistence of equality + positional deletes in one partition would need the equality path to also emit the internal columns and wrap in IcebergDvExec — worth treating as a follow-up.

Happy to help wire this up once #414 lands.

@osipovartem

Copy link
Copy Markdown
Contributor Author

Thanks, this convergence plan makes sense. I agree that v2 position-delete files and v3 deletion vectors should share the streaming bitmap-probe operator once #414 lands. I will keep this PR unchanged for now, then rebase it onto #414 and adapt the v2 loader while preserving the inclusive sequence rule, reserved field IDs, and collision-free internal column names. Equality/position-delete coexistence can remain an explicit follow-up.

@osipovartem
osipovartem force-pushed the upstream-position-delete-read branch 4 times, most recently from edc6a77 to bc4ce9d Compare September 23, 2026 13:42
@osipovartem
osipovartem force-pushed the upstream-position-delete-read branch from bc4ce9d to f854c8c Compare September 23, 2026 13:44
@JanKaul

JanKaul commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Looks great! Thanks a lot

@JanKaul
JanKaul merged commit 75c45df into JanKaul:main Sep 24, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants