Repository navigation
fix: repair native vector writes and data evolution read ranges - #1090
Conversation
leaves12138
left a comment
There was a problem hiding this comment.
Reviewed 5d02d3e, using the Python counterpart apache/paimon#10482 at 51f2c376. No blocking findings.
The FixedSizeList normalization preserves Java's positional vector elements and dimension checks without relaxing element types or visible NOT NULL constraints. Parent NULL masking and sliced payload-buffer reuse remain correct. The additional reader change correctly scopes vector selection to each prepared merge group; unavailable vector factories are rejected before writer construction can accept data.
Independent validation, with Rust 1.94.1, Python 3.10, PyArrow 19.0.1 and a freshly rebuilt extension:
- Core: 3,918 passed, 6 ignored; formatting check passed.
- Python bindings: 417 passed, using a Native-generated smoke warehouse rather than Java Docker fixtures.
- Reviewer cases: 115 input/NULL/slice/validation cases and 10 combined Python-planned split cases with unequal group sizes, FLOAT/DOUBLE, and full/partial/empty selections passed.
- Related Python suites: 969 passed, 7 skipped, 4 subtests passed, with all five Native CI flags enabled. Plan/read/write/commit counters and all six update-kind counters were positive. The skips require Python >= 3.11 for Vortex.
I also reproduced the original child-field schema rejection against an older extension, and the later-group vector NULL-fill bug against the previous normalization-only wheel; the rebuilt head passes both scenarios.
Validation boundaries: Java alignment was checked against source, not a Java/Python interoperability run or production storage. NULL-vector Classic reads encounter an independently reproduced pure-PyArrow 19 Parquet issue; Native values were checked instead. Nested required-element cases were read directly through the Rust binding because the Python target-schema parser already loses vector-element nullability. These baseline issues are not regressions introduced here.
The Python counterpart still installs Rust main; its complete Native CI needs to be rerun after this fix merges.
Purpose
Fix the native writer failure reported by apache/paimon#10482. Its Native CI run has 149 failures and 195 setup errors, all caused by rejecting otherwise valid Arrow FixedSizeList inputs for VECTOR fields.
The shared input normalizer already canonicalizes ARRAY and MAP child aliases, but omits FixedSizeList. PyArrow's child field differs from Paimon's
elementfield, so vectors fail schema validation before writing.Changes
Validation
The Python counterpart continues to install Rust main. This fix must merge before its Native CI can pass.
Additional failures uncovered after normalization
Final verification
Python follow-up commit: apache/paimon#10482. Its CI continues to use Rust main and therefore depends on this PR merging.