fix(upsert): make the match filter flat for composite keys - #4027
Open
azwanzuharimi wants to merge 3 commits into
Open
azwanzuharimi wants to merge 3 commits into
azwanzuharimi wants to merge 3 commits into
Conversation
An upsert with two or more join columns built one Or disjunct per key. PyArrow flattens that chain and overflows its stack from about 1,000 keys. The scan, the insert filter and the overwrite filter all crashed. Build one In per join column instead. Do the exact key match in Arrow with an anti join. Append back the rows that the overwrite filter removes but does not update. Closes apache#3508
azwanzuharimi
force-pushed
the
upsert-flat-match-filter
branch
from
September 29, 2026 08:34
0c60294 to
fbcc6fd
Compare
An empty table gave the index column the null type and the anti join rejected it. The insert path reaches this when one scan batch removes every source key and a later batch still matches the filter.
Match the check that get_rows_to_update already does, so a join column named __index gives a clear error instead of an Arrow join failure.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3508
Rationale for this change
Table.upsertwith two or more join columns builds oneOr(And(EqualTo, ...))disjunct per key. PyArrow flattens that chain and recurses over it. From about 1,000 keys the C++ stack overflows and the process dies without a Python traceback. All three uses of the filter crash: the scan for matching rows, the insert filter and the overwrite filter. A balanced Or tree does not help, because PyArrow flattens it again.This change keeps every filter flat:
create_match_filterreturns oneInper join column. For a composite key this filter can match more rows than the keys in the source.upsert_util.exclude_keysdoes the exact key match in Arrow with an anti join. The insert path uses it instead of an Arrow expression filter.Related work: #3509 groups keys by prefix and stays exact, but it still emits one disjunct per distinct prefix. #3420 bounds the scan filter, but the insert and overwrite paths keep the exact filter.
Results, measured locally on macOS with pyarrow 25.0.1. Target table 100,000 rows, upsert 20,000 rows (10,000 updates and 10,000 inserts). The check reads the table back and compares every row. rc 138 is a stack overflow crash. The script is available on request.
Case d rejects null keys on every branch. This PR does not change that.
Cost for a composite key: the overwrite reads the affected files one extra time. Rows that share a key column value with an updated key, but are not updated, are rewritten unchanged. In the worst case the filter matches every combination of the updated key values, for example every date times every id. Single column keys are not changed.
Are these changes tested?
Yes.
tests/table/test_upsert.pyadds:test_create_match_filter_composite_key_is_flattest_upsert_composite_key_keeps_rows_outside_source_keys: every key column value of the source exists in the target, but only some key tuples do. This guards against deleting too much.test_upsert_composite_key_large_batch: 5,000 key tuples. On main this crashes pytest with exit code 132.test_upsert_composite_key_all_source_keys_matched_across_files: the first data file removes every source key from the insert set, a later file still matches the filter.test_exclude_keys_rejects_reserved_column_nameAll 4239 unit tests pass.
Are there any user-facing changes?
Yes. An upsert on a composite key with more than about 1,000 keys no longer crashes. An upsert on a composite key can add one extra APPEND snapshot when the overwrite filter removes rows that are not updated. No API change.