Skip to content

GH-3574: Preserve null counts when min/max exceed limit - #3819

Open
efegokdemir wants to merge 2 commits into
apache:masterfrom
efegokdemir:codex/issue-3574-preserve-null-count
Open

efegokdemir wants to merge 2 commits into
apache:masterfrom
efegokdemir:codex/issue-3574-preserve-null-count

Conversation

@efegokdemir

Copy link
Copy Markdown

Summary

Preserve null_count when a column's min/max statistics exceed the metadata size limit.

The size guard is needed for min/max values, but null_count is a compact independent statistic and remains useful to readers. This keeps the null count available without reintroducing oversized min/max metadata.

Changes

  • Emit null_count before applying the min/max size limit.
  • Keep nan count and min/max omission behavior unchanged for oversized statistics.
  • Extend the binary statistics round-trip test to cover the preserved null count.

Testing

  • ./mvnw -pl parquet-hadoop -am -DskipITs -Dtest=TestParquetMetadataConverter -Dsurefire.failIfNoSpecifiedTests=false test
  • ./mvnw -pl parquet-hadoop -am -DskipTests spotless:check
  • git diff --check

All passed.

Notes

Fixes #3574

Signed-off-by: Efe Gökdemir <efe@rexcode.co.uk>

@divjotarora divjotarora left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @efegokdemir!

@divjotarora divjotarora left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@efegokdemir My previous comment was not addressed, the current code still sets NaN count stats after the withinLimit check. I suggest we set null + NaN stats unconditionally and guard only min/max by the withinLimit check.

Proposed code:

formatStats.setNull_count(stats.getNumNulls());
if (stats.isNanCountSet()) {
    formatStats.setNan_count(stats.getNanCount());
}

if (!withinLimit(stats, truncateLength)) {
    return formatStats;
}

Signed-off-by: Efe Gökdemir <gokdemirefe1903@gmail.com>
@efegokdemir

Copy link
Copy Markdown
Author

Moved nan_count alongside null_count before the withinLimit early return; min/max remain size-gated. Added testNanCountIsPreservedWhenMinMaxExceedLimit, which forces the floating-point stats size check over the limit and verifies null count, NaN count, and omitted min/max. The pushed remote HEAD is d015ab76dd7d2c9bf79f5795d971ee1c3a568f23, and I verified both changed files at that SHA. git diff --check passed. Focused Maven/Spotless validation could not run: the wrapper reports “Unable to locate a Java Runtime” in this environment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

null_count is omitted for large columns in parquet files

2 participants