Sitelet https://github.com/apache/paimon/issues/10378
Skip to content

[Bug] Java Parquet reads drop matching rows when a filtered nested field is pruned from the read type #10378

Description

@thswlsqls

Search before asking

  • I searched in the issues and found nothing similar.

Paimon version
master @ b0af456

Compute Engine
Java API (ReadBuilder). Flink/Spark SQL are not affected.

Minimal reproduce step
Create a Parquet table id INT, s ROW<a INT, b INT> and write (0, {10, 7}), (1, {20, 8}). Read with a nested filter on s.b and a read type that prunes s to ROW<a>, without executeFilter():

Predicate filter = new PredicateBuilder(rowType)
        .equal(new NestedFieldTransform(new FieldRef(1, "s", sType), singletonList("b")), 7);
table.newReadBuilder().withFilter(filter).withReadType(/* id, s ROW<a> */);

What doesn't meet your expectations?
No rows are returned. Row 0 must be returned, as it is when the read type keeps s.b.

Anything else?
#9858 drops filters on unprojected columns in the ParquetReaderFactory constructor, but checks only top-level names. Since #9423 pushes nested predicates to Parquet, s.b = 7 passes because s is projected; Parquet reads the unrequested s.b as all-null and skips the row group. executeFilter() is safe because it widens the read type to the filter fields.

Are you willing to submit a PR?

  • I'm willing to submit a PR!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions