Search before asking
Paimon version
master @ b0af456
Compute Engine
Java API (ReadBuilder). Flink/Spark SQL are not affected.
Minimal reproduce step
Create a Parquet table id INT, s ROW<a INT, b INT> and write (0, {10, 7}), (1, {20, 8}). Read with a nested filter on s.b and a read type that prunes s to ROW<a>, without executeFilter():
Predicate filter = new PredicateBuilder(rowType)
.equal(new NestedFieldTransform(new FieldRef(1, "s", sType), singletonList("b")), 7);
table.newReadBuilder().withFilter(filter).withReadType(/* id, s ROW<a> */);
What doesn't meet your expectations?
No rows are returned. Row 0 must be returned, as it is when the read type keeps s.b.
Anything else?
#9858 drops filters on unprojected columns in the ParquetReaderFactory constructor, but checks only top-level names. Since #9423 pushes nested predicates to Parquet, s.b = 7 passes because s is projected; Parquet reads the unrequested s.b as all-null and skips the row group. executeFilter() is safe because it widens the read type to the filter fields.
Are you willing to submit a PR?
Search before asking
Paimon version
master @ b0af456
Compute Engine
Java API (
ReadBuilder). Flink/Spark SQL are not affected.Minimal reproduce step
Create a Parquet table
id INT, s ROW<a INT, b INT>and write(0, {10, 7}),(1, {20, 8}). Read with a nested filter ons.band a read type that prunesstoROW<a>, withoutexecuteFilter():What doesn't meet your expectations?
No rows are returned. Row
0must be returned, as it is when the read type keepss.b.Anything else?
#9858 drops filters on unprojected columns in the
ParquetReaderFactoryconstructor, but checks only top-level names. Since #9423 pushes nested predicates to Parquet,s.b = 7passes becausesis projected; Parquet reads the unrequesteds.bas all-null and skips the row group.executeFilter()is safe because it widens the read type to the filter fields.Are you willing to submit a PR?