Columnar readers keep every column, a stable Arrow type, and exact DECIMALs (#113, #114, #115) - #132
Merged
Merged
Conversation
…rrow type, and exact DECIMALs to_columns, to_dataframe, and to_arrow took the column set from the first row and pinned the first batch's set for the rest, so a property the first row lacked was dropped. Each batch now reports the union of its rows' names and the readers take the union over the batches in order of first appearance, with nulls for a batch that lacks a column. Finding the names costs one pass over each row's names (about 25% on a 200k x 12 to_columns), so to_columns and to_arrow take an optional columns= that reads exactly those columns. The CSV export's header is the union of the first batch, and a column that first appears later raises an ArcadeDBError that names it. An Arrow column's type no longer follows the batch size: a chunk that is all null or holds only empty lists takes the type of the other chunks, and a column that mixes types inside one batch becomes strings instead of raising ArrowInvalid. BigDecimal has its own column type in ColumnBatcher, decoded to Decimal objects (an object array) and to a decimal Arrow column, so no digit is lost to a double, the dtype no longer follows the data, and a value above 2**63 no longer overflows. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #113, #114, and #115 (the columnar readers
to_columns,to_dataframe, andto_arrow, and the header ofexport_to_csv). Like #131 this changes the Java bridge (ColumnBatcher), so the tests need the bridge jar built from this branch, which the wheel build does.#113, a property the first row lacks was dropped.
ColumnBatchertook the column set from the first row, and the Python side pinned the first batch's set for every later batch. Each batch now reports the union of its rows' property names in order of first appearance, and the Python side takes the union over the batches in order of first appearance, filling a batch that lacks a column with nulls in the shape that column has elsewhere (NaN, NaT, or None;pa.nullsfor Arrow). The three readers returnn, kind, reasonfor the issue's three events at every batch size.export_to_csvtakes its header from the union of the first batch's rows (and of every row for a list of dicts); a column that first appears after the first batch of 10,000 rows now raisesArcadeDBErrornaming it and saying to passfieldnames, instead of the writer's bare "dict contains fields not in fieldnames" (the header is already written by then, so the partial file remains).The union costs one pass over each row's property names, because
Result.getPropertyNames()builds a set per row. On a 200,000-row, twelve-propertyto_columns()I measured about 25% (medians 0.91 to 1.01 s on the old bridge jar against 1.15 to 1.24 s on the new, three alternating rounds pinned to cores 0 to 3 of the laptop; relative only, the laptop was running other work). Soto_columns()andto_arrow()take an optionalcolumns=[...]that reads exactly those columns as a projection would and skips the pass; a row lacking one reads null and a property not listed is left out.#114, an Arrow column's type followed the batch size. A chunk that says nothing about its column's type (every value null, or a
list<null>from only empty lists) now takes the type of the other chunks instead of forcing the whole column to strings;levelisint64andtagsislist<string>at batch sizes 25,000, 2, and 1 for the issue's five readings. A column that mixes types inside one batch (an int in one row, a string in the next) raised a rawArrowInvalid; it is now the string column the cross-batch case already produced.#115, a DECIMAL column came back as JSON numbers.
BigDecimalgets its own column type inColumnBatcher(the string layout, each value written withtoPlainString), decoded toDecimalobjects: an object array forto_columns()andto_dataframe(), adecimal128Arrow column (decimal256above 38 digits, an exact string above 76), with differing precision and scale across batches unified to one decimal type. The issue's four cases (36 digits, whole amounts only, whole and fractional, a whole amount above 2**63) plus one with a null round-trip exactly at batch sizes 25,000 and 1.The tests are in
tests/test_columnar_readers.py(31) and are described in the testing docs. Verified: 29 of the 31 fail on the bridge jar and Python ofmain(the two that pass are 25,000-row cases that already worked) and all pass on this branch, the existing arrow, resultset, and exporter tests pass, and the full suite passes (540 passed, 3 skipped, 3 xfailed), bandit clean at CI strictness.🤖 Generated with Claude Code
https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL