Sitelet https://github.com/humemai/arcadedb-embedded-python/pull/132
Skip to content

Columnar readers keep every column, a stable Arrow type, and exact DECIMALs (#113, #114, #115) - #132

Merged
tae898 merged 1 commit into
mainfrom
fix-columns
Oct 3, 2026
Merged

tae898 merged 1 commit into
mainfrom
fix-columns

Conversation

@tae898

@tae898 tae898 commented Oct 3, 2026

Copy link
Copy Markdown

Fixes #113, #114, and #115 (the columnar readers to_columns, to_dataframe, and to_arrow, and the header of export_to_csv). Like #131 this changes the Java bridge (ColumnBatcher), so the tests need the bridge jar built from this branch, which the wheel build does.

#113, a property the first row lacks was dropped. ColumnBatcher took the column set from the first row, and the Python side pinned the first batch's set for every later batch. Each batch now reports the union of its rows' property names in order of first appearance, and the Python side takes the union over the batches in order of first appearance, filling a batch that lacks a column with nulls in the shape that column has elsewhere (NaN, NaT, or None; pa.nulls for Arrow). The three readers return n, kind, reason for the issue's three events at every batch size. export_to_csv takes its header from the union of the first batch's rows (and of every row for a list of dicts); a column that first appears after the first batch of 10,000 rows now raises ArcadeDBError naming it and saying to pass fieldnames, instead of the writer's bare "dict contains fields not in fieldnames" (the header is already written by then, so the partial file remains).

The union costs one pass over each row's property names, because Result.getPropertyNames() builds a set per row. On a 200,000-row, twelve-property to_columns() I measured about 25% (medians 0.91 to 1.01 s on the old bridge jar against 1.15 to 1.24 s on the new, three alternating rounds pinned to cores 0 to 3 of the laptop; relative only, the laptop was running other work). So to_columns() and to_arrow() take an optional columns=[...] that reads exactly those columns as a projection would and skips the pass; a row lacking one reads null and a property not listed is left out.

#114, an Arrow column's type followed the batch size. A chunk that says nothing about its column's type (every value null, or a list<null> from only empty lists) now takes the type of the other chunks instead of forcing the whole column to strings; level is int64 and tags is list<string> at batch sizes 25,000, 2, and 1 for the issue's five readings. A column that mixes types inside one batch (an int in one row, a string in the next) raised a raw ArrowInvalid; it is now the string column the cross-batch case already produced.

#115, a DECIMAL column came back as JSON numbers. BigDecimal gets its own column type in ColumnBatcher (the string layout, each value written with toPlainString), decoded to Decimal objects: an object array for to_columns() and to_dataframe(), a decimal128 Arrow column (decimal256 above 38 digits, an exact string above 76), with differing precision and scale across batches unified to one decimal type. The issue's four cases (36 digits, whole amounts only, whole and fractional, a whole amount above 2**63) plus one with a null round-trip exactly at batch sizes 25,000 and 1.

The tests are in tests/test_columnar_readers.py (31) and are described in the testing docs. Verified: 29 of the 31 fail on the bridge jar and Python of main (the two that pass are 25,000-row cases that already worked) and all pass on this branch, the existing arrow, resultset, and exporter tests pass, and the full suite passes (540 passed, 3 skipped, 3 xfailed), bandit clean at CI strictness.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

…rrow type, and exact DECIMALs

to_columns, to_dataframe, and to_arrow took the column set from the first row
and pinned the first batch's set for the rest, so a property the first row
lacked was dropped. Each batch now reports the union of its rows' names and the
readers take the union over the batches in order of first appearance, with
nulls for a batch that lacks a column. Finding the names costs one pass over
each row's names (about 25% on a 200k x 12 to_columns), so to_columns and
to_arrow take an optional columns= that reads exactly those columns. The CSV
export's header is the union of the first batch, and a column that first
appears later raises an ArcadeDBError that names it.

An Arrow column's type no longer follows the batch size: a chunk that is all
null or holds only empty lists takes the type of the other chunks, and a column
that mixes types inside one batch becomes strings instead of raising
ArrowInvalid.

BigDecimal has its own column type in ColumnBatcher, decoded to Decimal objects
(an object array) and to a decimal Arrow column, so no digit is lost to a
double, the dtype no longer follows the data, and a value above 2**63 no longer
overflows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

to_columns, to_dataframe, and to_arrow drop every property that the first row lacks, and export_to_csv raises after writing part of the file

1 participant