Level: bindings. Two defects in ResultSet.to_arrow. First, when one batch of a column holds only NULLs (or only empty lists), _unify_arrow_chunk_types gives up on the column and casts every batch to string. A column with the values 10, 20, NULL, NULL, and 50 is int64 at the default batch_size, and is string with the values '10', '20', None, None, and '50' at batch_size=2. A list column becomes the strings "['a']", '[]', and so on at batch_size=1. So the schema of the table depends on a tuning parameter and on where the NULLs fall. Second, a column that mixes types inside one batch (an int in one row and a str in another, which a schemaless document allows) makes to_arrow() raise pyarrow.lib.ArrowInvalid, a raw pyarrow error, after the result set was consumed, while to_list, to_columns, and to_dataframe return the values.
Repro
The script prints to stdout. The JVM's own log lines go to stderr and are discarded by the command below.
uv venv .venv
uv pip install --python .venv/bin/python arcadedb_embedded-26.10.1.dev0-cp312-cp312-manylinux_2_34_x86_64.whl numpy pandas pyarrow
.venv/bin/python repro-7-to-arrow-types.py 2>/dev/null
"""to_arrow: the type of a column depends on batch_size, and a column of mixed types raises a raw pyarrow error."""
import os
import tempfile
import arcadedb_embedded as arcade
from arcadedb_embedded.jvm import start_jvm
start_jvm(heap_size="2g")
db = arcade.create_database(os.path.join(tempfile.mkdtemp(prefix="hb7-"), "db"))
db.command("sql", "CREATE DOCUMENT TYPE Reading")
with db.transaction():
# rows 3 and 4 carry no "level" at all (a legal schemaless document), rows 1, 2 and 5 carry an int
for n, level, tags in [(1, 10, ["a"]), (2, 20, []), (3, None, ["b", "c"]), (4, None, []), (5, 50, ["d"])]:
if level is None:
db.command("sql", "INSERT INTO Reading SET n = ?, tags = ?", n, tags)
else:
db.command("sql", "INSERT INTO Reading SET n = ?, level = ?, tags = ?", n, level, tags)
Q = "SELECT FROM Reading ORDER BY n"
for batch_size in (25_000, 2, 1):
t = db.query("sql", Q).to_arrow(batch_size=batch_size)
print(f"to_arrow(batch_size={batch_size:>6}):")
for col in ("level", "tags"):
print(f" {col:<6} {str(t.schema.field(col).type):<12} {t.column(col).to_pylist()}")
cols = db.query("sql", Q).to_columns(batch_size=batch_size)
print(f"to_columns(batch_size={batch_size:>6}): level is a {type(cols['level']).__name__}: {list(cols['level'])}")
print()
with db.transaction():
db.command("sql", "INSERT INTO Reading SET n = 6, mixed = 1")
db.command("sql", "INSERT INTO Reading SET n = 7, mixed = 'seven'")
for name, read in [("to_list", lambda r: r.to_list()), ("to_columns", lambda r: r.to_columns()), ("to_dataframe", lambda r: r.to_dataframe()), ("to_arrow", lambda r: r.to_arrow())]:
try:
read(db.query("sql", "SELECT n, mixed FROM Reading WHERE n >= 6 ORDER BY n"))
print(f"{name:<13}: ok")
except Exception as e:
print(f"{name:<13}: {type(e).__module__}.{type(e).__name__}: {e}")
db.close()
Output
Wheel arcadedb_embedded-26.10.1.dev0-cp312-cp312-manylinux_2_34_x86_64.whl with its bundled JRE (OpenJDK 25.0.4.1), JPype 1.7.1, and Python 3.12.13. The engine build the wheel reports through com.arcadedb.Constants is 26.10.1-SNAPSHOT (build 04bb03ebec07fce19101595d65ca1e3bdf516755), an ancestor of upstream cecb369df9. The repository is at f92ac1381c8335249a6fde644cbf04594f6d6391, and the installed Python sources are identical to bindings/python/src/arcadedb_embedded/ at that commit.
to_arrow(batch_size= 25000):
level int64 [10, 20, None, None, 50]
tags list<item: string> [['a'], [], ['b', 'c'], [], ['d']]
to_columns(batch_size= 25000): level is a ndarray: [np.float64(10.0), np.float64(20.0), np.float64(nan), np.float64(nan), np.float64(50.0)]
to_arrow(batch_size= 2):
level string ['10', '20', None, None, '50']
tags list<item: string> [['a'], [], ['b', 'c'], [], ['d']]
to_columns(batch_size= 2): level is a list: [10, 20, None, None, 50]
to_arrow(batch_size= 1):
level string ['10', '20', None, None, '50']
tags string ["['a']", '[]', "['b', 'c']", '[]', "['d']"]
to_columns(batch_size= 1): level is a list: [10, 20, None, None, 50]
to_list : ok
to_columns : ok
to_dataframe : ok
to_arrow : pyarrow.lib.ArrowInvalid: Could not convert 'seven' with type str: tried to convert to int64
A property test (hypothesis 6.168.3, seed 20261003) over seven column types with NULLs (integer, float, bool, str, datetime, date, and list; the DECIMAL and map columns are left out because of their own defects) finds the same thing without any hand-built data. At the default batch size it passes 200 examples. With batch_size drawn from 1 to 5 it fails and shrinks to two rows, {'l': None} and {'l': []}, at batch_size=1, where the column comes back as the string '[]'.
Cause
ColumnBatcher types a column whose values in the batch are all NULL as a string column (ColumnBatcher.java L121), and a list column is read as pa.array(json.loads(...)) (results.py L616), which types a batch of empty lists as list<null>. _unify_arrow_chunk_types then widens only the pair int64 and float64 (results.py L100-L104) and sends every other combination, including string with int64 and list<null> with list<string>, to _cast_all_to_string (results.py L106). An all-NULL batch carries no type information, so it should not count as a conflicting type.
The mixed-type failure is the same pa.array(json.loads(...)) call (results.py L616): the unification handles batches that disagree with each other, but a single batch of mixed values never reaches it because the batch itself cannot be turned into an array.
Expected
The Arrow type of a column does not depend on batch_size: an all-NULL batch takes the type of the other batches, and a column that really mixes types gets one defined fallback (strings, as _unify_arrow_chunk_types already chooses for conflicting batches) in a single batch as well, or an ArcadeDBError, but not a raw pyarrow exception.
Proposed fix
In _unify_arrow_chunk_types, treat a chunk that is all NULL, or a list<null> chunk, as compatible with the other chunks and cast it to their type, and fall back to the existing string cast only when the remaining chunks still disagree. In decode_batch, catch pa.ArrowException around pa.array(json.loads(...)) and build the string column the same way _cast_all_to_string does. The first half is a sketch I checked with a run-time monkeypatch that leaves the repository untouched: with it level is int64 and tags is list<item: string> at every batch size. The second half (mixed types in one batch) is not covered by that check, and the last line of the same output still shows the ArrowInvalid.
Related issues
ArcadeData#7108 (closed) added _unify_arrow_chunk_types for a column that is int64 in one batch and float64 in another, and says "anything mixed -> string" is an acceptable choice. This report is different: an all-NULL batch is not a mixed type and should not trigger the string fallback, and a column that mixes types within one batch still raises.
In this repository's tracker no issue is about this. The searches for to_arrow, Arrow, and pyarrow found nothing.
Related reports from the same review: #113 (columns dropped), #115 (DECIMAL columns).
Level: bindings. Two defects in
ResultSet.to_arrow. First, when one batch of a column holds only NULLs (or only empty lists),_unify_arrow_chunk_typesgives up on the column and casts every batch to string. A column with the values 10, 20, NULL, NULL, and 50 isint64at the defaultbatch_size, and isstringwith the values '10', '20', None, None, and '50' atbatch_size=2. A list column becomes the strings"['a']",'[]', and so on atbatch_size=1. So the schema of the table depends on a tuning parameter and on where the NULLs fall. Second, a column that mixes types inside one batch (an int in one row and a str in another, which a schemaless document allows) makesto_arrow()raisepyarrow.lib.ArrowInvalid, a raw pyarrow error, after the result set was consumed, whileto_list,to_columns, andto_dataframereturn the values.Repro
The script prints to stdout. The JVM's own log lines go to stderr and are discarded by the command below.
uv venv .venv uv pip install --python .venv/bin/python arcadedb_embedded-26.10.1.dev0-cp312-cp312-manylinux_2_34_x86_64.whl numpy pandas pyarrow .venv/bin/python repro-7-to-arrow-types.py 2>/dev/nullOutput
Wheel
arcadedb_embedded-26.10.1.dev0-cp312-cp312-manylinux_2_34_x86_64.whlwith its bundled JRE (OpenJDK 25.0.4.1), JPype 1.7.1, and Python 3.12.13. The engine build the wheel reports throughcom.arcadedb.Constantsis26.10.1-SNAPSHOT (build 04bb03ebec07fce19101595d65ca1e3bdf516755), an ancestor of upstreamcecb369df9. The repository is atf92ac1381c8335249a6fde644cbf04594f6d6391, and the installed Python sources are identical tobindings/python/src/arcadedb_embedded/at that commit.A property test (hypothesis 6.168.3, seed 20261003) over seven column types with NULLs (integer, float, bool, str, datetime, date, and list; the DECIMAL and map columns are left out because of their own defects) finds the same thing without any hand-built data. At the default batch size it passes 200 examples. With
batch_sizedrawn from 1 to 5 it fails and shrinks to two rows,{'l': None}and{'l': []}, atbatch_size=1, where the column comes back as the string'[]'.Cause
ColumnBatchertypes a column whose values in the batch are all NULL as a string column (ColumnBatcher.java L121), and a list column is read aspa.array(json.loads(...))(results.py L616), which types a batch of empty lists aslist<null>._unify_arrow_chunk_typesthen widens only the pair int64 and float64 (results.py L100-L104) and sends every other combination, includingstringwithint64andlist<null>withlist<string>, to_cast_all_to_string(results.py L106). An all-NULL batch carries no type information, so it should not count as a conflicting type.The mixed-type failure is the same
pa.array(json.loads(...))call (results.py L616): the unification handles batches that disagree with each other, but a single batch of mixed values never reaches it because the batch itself cannot be turned into an array.Expected
The Arrow type of a column does not depend on
batch_size: an all-NULL batch takes the type of the other batches, and a column that really mixes types gets one defined fallback (strings, as_unify_arrow_chunk_typesalready chooses for conflicting batches) in a single batch as well, or anArcadeDBError, but not a raw pyarrow exception.Proposed fix
In
_unify_arrow_chunk_types, treat a chunk that is all NULL, or alist<null>chunk, as compatible with the other chunks and cast it to their type, and fall back to the existing string cast only when the remaining chunks still disagree. Indecode_batch, catchpa.ArrowExceptionaroundpa.array(json.loads(...))and build the string column the same way_cast_all_to_stringdoes. The first half is a sketch I checked with a run-time monkeypatch that leaves the repository untouched: with itlevelisint64andtagsislist<item: string>at every batch size. The second half (mixed types in one batch) is not covered by that check, and the last line of the same output still shows theArrowInvalid.Related issues
ArcadeData#7108 (closed) added
_unify_arrow_chunk_typesfor a column that is int64 in one batch and float64 in another, and says "anything mixed -> string" is an acceptable choice. This report is different: an all-NULL batch is not a mixed type and should not trigger the string fallback, and a column that mixes types within one batch still raises.In this repository's tracker no issue is about this. The searches for
to_arrow,Arrow, andpyarrowfound nothing.Related reports from the same review: #113 (columns dropped), #115 (DECIMAL columns).