Conversation
PlDfStatsV2 should compute value_counts for every column in one pl.collect_all instead of one Series at a time, hand each column its result through StatPipeline.process_df, and derive mode from it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#997) PlDfStatsV2 computes value_counts for every column in one pl.collect_all before the per-column pass and seeds each column's accumulator with it through the new StatPipeline.process_df(column_initial_stats=...). process_column skips a stat whose every provided key was seeded, so the batch is the only provider of value_counts on that path. pl_base_summary_stats is split: pl_value_counts provides value_counts per Series (for StatPipeline callers that pass a raw series), and pl_base_summary_stats takes value_counts from the accumulator and derives mode from its first row instead of a second group-by. Object columns are left out of the batch because collect_all panics on them; a batch that raises falls back to the per-Series path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
📦 TestPyPI package publishedpip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37085735205or with uv: uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37085735205MCP server for Claude Codeclaude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37085735205" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table📖 Docs preview🎨 Storybook preview |
|
Review of this PR found two issues, both addressed.
Local run before pushing: ruff clean, unit suite 1174 passed. Fifteen lazy-widget tests failed locally with |
MeasurementsTimed
Peak RSS ( The search-rerun case on the small file measured 0.41 s on main here, not the ~588 ms expected; I did not find what accounts for the difference. CorrectnessCompared
Every other key, including Summary: the large-file stats phase is about twice as fast, at the cost of about 3 GB more peak memory on a 43-column, 10.8M-row frame. The only output differences are |
|
Closing. Batching value_counts across columns with collect_all tunes the in-memory path and costs about 3 GB more peak memory at 10.8M rows; stats for the out-of-core backend come from a single lazy select. The branch stays. |
Part of #997. One of several candidate approaches; the alternatives are taking
modefrom the first row ofvalue_countsonly (perf/997-mode-from-value-counts) and the batch pipeline / stats pre-pass work under #999.Problem
pl_base_summary_statsrunsvalue_countsandmodeon onepl.Seriesat a time, so polars never parallelises them across columns. On the 10.8M-row, 43-column parquet from #997 that is 6.8 s ofvalue_countsplus 2.5 s ofmode, against 2.8 s when the same 43value_countsare collected together withpl.collect_all.Approach
PlDfStatsV2computesvalue_countsfor every column in onepl.collect_allbeforeprocess_dfruns, converts each result to thepd.Seriesshape the rest of the DAG expects, and hands them toStatPipeline.process_dfas per-column initial stats.pl_base_summary_statstakesvalue_countsfrom the accumulator and derivesmodefrom its first row instead of running a second group-by. A per-series provider,pl_value_counts, stays inPL_ANALYSIS_V2for callers that useStatPipelinedirectly; when a column'svalue_countsis supplied up front the pipeline skips that stat, so the key is only ever provided once.What changes
StatPipeline.process_dfacceptscolumn_initial_stats(orig column name to stats dict), merged into each column'sinitial_stats.StatPipeline.process_columnskips a stat whose every provided key was supplied throughinitial_stats.polars_utils.batch_value_counts(df, columns)runs the one-pass collect and returns{column: pd.Series};vc_frame_to_pdis the shared conversion used by both paths. Object columns are left out (polars panics on them insidecollect_all).pl_stats_v2:pl_value_countsprovidesvalue_counts;pl_base_summary_statsrequires it and derivesmodefrom it.PlDfStatsV2runs the batch when the pipeline providesvalue_counts, falls back to the per-series path if the collect raises, and reuses the batch inadd_analysis.The 50k-row sample in
PlDfStatsV2.get_operating_dfis unchanged; #992 covers it.Tests
tests/unit/test_pl_stats_v2.py::TestPlBatchValueCounts:pl.Series.value_countsis monkeypatched with a counter andPlDfStatsV2on a 5-column frame makes zero per-Series calls; the batched pipeline output matches the per-series output key by key on a frame with no count ties; a suppliedvalue_countspre-empts the per-series stat (mode and distinct_count follow it);process_columnwith a raw series and no initial value still producesvalue_countsandmode; Object columns are left out of the batch and still get stats.Compatibility
pl_base_summary_statsnow requiresvalue_counts, and the provider of that key moved to the newpl_value_countsstat. A pipeline built by hand from the old public name, for example a list copied from an olderPL_ANALYSIS_V2such as[pl_typing_stats, _type, pl_base_summary_stats, ...], raisesDAGConfigError: No function provides 'value_counts'at construction. Addingpl_value_countsto the list fixes it. In-repo callers usePL_ANALYSIS_V2and are unaffected, and a custom stat that provides bothvalue_countsandmodestill works. The existing tests that builtStatPipeline([pl_base_summary_stats])directly were updated to includepl_value_counts.Trade-offs
modeamong tied values was unspecified before and still is; it now followsvalue_countsorder, somodeandmost_freqalways agree.modefor datetime and duration columns comes back aspd.Timestamp/pd.Timedelta(subclasses of the stdlib types, and whatmost_freqalready returned) rather thandatetime/timedelta.collect_allraises for any column the whole batch is dropped and every column goes through the per-series path, so a frame with one unsupported column pays the old cost plus a failed attempt.value_countspassed toPlDfStatsV2is pre-empted by the batch;StatPipelineused directly is unaffected.🤖 Generated with Claude Code