This directory contains the canonical numerical-parity, timing and execution-architecture evidence used by the repository and Pages site.
headline_result_v2.json is the current competitive source of truth.
The committed result contains 19 total rows, 19 parity-eligible comparisons, 19 Flow wins, 0 scikit-learn wins, 0 ties, 0 parity-unresolved rows and 0 measurement-unresolved rows.
A competitive speed claim is only emitted when a row has resolved timing, a declared millisecond unit, comparable benchmark semantics and a passing estimator-specific parity contract.
The benchmark contract is milliseconds end-to-end.
The Python v2 runner uses high-resolution perf_counter_ns()-backed adaptive measurement and emits explicit TIMING_UNIT|ms. Flow emits millisecond timings. run_headline.py repeats complete runs and aggregates process-level measurements with medians and IQRs.
Comparison tools do not apply hidden seconds-to-milliseconds conversions. Missing or unexpected timing units are rejected.
Older published data that stored second-valued Python literals in fields labelled ms is retained only as historical audit material. It is not part of canonical v2.
Canonical v2 consumes persisted train/test fixtures shared by Python and Flow. CI verifies that the JSON split indices and Flow binary split fixtures are byte-for-byte equivalent.
parity_contract.json defines the semantic contract for all 19 rows. Supervised rows compare their declared predictive metric under explicit tolerances. KMeans uses adjusted Rand index plus inertia rather than label-mapped accuracy. PCA additionally checks explained-variance ratios, singular values, reconstruction error and sign-aligned components.
Digits KMeans is classified as approximately equivalent under the tolerances declared in parity_contract.json, with no estimator-specific exception. audit_kmeans_semantics.py re-derives Flow's initial centre indices from a Python mirror of Flow's own MT19937 and greedy k-means++ and checks them against scikit-learn's initializer for all ten n_init restarts. It also records the differences that survive that alignment: the convergence statistic, the point at which inertia is reported, empty-cluster relocation and the n_init selection rule. Each was substituted in turn and none moves a canonical row.
Each contract row records its configuration twice, once in Flow's vocabulary and once in scikit-learn's. Comparing the two dictionaries key by key reported a difference every time the two projects spelled the same setting differently, and the differences that were real sat buried among them.
config_equivalence.py holds the equivalences the
comparator may apply, and generate_disparity_report.py records every applied
one in that row's configuration_equivalences so nothing is dropped silently.
Three rules:
| Rule | Example | Condition |
|---|---|---|
| Cross-vocabulary mapping | sklearn C=1.0 against Flow l2=0.008333333 |
the recorded numbers satisfy l2 = 1 / (C * n_train) to 1e-5 relative, with n_train read from split_indices.json |
| Absent equals explicitly disabled | Flow penalty=none on LinearRegression |
one side records the feature off and the other has no such parameter |
| Solver-private parameter | Flow learning_rate against sklearn lbfgs; sklearn dual against Flow's LinearSVC |
the knob belongs to one side's solver and the counterpart's solver does not have it |
All three apply only to a parameter present on exactly one side. When both sides
record a parameter, the values are compared and any difference is reported, so
max_iter 200 vs 1000 and optimizer lbfgs_no_line_search vs lbfgs stay
visible. A cross-vocabulary mapping that does not hold numerically is reported
with the value the relation demanded attached; that is the mismatch #430 found.
A score tolerance says two implementations agree on the answer. It says nothing
about whether they agree on the model. Both runners therefore emit DETAIL
records alongside each RESULT line:
DETAIL|<algorithm>|<dataset>|<field>|<value or comma-separated values>
generate_disparity_report.py pairs every field
both runners emitted for a canonical row and writes the deltas into that row's
model_state_diagnostics. Scalars give <field>_abs_diff and
<field>_relative_diff; equal-length vectors give <field>_max_abs_diff,
<field>_max_relative_diff and <field>_first_divergent_index, which is -1
when the vectors match elementwise and otherwise points at the first position
they disagree on; a length disagreement is recorded as
<field>_length_abs_diff. This runs off the frozen DETAIL records rather than
the parity comparison, so a row that misses its score tolerance still keeps its
state evidence. model_state_coverage.py audits which
canonical rows have any state evidence at all.
What each estimator emits:
| Row | Learned state |
|---|---|
| LogisticRegression, LinearSVC | sorted class labels, coefficient Frobenius norm, per-class row norms, intercepts |
| KernelSVC_RBF | gamma, C, one-vs-one pair labels and sizes, support-vector count per pair, bounded (at-C) support count per pair, summed absolute dual coefficients per pair, intercept per pair |
| DecisionTree | node and leaf counts, realized depth, root split feature/threshold/impurity, both depth-1 splits, node count per depth, split-feature histogram, sample-weighted mean leaf depth, full preorder split feature and threshold vectors |
| RandomForest | per-tree node/leaf counts, depths, root splits and impurities, per-tree bootstrap index checksum and unique-sample fraction, per-tree feature-subsample seed, mean vote margin, mean top-vote fraction, unanimous-vote fraction |
| GaussianNB | sorted class labels, class priors, per-class mean and variance row norms, Frobenius norms, variance min/max |
| KMeans | inertia, iteration count, sorted center norms, sorted training cluster sizes |
| PCA | explained-variance ratios, singular values, reconstruction MSE, first two components |
| Ridge, Lasso, LinearRegression | full coefficient vector, intercept, coefficient L2 norm and absolute sum, zero-coefficient count |
| KernelRidge_RBF | alpha, gamma, training size, dual-coefficient norm, absolute sum, min, max and mean |
Large structures are summarised rather than dumped, with one exception. A single
decision tree emits its full preorder split vectors, because node ids are
assigned parent-before-children and left-subtree-first on both sides, so index
i names the same position in both trees and first_divergent_index is then a
literal pointer at the first structurally different node. Forests emit per-tree
scalars and a bootstrap checksum instead of every bootstrap index; kernel SVMs
emit per-pair aggregates instead of every dual coefficient. Per-class and
per-pair vectors are ordered by class label so they line up with sklearn
regardless of the order Flow discovered the classes in.
Emission happens outside the timing windows and reads only fitted state, so it does not move any canonical metric or timing.
The committed model_state_diagnostics for Ridge, Lasso and LinearRegression
on diabetes record coefficient differences of 7e-5, 6.9% and one index past
the 1e-6 threshold. Issue #472 traced all three to the scikit-learn side of
the comparison, measured against f64 references recomputed from the committed
fixtures:
- Lasso: Flow's objective sits 3e-12 from the exact optimum; sklearn stopped at 246 iterations with its duality gap 1750x inside its own threshold, 7.7e-5 above the optimum. The 6.9% is the smallest coefficient, where Flow matches the optimum to all printed digits.
- Ridge: the conventions match to 2e-14 in f64. sklearn's f32 Cholesky lands 964 ulps from the exact solution at the divergent index; Flow lands 8.5.
- LinearRegression: the centred design has condition number 20.7 through the
collinear S1/S2 pair, and f32-ulp jitter in the inputs reproduces the
observed per-index divergence. Flow is closer to the exact solution at 9 of
10 indices. Flow's QR runs in f64; sklearn's
lstsqran on the f32 input.
All three rows pass their score gates. The full measurements are in #472.
Both runners read the same fixture bytes. A DETAIL field that disagrees
between a local run and the committed evidence therefore means one side computed
a different number from the same input.
Issue #461 recorded one such disagreement.
DETAIL|RandomForest|digits|unanimous_vote_fraction is 0.197222222 in the
committed sklearn_results_v2.txt and was observed as 0.205556 locally, three
of 360 test rows, with the per-tree node counts matching exactly. Summation order
in the scaler was the suspected cause. It turns out to be the dtype the sklearn
side scales in.
standard_scaler_fit accumulates the mean and the variance in f64 and uses the
two-pass form, mean first and then squared deviations. On the digits fixture its
f32 mean and std are bit-identical in all 64 columns to the values obtained
from exact rational arithmetic over the integer pixel data. The tightest column's
exact std sits 4.9e-10 (relative) from an f32 rounding boundary, roughly four
orders of magnitude above the error an f64 sum over 1437 rows can carry, so no
reassociation of that sum can move the rounded f32 result. The margin is not
generous. An f32 two-pass over the same data moves the rounded std in 61 of 64
columns, and an f32 E[x^2] - E[x]^2 moves it in 23 of 64.
bench_sklearn_v2.py casts the digits matrix to float32 before fitting, so
StandardScaler.transform evaluates (x - mean_) / scale_ in float32 with a
single rounding, and the forest then votes unanimously on 71 of 360 rows:
0.197222222, the committed value. Scaling the same split in float64 and
rounding to float32 afterwards changes 7412 of the 23040 test cells in their
last bits, three rows flip, and the fraction becomes 0.205555556. Both are
reproducible on demand and neither depends on the platform. mean_ and scale_
themselves are float64 in both cases and agree bit for bit.
So the committed snapshot is correct as recorded. A local run that disagrees with it should be checked first for the dtype of the array it scaled.
FLOW_HEADLINE_COMMAND="flow run benchmarks/bench_flow_v2.flow" \
python benchmarks/run_headline.py --repeats 7The runner writes the raw Python/Flow result files, environment metadata, row-level comparison output and generated headline summary.
The benchmark result is joined to a separate sklearn execution-architecture pipeline so a Flow win or loss can be interpreted against what sklearn actually executes.
The committed evidence currently contains:
- 491 estimator-operation inventory rows
- 32 runtime-attribution rows
- 491 optimization-roadmap rows
- 146 native/mixed hotspot dispositions
- 8 whole-estimator experiments
- substrate and speedup evidence for all 19 canonical rows
run_architecture_pipeline.py regenerates the inventory, profiles, optimization ranking, native-hotspot audit and architecture/performance map against the pinned sklearn dependency set in requirements-architecture.txt.
CI detects estimator-surface drift: newly added or removed sklearn estimator operations must be reflected in the committed generated inventory rather than silently changing the analysis.
The main generated views are:
SKLEARN_EXECUTION_INVENTORY.mdARCHITECTURE_PERFORMANCE_MAP.mdOPTIMIZATION_ROADMAP.mdNATIVE_HOTSPOT_AUDIT.md
A speedup is sklearn_ms / flow_ms. Values above 1x mean Flow is faster; values below 1x mean scikit-learn is faster.
The current architecture map shows a useful but non-causal pattern: Flow wins every headline row in all three substrate classes, and the margin varies with the class, from a mean of 31.14x on Python-bound rows down to 2.94x on external-native-bound ones. That ordering is evidence for prioritization. It is not proof that execution substrate alone determines performance.
Mature BLAS/LAPACK, liblinear, libsvm and other native backends are treated as native competitors. The optimization roadmap deliberately prefers retaining those kernels unless benchmark and parity evidence justify replacement.
The canonical rows run on iris, digits and diabetes and stop at 1797 samples. bench_scaled.flow and bench_scaled_sklearn.py run five estimators at 100, 1000 and 10000 rows against 8 and 32 features. CI runs both sides every push and reports the comparison without gating the build.
One run does not settle a row. At 100 and 1000 samples a fit finishes in well under a millisecond, and the runner moves that by more than the Flow-versus-sklearn difference: Lasso at 1000 rows and 32 features was recorded at 3.83x on one run and 0.92x on another, on code that differs in nothing touching Lasso.
So the published claim is the spread. summarize_scaled_ci.py folds each run's scaled_comparison.json into scaled_ci_history.json, which records every observation per row along with its minimum, median and maximum:
gh run download <id> -n scaled-benchmark-<id> -D /tmp/<id>
python benchmarks/summarize_scaled_ci.py <id>=/tmp/<id>/scaled_comparison.json
Runs merge by id, so adding a new one extends the history and re-adding an existing one replaces it.
scaled_flow_baseline.json is a different artifact for a different job. It compares Flow against its own earlier CI timings and does fail the build, with a 20% relative tolerance and a 0.25 ms absolute floor.
Each of its rows is the slowest observation across the runs it was built from, so the gate fires when the code is slower than it has ever legitimately been and stays quiet when a run is merely unlucky. Taken from one run it does the opposite: RandomForest at 1000 rows and 8 features has been measured at 1.50, 1.58, 1.63, 2.31, 2.88 and 3.32 ms on identical code, and a baseline taken from the 1.58 run failed the build on the 2.88 one. Rebuild it with the same script:
python benchmarks/summarize_scaled_ci.py --baseline benchmarks/scaled_flow_baseline.json \
<id>=/tmp/<id>/scaled_comparison.json ...
Each run directory needs scaled_flow.json beside scaled_comparison.json. Build it from CI artifacts; a developer machine's numbers would make a fast laptop the standard CI has to meet.
The canonical benchmark races 12 estimators. lib/scikit exports 203, and a
claim about Flow against scikit-learn says little while the other 191 are
unmeasured. Four pieces cover the rest, all driven from one registry so the two
sides cannot drift apart.
estimator_coverage.py parses every exported *_fit
out of lib/scikit, resolves its arguments from a table keyed by parameter
name, finds the scikit-learn class by name, and sorts each estimator into a
bucket. Nothing is dropped silently; every entry that is not raced carries its
reason.
| bucket | meaning |
|---|---|
runnable |
arguments resolved and a scikit-learn counterpart exists |
shaped |
its fit does not begin with a feature matrix, so the registry carries the call written out, and the row is raced and ranked like any other |
different_shape |
the Flow function does part of what the scikit-learn class does, such as a voting estimator that takes models already fitted |
flow_only |
Flow implements it and scikit-learn has no equivalent |
simplified |
the implementation's own comments call it a simplified stand-in |
blocked |
the signature is not resolved yet, with the missing parameter named |
A shaped entry gives the Flow call in terms of the variables the generated
harness declares, what the scikit-learn side fits and transforms, and where
needed a preamble that builds the input: a pipeline of a scaler and a logistic
regression, a corpus for the two text vectorizers, a list of dicts for the dict
vectorizer. The corpus lives in the registry so the Flow file and the
scikit-learn harness read one copy of it.
The simplified bucket is detected from the source rather than listed, so it
stays true as the implementations are filled in. It matters: spectral_biclustering
thresholds row and column means where scikit-learn does an SVD and k-means, and
came out at 25106x. Three of the four it catches would otherwise have been the
widest wins in the whole matrix.
generate_estimator_bench.py emits the Flow
timing blocks from that registry, split across several files because Flow issue
#469 miscompiles some functions once a program grows past a certain size. Each
block prints one line and flushes it. Without the flush, one estimator trapping
takes the whole file's buffered output with it and the run looks empty rather
than partial, which is how a RANSAC crash first presented.
bench_estimators_sklearn.py times the
scikit-learn side of the same registry, and
compare_estimators.py joins them.
python benchmarks/estimator_coverage.py
python benchmarks/generate_estimator_bench.py
python benchmarks/run_estimator_bench.py --rounds 3
python benchmarks/bench_estimators_sklearn.py
python benchmarks/compare_estimators.py benchmarks/estimator_flow_raw.txt
python benchmarks/check_estimator_matrix.py
run_estimator_bench.py checks every chunk's exit
status, so a chunk that dies takes its own row count down with it in the
summary rather than disappearing, and reuses the binary the first round leaves
behind so later rounds pay for timing instead of for compiling the library
again. check_estimator_matrix.py fails on a
ranked row slower than scikit-learn, on a row that reported no timing, and on a
ranked count that has quietly shrunk. The Wide estimator matrix job in
.github/workflows/flow.yml runs all of it on a runner that is not competing
with anything, which is where a published number belongs. A developer machine
under load recorded the same row at 1.72x and at 0.96x in consecutive runs.
Read the result for what it is. These rows have no parity contract, no declared
tolerances and no disparity report: each library runs its own defaults over the
same data. That answers whether a Flow implementation is in the same
performance league and says nothing about whether it computes the same thing.
The per-estimator numerical contract remains the canonical benchmark's job, and
estimator_comparison.json says so in its own contract field.
The breadth is worth having for correctness as much as for speed. The first run
of it found a heap-buffer-overflow in _solve_lstsq_qr, which assumed a design
has at least as many rows as columns; RANSAC reaches it by fitting 5 sampled
rows against the 10 features of diabetes.
publish_headline_v2.py validates the committed canonical benchmark and architecture map, then copies the JSON artifacts into docs/ for the static site. The public benchmark and architecture pages render those artifacts directly instead of embedding hand-maintained timing claims.