End-to-end fit + predict comparisons won by Flow.
canonical v2 / parity + disparity benchmark
Eligibility never means identity.
All 19 canonical rows are measured and currently eligible for comparison, but numerical, semantic and runtime disparities remain first-class evidence. This page renders the committed benchmark and disparity artifacts directly so differences cannot disappear merely because a row passes its contract.
End-to-end comparisons won by scikit-learn.
Rows admitted to the competitive denominator.
Rows whose fitted state, score, configuration or semantics genuinely diverge, above float-noise floors. Runtime differences are tracked per row but not counted here.
runtime overview
The plots are generated from the canonical JSON.
Each runtime plot shows end-to-end fit + predict time on a log scale. The plots use the same rows as the table below and therefore update whenever the frozen canonical result changes.
All 19 speed ratios
scikit-learn total time divided by Flow total time. The vertical 1× line separates Flow wins from scikit-learn wins.
Iris total runtime
Digits total runtime
Diabetes total runtime
persistent disparity
Passing parity does not erase the gap.
The disparity plot normalizes each row's principal numerical difference against its effective tolerance where a tolerance is available. A value near 1 means the row is close to the acceptance boundary. Semantic/configuration differences are tracked in the same artifact and remain visible in the table.
Numerical disparity relative to tolerance
The dashed line is the acceptance boundary. Values can remain non-zero even for eligible rows.
all canonical rows
No selected-win table.
Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow. Strict diagnostic status is kept separate from final eligibility.
| Algorithm | Dataset | Final parity | Strict diagnostic | Winner | score |Δ| | sklearn ms | Flow ms | speedup |
|---|
larger data
The canonical rows all fit in 1797 samples.
A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features. One run of it does not settle a row: at 100 and 1000 samples a fit finishes in well under a millisecond, and the CI runner moves that by more than the difference being measured. Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso. The table is therefore the spread across consecutive runs rather than one run's number, sorted with the narrowest margins first.
| Algorithm | samples | features | runs won | median | range |
|---|
This matrix is reported without gating the build. What does gate is benchmarks/scaled_flow_baseline.json, a Flow-against-itself comparison refreshed from a CI artifact.
the rest of the library
The library is 203 estimators.
The canonical rows above race twelve estimators under a parity contract. lib/scikit exports 203, and a statement about Flow against scikit-learn covers six percent of it while the rest go unmeasured. A separate registry maps every exported fit to its scikit-learn counterpart and times both sides on the same data.
Ranked against a named scikit-learn class.
Of the rows that produce a ratio.
A second function beside each of these does the whole job, and that one is the row that races the class.
Flow implements it and scikit-learn has no equivalent.
The widest ratios say more about the defaults than about the code. Each side runs its own. A row can differ by three orders of magnitude simply because one library does far more work at its defaults. That is a difference in the job, and the ratio does not measure how fast either one is at the same job. Iris also carries no missing values, which leaves the imputers nothing to impute on the Flow side while scikit-learn still runs its full round robin. The narrow rows at the top of the table are the informative ones.
| Flow estimator | scikit-learn | Flow ms | scikit-learn ms | speedup |
|---|
methodology
Correctness, disparity and timing are separate dimensions.
The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner.
Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia. The persistent disparity artifact preserves raw numerical gaps and known semantic/configuration differences even after the estimator-specific eligibility contract succeeds.
historical deployment evidence
Footprint and startup remain separate experiments.
The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.
trajectory
Flow versus Python, across freezes.
Each row's speedup at the previous freeze and at the latest one. A speedup can move because Flow changed or because scikit-learn's side changed on that runner. When a row moves by more than 10%, the last column names which side's own time moved more, from the committed absolute timings.
reproduce