Sitelet https://godofecht.github.io/flow-scikit/benchmarks.html

canonical v2 / parity + disparity benchmark

Eligibility never means identity.

All 19 canonical rows are measured and currently eligible for comparison, but numerical, semantic and runtime disparities remain first-class evidence. This page renders the committed benchmark and disparity artifacts directly so differences cannot disappear merely because a row passes its contract.

Flow wins...

End-to-end fit + predict comparisons won by Flow.

sklearn wins...

End-to-end comparisons won by scikit-learn.

parity eligible...

Rows admitted to the competitive denominator.

substantive disparities...

Rows whose fitted state, score, configuration or semantics genuinely diverge, above float-noise floors. Runtime differences are tracked per row but not counted here.

TIMING_UNIT|msend-to-endseed=4280/20 persisted split2% practical tie thresholddisparity retained after eligibility
KMeans note: Digits KMeans is eligible under the same declared contract as every other clustering row. Its seeded k-means++ initialization now matches scikit-learn's, so the strict diagnostic and the final eligibility decision agree. The convergence statistic, the point at which inertia is reported, empty-cluster relocation and the n_init selection rule still differ and stay visible in the disparity artifact.

runtime overview

The plots are generated from the canonical JSON.

Each runtime plot shows end-to-end fit + predict time on a log scale. The plots use the same rows as the table below and therefore update whenever the frozen canonical result changes.

All 19 speed ratios

scikit-learn total time divided by Flow total time. The vertical 1× line separates Flow wins from scikit-learn wins.

Iris total runtime

scikit-learnFlow

Digits total runtime

scikit-learnFlow

Diabetes total runtime

scikit-learnFlow

persistent disparity

Passing parity does not erase the gap.

The disparity plot normalizes each row's principal numerical difference against its effective tolerance where a tolerance is available. A value near 1 means the row is close to the acceptance boundary. Semantic/configuration differences are tracked in the same artifact and remain visible in the table.

Numerical disparity relative to tolerance

The dashed line is the acceptance boundary. Values can remain non-zero even for eligible rows.

all canonical rows

No selected-win table.

Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow. Strict diagnostic status is kept separate from final eligibility.

AlgorithmDatasetFinal parityStrict diagnosticWinnerscore |Δ|sklearn msFlow msspeedup

larger data

The canonical rows all fit in 1797 samples.

A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features. One run of it does not settle a row: at 100 and 1000 samples a fit finishes in well under a millisecond, and the CI runner moves that by more than the difference being measured. Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso. The table is therefore the spread across consecutive runs rather than one run's number, sorted with the narrowest margins first.

Algorithmsamplesfeaturesruns wonmedianrange

This matrix is reported without gating the build. What does gate is benchmarks/scaled_flow_baseline.json, a Flow-against-itself comparison refreshed from a CI artifact.

the rest of the library

The library is 203 estimators.

The canonical rows above race twelve estimators under a parity contract. lib/scikit exports 203, and a statement about Flow against scikit-learn covers six percent of it while the rest go unmeasured. A separate registry maps every exported fit to its scikit-learn counterpart and times both sides on the same data.

raced...

Ranked against a named scikit-learn class.

Flow faster...

Of the rows that produce a ratio.

older entry point...

A second function beside each of these does the whole job, and that one is the row that races the class.

no counterpart...

Flow implements it and scikit-learn has no equivalent.

Read this for what it is. These rows carry no parity contract, no declared tolerances and no disparity report. Each library runs its own defaults over the same data, which answers whether an implementation is in the same performance league and says nothing about whether it computes the same thing. The canonical rows above are where numerical equivalence is established. These timings come from the CI runner, where a job times both sides and fails the build if any ranked row is slower than scikit-learn.

The widest ratios say more about the defaults than about the code. Each side runs its own. A row can differ by three orders of magnitude simply because one library does far more work at its defaults. That is a difference in the job, and the ratio does not measure how fast either one is at the same job. Iris also carries no missing values, which leaves the imputers nothing to impute on the Flow side while scikit-learn still runs its full round robin. The narrow rows at the top of the table are the informative ones.

Flow estimatorscikit-learnFlow msscikit-learn msspeedup

methodology

Correctness, disparity and timing are separate dimensions.

The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner.

Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia. The persistent disparity artifact preserves raw numerical gaps and known semantic/configuration differences even after the estimator-specific eligibility contract succeeds.

historical deployment evidence

Footprint and startup remain separate experiments.

The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.

trajectory

Flow versus Python, across freezes.

Each row's speedup at the previous freeze and at the latest one. A speedup can move because Flow changed or because scikit-learn's side changed on that runner. When a row moves by more than 10%, the last column names which side's own time moved more, from the committed absolute timings.

reproduce

Read the source artifacts.

Canonical result ↗ Disparity report ↗