Sitelet https://godofecht.github.io/flow-scikit/

compiled classical machine learning / Flow

Rebuild the estimator stack. Measure what actually changes.

flow-scikit reimplements classical ML in Flow and tests it against scikit-learn with parity-gated timings, explicit execution-substrate analysis, native deployment and reproducible benchmark artifacts.

  • 19/19 parity eligible
  • native binary
  • substrate-aware appraisal
canonical parity19 / 19

Every canonical row has resolved parity and measurement status.

Flow timing wins19 / 19

Current canonical end-to-end result; all rows remain visible.

estimator operations mapped491

Generated sklearn execution-substrate inventory.

runtime profiles32

Mixed-stack attribution rows feeding the optimization roadmap.

the important distinction

"Python versus compiled" is too crude a model for scikit-learn.

Its public interface is Python, but hot paths may execute in NumPy/SciPy, BLAS/LAPACK, sklearn-owned compiled code, or external native libraries such as liblinear and libsvm. The meaningful question is therefore not simply whether Flow beats Python, but which execution layer Flow is replacing, retaining, or compiling around.

Inspect the execution map →
Python-boundInterpreter and orchestration work can be direct compilation targets.
Mixed / boundary-heavyWhole-estimator compilation can remove crossings and temporary allocations.
BLAS / LAPACK-boundRetain mature kernels unless evidence supports replacement.
sklearn-owned nativeBenchmark Flow against the compiled implementation directly.
External nativeliblinear and libsvm are compiled competitors. Beating them is a different claim from beating an interpreter.

current appraisal

The architecture result is more useful than a blanket speed claim.

The committed map groups every canonical row by sklearn fit substrate. Flow wins all of them, and the margin follows the class. That is evidence that execution substrate is predictive enough to guide optimization work, while still being far from a causal proof by itself.

What the evidence supports today: Flow wins every canonical row, including the ones where scikit-learn calls liblinear or libsvm, and ships a much smaller native deployment. Those native solvers remain close baselines rather than beaten ones: the narrowest canonical margins sit near 1.5x, and a row that close can land either way on a different BLAS and core count. The roadmap ranks work from measured substrate, runtime attribution, benchmark results and implementation readiness rather than from source-language folklore.

evidence chain

Correctness → timing → substrate → attribution → roadmap.

The repository now keeps each stage machine-readable and reproducible.

01 / parity

Resolve all 19 canonical rows.

Competitive timing only follows estimator-specific numerical checks.

Parity evidence →
02 / coverage

Measure the whole library.

The canonical rows race twelve estimators. 203 are exported, and every one of them is now either raced, or carries a written reason why not.

Estimator coverage →
03 / architecture

Map what sklearn actually executes.

491 estimator-operation rows are classified with evidence and drift detection.

Architecture map →
04 / priorities

Turn evidence into an optimization queue.

Runtime profiles, native-hotspot dispositions and whole-estimator experiments feed a generated roadmap.

Roadmap ↗

minimal example

Fit, predict, inspect.

The library remains a native Flow implementation rather than a Python compatibility layer.

import "lib/scikit/scikit.flow"

let model = knn_classifier_fit(X_train, y_train, 5)
let predictions = knn_classifier_predict(model, X_test)
let score = accuracy_score(y_test, predictions, n_test)

open evidence

Read the benchmark. Inspect the substrate. Reproduce the result.

Open flow-scikit on GitHub ↗