Sitelet https://github.com/godofecht/flow-scikit/compare/main...real-algorithms
Skip to content
Permalink

Comparing changes

Choose two branches to see what’s changed or to start a new pull request. If you need to, you can also or learn more about diff comparisons.

Open a pull request

Create a new pull request by comparing changes across two branches. If you need to, you can also . Learn more about diff comparisons here.
base repository: godofecht/flow-scikit
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: main
Choose a base ref
...
head repository: godofecht/flow-scikit
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: real-algorithms
Choose a head ref
Checking mergeability… Don’t worry, you can still create the pull request.
  • 2 commits
  • 20 files changed
  • 1 contributor

Commits on Sep 29, 2026

  1. Run the algorithm in the four rows that were standing in for one

    Four implementations said in their own comments that they were simplified, so
    the registry timed them and showed the numbers without ranking them, and the
    page carried a paragraph explaining why three of the four would otherwise have
    been its widest wins. They were wide because they were not doing the work.
    
    spectral_coclustering bucketed each row sum between the smallest and the
    largest. It runs Dhillon's algorithm now: normalize by the square roots of the
    row and column sums, take the singular vectors after the trivial one, and
    cluster rows and columns together in the space they span. 0.03 ms against
    scikit-learn's 8.09.
    
    spectral_biclustering bucketed row means and column means the same way. It
    runs Kluger's: Sinkhorn-Knopp, then rank the singular vectors by how well a one
    dimensional k-means fits each one, keep the best few, and cluster rows and
    columns separately. 0.09 ms against 35.17.
    
    elliptic_envelope drew random subsets and scored each by the product of the
    diagonal of its covariance, which is the determinant only when the features do
    not correlate, and stopped at the draw. The concentration step is the whole
    content of FastMCD, so it takes the h points closest in Mahalanobis distance,
    recomputes, and repeats until the determinant stops falling. The determinant
    comes from a Cholesky factor, and predict uses the full distance rather than
    the diagonal. 2.03 ms against 19.46.
    
    nca computed neither the objective nor its gradient. Its distance summed the
    squares of the individual products instead of squaring the projected
    difference, which is not a distance in the projected space, and its gradient
    kept only the term that pulls same-class points together, with nothing pushing
    back. Both are written out now. 181 ms against 856.
    
    Two test files cover them, by behaviour rather than by shape: the block
    structure has to come back out of both biclusterings, every planted outlier has
    to be found with no false alarms, and the projection has to separate the
    classes better than the identity did. Each was checked by inverting an
    assertion and confirming a non-zero exit.
    godofecht committed Sep 29, 2026
    Configuration menu
    Copy the full SHA
    3d3dbc3 View commit details
    Browse the repository at this point in the history
  2. Race the five rows that were doing half the job

    Five Flow functions covered part of what the scikit-learn class they are named
    after does, so the registry kept them out of the ranking and said why. The
    missing half is written now, as a second entry point beside each of them, and
    the five are raced.
    
    voting_classifier_full_fit and voting_regressor_full_fit fit the ensemble from
    the data, where the functions beside them take estimators that are already
    fitted and their cost is the vote alone. Three trees on both sides, because a
    vote of one is not what either library is being asked for.
    
    stacking_classifier_full_fit cross-validates its base estimators to build the
    meta-features, fits the meta-learner on those, and refits the base estimators
    on all of the data. Out of fold predictions are the point: a base estimator's
    prediction for a row it helped fit would tell the meta-learner more than it
    will know at prediction time.
    
    self_training_classifier_full_fit refits its base estimator on every round,
    which is the expensive half the function beside it leaves to its caller.
    
    incremental_pca_fit walks the design in batches, sized as scikit-learn sizes
    its own, where the partial fit beside it is one step of that.
    
    Racing incremental PCA turned up that it did not work. Its partial fit
    accumulated the same per-feature quantity for every component, so every
    component came out as the same vector, and the step it called a power
    iteration multiplied a component by itself rather than by a covariance. It
    keeps a running scatter matrix now, which combines across batches exactly once
    the shift between two means is accounted for, and takes its components from
    the leading eigenvectors of that. On the test fixture the leading component
    carried a spread of 640 against the second's 4, where the two were previously
    identical to the digit.
    
    Local timings against scikit-learn on the same data:
    
        incremental_pca                 0.010 ms  against  1.01
        voting_classifier_full          0.053     against  2.23
        self_training_classifier_full   0.266     against 19.83
        stacking_classifier_full        0.529     against 16.07
        voting_regressor_full           1.184     against  3.74
    
    A test file covers all five by behaviour: the two ensembles and the stack have
    to classify or regress the fixture, self training has to label the unlabelled
    rows and get them right, and the batched decomposition has to see every row and
    put more spread in the leading component than the second. It was checked as a
    real gate by inverting an assertion.
    godofecht committed Sep 29, 2026
    Configuration menu
    Copy the full SHA
    b49b5dd View commit details
    Browse the repository at this point in the history
Loading