Sitelet https://github.com/kubernetes/kops/pull/18675
Skip to content

Cap the kops-controller bootstrap retry backoff, and verify the e2e upgrade actually rolls - #18675

Merged
kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:fix-bootstrap-backoff-and-upgrade-scenario
Aug 9, 2026
Merged

kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:fix-bootstrap-backoff-and-upgrade-scenario

Conversation

@rifelpet

@rifelpet rifelpet commented Aug 8, 2026 •

Copy link
Copy Markdown
Member

Two fixes found while working through the kops-azure TestGrid after #18604. Neither is Azure-specific in its fix, though both were surfaced by Azure runs.

1. Cap the kops-controller bootstrap retry backoff

vfs.RetryWithBackoff never read wait.Backoff.Cap, and pkg/kopscontrollerclient paired that with Factor: 2 and Steps: 100. The interval doubles without limit, so attempt 11 lands 17 minutes in and attempt 12 at 34 minutes — a node whose control plane is slow to become reachable stops retrying at any useful rate.

Seen on e2e-kops-azure-conformance-1-34/2080139096432316416: the VMSS create stalled ~17.5 minutes, which delayed the role assignments the VMs need in order to read their nodeup config, so every worker spent that window failing with AuthorizationPermissionMismatch. kops-controller began serving at 04:21:29, but the first bootstrap request did not arrive until 04:34:25 — 23 seconds after cluster validation had already given up — and succeeded on its first try:

I0723 04:21:29.707790 server.go:154] kops-controller listening on :3988
I0723 04:34:26.183065 server.go:246] performed successful callback challenge with 10.0.0.7:3987
I0723 04:34:26.649916 server.go:287] bootstrap 51.105.40.77:8980 success

All 4 workers were provisioningState: Succeeded in Azure but never joined; validation failed on machine "..." has not yet joined cluster. Nothing else was wrong — the eventual success proves NSG, NAT gateway, LB rule 3988 and attestation were all fine.

RetryWithBackoff now honors Cap, and the bootstrap client sets a 30s cap. Every other caller uses between 4 and 20 steps, where the unbounded growth is already bounded in practice, and none of them set Cap — so this is the only call site whose behavior changes.

2. Assert the A→B upgrade actually rolled the nodes

A cluster that was never rolled validates cleanly, so an upgrade that silently did nothing still reports green.

That is exactly what was happening on e2e-kops-azure-upgrade-dns-none, which failed 6 of 17 runs. In all six, both rolling-update phases printed No rolling-update required and every node stayed on the old kubelet version — the upgrade never happened. The job only went red because an addon rollout happened to still be in flight when validation ran; had it settled twenty seconds earlier, the run would have passed on an un-upgraded cluster, and some of the eleven "passing" runs may have been exactly that.

The underlying cause was fixed in #18669, so this PR does not re-fix it. It adds the guard that would have caught it in July instead of letting it hide behind an intermittent validation failure for three weeks: every node's kubelet must report K8S_VERSION_B.

An earlier revision of this PR also added --wait 15m to the validate cluster call. That was wrong and has been dropped — as @hakman pointed out in review, the cluster should already be stable at that point, since a rolling update's last action is a full cluster validation (pkg/instancegroups/instancegroups.go:262). e2e-kops-aws-upgrade-dns-none is 52/52 on the bare validate cluster, and waiting would have masked the very regression class this PR is meant to catch.

Testing

New unit tests in util/pkg/vfs/context_test.go cover the cap clamping. go build ./... and the affected package tests pass. The scenario change is only exercised by the periodic upgrade jobs.

Draft because I have not run make ci / verify-golangci-lint locally yet.

Fixed an issue where a node could stop retrying kops-controller bootstrap at a useful rate when the control plane was slow to become reachable.

/kind bug
/cc @hakman

RetryWithBackoff never read wait.Backoff.Cap, and the bootstrap client
paired that with Factor 2 and Steps 100, so the retry interval doubled
without limit: attempt 11 lands 17 minutes in, attempt 12 at 34 minutes.

A node whose control plane is slow to become reachable therefore stops
retrying at any useful rate. In one Azure run the scale set create stalled
long enough to delay the role assignments the VMs need to read their nodeup
config. By the time kops-controller was serving, the nodes were asleep: the
first bootstrap request arrived 23 seconds after cluster validation had
already given up, and succeeded on its first try.

Honor Cap in RetryWithBackoff and set a 30s cap on the bootstrap client.
Every other caller uses between 4 and 20 steps, where the unbounded growth
is already bounded in practice, and none of them set Cap.
@kubernetes-prow

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@kubernetes-prow
kubernetes-prow Bot requested a review from hakman August 8, 2026 15:19
@kubernetes-prow kubernetes-prow Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. kind/bug Categorizes issue or PR as related to a bug. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. size/L Denotes a PR that changes 100-499 lines, ignoring generated files. labels Aug 8, 2026
@rifelpet
rifelpet force-pushed the fix-bootstrap-backoff-and-upgrade-scenario branch from a2569a9 to 477eee7 Compare August 8, 2026 15:32
@rifelpet

rifelpet commented Aug 8, 2026

Copy link
Copy Markdown
Member Author

/test all
/test pull-kops-azure-upgrade-dns-none
/test pull-kops-aws-upgrade-dns-none
/test pull-kops-gce-upgrade-dns-none
/test pull-kops-azure-dns-none-ha

Comment thread tests/e2e/scenarios/lib/upgrade.sh Outdated
A cluster that was never rolled validates cleanly, so an upgrade that
silently did nothing still reports green. That is how a broken Azure
rolling update went unnoticed for three weeks: reconcile printed "No
rolling-update required" for both phases, every node stayed on the old
kubelet, and the job went red only when an unrelated addon rollout
happened to still be in flight when validation ran.

Assert that every node's kubelet reports the target version, so that case
fails for the right reason. K8S_VERSION_B is rewritten into a release-dev
URL when it is "ci", so compare on its last path segment, which is what
kubelet reports.
@rifelpet
rifelpet force-pushed the fix-bootstrap-backoff-and-upgrade-scenario branch from 477eee7 to af93dc3 Compare August 8, 2026 22:23
@rifelpet

rifelpet commented Aug 8, 2026

Copy link
Copy Markdown
Member Author

/test all
/test pull-kops-azure-upgrade-dns-none
/test pull-kops-aws-upgrade-dns-none
/test pull-kops-gce-upgrade-dns-none
/test pull-kops-azure-dns-none-ha

@rifelpet
rifelpet marked this pull request as ready for review August 8, 2026 22:32
@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 8, 2026
@kubernetes-prow kubernetes-prow Bot added the lgtm "Looks good to me", indicates that a PR is ready to be merged. label Aug 9, 2026
@kubernetes-prow

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: hakman

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow kubernetes-prow Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 9, 2026
@kubernetes-prow
kubernetes-prow Bot merged commit 3b480f4 into kubernetes:master Aug 9, 2026
31 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. kind/bug Categorizes issue or PR as related to a bug. lgtm "Looks good to me", indicates that a PR is ready to be merged. size/L Denotes a PR that changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants