Sitelet https://github.com/kubernetes/kops/pull/18510
Skip to content

Make cloud-controller-manager pods tolerate all taints - #18510

Merged
kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:ccm-tolerate-all-taints
Jun 29, 2026
Merged

kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:ccm-tolerate-all-taints

Conversation

@rifelpet

Copy link
Copy Markdown
Member

What

Set tolerations: [{operator: Exists}] on every cloud-controller-manager pod that runs with hostNetwork: true (AWS, Azure, DigitalOcean, GCP, Scaleway), so the CCM can always be scheduled onto a node regardless of which taints are currently present. The Scaleway CCM additionally gains a node-role.kubernetes.io/control-plane nodeAffinity, since broadening its tolerations means tolerations can no longer keep it off worker nodes.

The Azure cloud-node-manager DaemonSet keeps nodeSelector: kubernetes.io/os: linux (it is a per-node agent and must run on every node); only its tolerations are broadened.

This matches the pattern already used by node-level DaemonSets that must run everywhere — the Cilium agent, CNI plugins, and CSI node drivers all use - operator: Exists.

Why

e2e-kops-do-gossip is intermittently failing cluster validation, e.g.:
https://prow.k8s.io/view/gs/kubernetes-ci-logs/logs/e2e-kops-do-gossip/2068280267818143744

The deadlock

With --cloud-provider=external, a node has no InternalIP until the cloud-controller-manager initializes it. The Cilium agent needs that node IP — without it the agent dies with unable to determine direct routing device (k8sNodeIP=""), so it never becomes ready.

Meanwhile the cilium-operator (set-cilium-node-taints: "true") adds node.cilium.io/agent-not-ready:NoSchedule to any node whose Cilium agent isn't ready yet. The DO CCM DaemonSet did not tolerate that taint. So once the operator taints the control-plane node before the CCM pod is bound there:

  1. CCM can no longer schedule onto the control-plane node → node never initialized → no InternalIP
  2. Cilium agent can't start without the node IP → never becomes ready
  3. operator never removes the agent-not-ready taint → back to (1)

A permanent deadlock. The CCM pod never appears at all and the cluster fails to validate.

Why it's a race, and why only some clouds lose it

The agent-not-ready taint only blocks new scheduling — it never evicts an already-bound pod. So the outcome hinges entirely on whether the CCM pod is bound to the control-plane node before the cilium-operator taints it.

This is decided during control-plane bootstrap by which static-pod component wins a restart-timing race:

  • The kube-scheduler and kube-controller-manager both crash-loop on startup until the kube-apiserver has published the extension-apiserver-authentication configmap (identical behavior on every cloud).
  • On a healthy/fast control plane, the scheduler is up and caches-synced before the daemonset-controller creates the CCM pod, so the CCM pod binds immediately — well before the operator's taint controller (which has a built-in ~11s leader-election/startup delay) ever runs.
  • On a slower/less-stable control plane, the apiserver restarts an extra time, the scheduler crash-loops one or two extra cycles, and CrashLoopBackOff (10s → 20s → 40s → 80s) turns that into a large gap. The CCM/cilium pods sit Pending with no scheduler; when the scheduler finally comes up it binds the all-tolerating Cilium pod to the control-plane node first, the operator taints the node, and the CCM pod is locked out.

Observed on the failing DO run vs. a passing AWS run:

apiserver starts scheduler restarts KCM healthy scheduler healthy gap result
AWS 2 2 09:54:44 09:54:44 ~0s CCM binds, wins
DO 3 4 10:38:53 10:39:38 ~45s operator taints first, deadlock

Because the margin is environmental (control-plane bring-up speed, not config), the same DO job passes on runs where the control plane comes up quickly — which is exactly why it's flaky rather than always-failing. The placement constraints, node-IP handling, Cilium routing mode (tunnel/ipam: kubernetes), and addon apply order are all identical between the passing and failing clouds, so none of those explain the difference — it is purely the bootstrap restart-timing race.

Every external CCM in kops shares this latent exposure; they just run on clouds whose control planes usually win the race. Making the CCM tolerate all taints removes the dependency on winning that window entirely.

Testing

  • Regenerated integration golden files (hack/update-expected.sh); manifestHashes update so existing clusters re-apply the new manifests automatically.
  • go test ./cmd/kops/ and the bootstrapchannelbuilder tests pass.

🤖 Generated with Claude Code

@kubernetes-prow

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@kubernetes-prow kubernetes-prow Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. labels Jun 26, 2026
@kubernetes-prow
kubernetes-prow Bot requested review from hakman and zetaab June 26, 2026 01:51
@rifelpet
rifelpet marked this pull request as ready for review June 26, 2026 01:52
@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jun 26, 2026
@rifelpet

Copy link
Copy Markdown
Member Author

/hold for feedback

@kubernetes-prow kubernetes-prow Bot added the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Jun 26, 2026
@rifelpet

Copy link
Copy Markdown
Member Author

/hold cancel

@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Jun 28, 2026
@kubernetes-prow kubernetes-prow Bot added the lgtm "Looks good to me", indicates that a PR is ready to be merged. label Jun 29, 2026
@kubernetes-prow

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: hakman

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow kubernetes-prow Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Jun 29, 2026
@kubernetes-prow
kubernetes-prow Bot merged commit 95f5a62 into kubernetes:master Jun 29, 2026
29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/addons cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. lgtm "Looks good to me", indicates that a PR is ready to be merged. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants