Sitelet https://github.com/kubernetes/kops/pull/18684
Skip to content

Fix node bootstrap in IPv6-only clusters with --dns=none - #18684

Merged
kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:ipv6-dns-none-bootstrap
Aug 13, 2026
Merged

kubernetes-prow[bot] merged 2 commits into
kubernetes:masterfrom
rifelpet:ipv6-dns-none-bootstrap

Conversation

@rifelpet

Copy link
Copy Markdown
Member

What

Nodes never join an IPv6-only cluster that uses --dns=none. Two independent bugs; one commit each.

Why

e2e-kops-aws-ipv6-karpenter started failing as soon as it began actually creating an IPv6 cluster (#18681 plus kubernetes/test-infra#37675 — before that it was silently building an IPv4 cluster). Every node fails to bootstrap:

Post "https://172.20.6.26:3988/bootstrap": dial tcp 172.20.6.26:3988: connect: network is unreachable

The generated EC2NodeClass userData shows why:

APIServerIPs:
- 172.20.6.26
- 2a05:d01c:f3c:6402:979a:1f3c:2851:c9a2
ConfigServer:
  servers:
  - https://172.20.6.26:3988/
  - https://[2a05:d01c:f3c:6402:979a:1f3c:2851:c9a2]:3988/
  1. The IPv4 address should not be there. With --dns=none, nodeupconfigbuilder fills APIServerIPs from the API NLB's ENI addresses, keeping anything inside the network CIDR or any IPv6 address — so a dualstack NLB contributes both. Nodes in an IPv6-only cluster are in subnets with no IPv4 CIDR and cannot reach 172.20.6.26 at all. These addresses also become the /etc/hosts entries for api.internal and kops-controller.internal.

  2. nodeup never reaches the second server. getNodeConfigFromServers walks the server list, but Client.Query retries one server 100 times with a capped backoff — roughly 50 minutes. FindAddresses sorts the addresses as strings, so the IPv4 one sorts first and absorbs the entire budget. The multi-server list is effectively single-server whenever the first entry is down.

Either fix alone unblocks the job; both are worth having.

Coverage gap

IPv6 + dns=none is untested. Every IPv6 job passes --dns=public, and no job in kops-periodics-dns-none.yaml / kops-presubmits-dns-none.yaml is IPv6. e2e-kops-aws-ipv6-karpenter hit it only because it omits --dns=public and so picked up the dns=none default from setupDNSTopology.

This PR adds TestMinimalIPv6DNSNone, the first ipv6 + dns=none integration test. It does not exercise the fix directly — the integration test mock never populates load balancer addresses, so APIServerIPs comes out empty and the config falls back to the DNS name (minimal-dns-none has the same gap). The address selection is therefore pulled out into selectControlPlaneIPs and unit tested.

Testing

  • go test ./cmd/kops/... ./pkg/nodemodel/... ./upup/pkg/fi/nodeup/... ./util/pkg/vfs/... passes.
  • Not yet verified end to end. Once this is out of draft: /test pull-kops-e2e-aws-ipv6-karpenter — the only presubmit that covers this combination. Note it is always_run: false with its run_if_changed commented out, so it has to be requested explicitly.

With --dns=none, nodeupconfigbuilder puts the API load balancer's addresses
into BootConfig.APIServerIPs, which become the kops-controller bootstrap
endpoints and the /etc/hosts entries for api.internal and
kops-controller.internal. On AWS an address qualifies if it is inside the
network CIDR or if it is IPv6, so a dualstack NLB contributes both its private
IPv4 address and its IPv6 address.

Nodes in an IPv6-only cluster sit in subnets that have no IPv4 CIDR, so the
IPv4 address is unroutable from them:

  Post "https://172.20.6.26:3988/bootstrap": dial tcp 172.20.6.26:3988:
  connect: network is unreachable

FindAddresses sorts the addresses as strings, which puts the IPv4 one first,
so that is the endpoint nodeup spends its retry budget on and no node ever
joins. Drop IPv4 addresses for IPv6-only clusters.

This combination has no coverage today: every IPv6 job passes --dns=public,
and no dns-none job is IPv6. e2e-kops-aws-ipv6-karpenter reached it first
because it does not pass --dns=public and so picked up the dns=none default.
Add an ipv6 + dns=none integration test, and pull the address selection out
into selectControlPlaneIPs so it can be unit tested -- the integration test
mock does not populate load balancer addresses, so it cannot exercise this.
BootConfig.ConfigServer.Servers can hold more than one endpoint, and
getNodeConfigFromServers walks the list -- but Client.Query retries a single
server 100 times with a capped exponential backoff, roughly 50 minutes. Any
server that is permanently unreachable, such as an IPv4 address in an
IPv6-only cluster, therefore consumes the whole budget and the remaining
entries are never tried.

Give the client a configurable Backoff, and have getNodeConfigFromServers set
a short one so each server gets about a minute before we move on. Keep asking
until the overall bootstrap timeout, cycling the list rather than giving up
after one pass, so a control plane that is slow to come up is still waited
out. BootstrapClientTask keeps the previous behaviour: it has a single server
and does not set Backoff.
@kubernetes-prow

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@kubernetes-prow kubernetes-prow Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. labels Aug 13, 2026
@kubernetes-prow
kubernetes-prow Bot requested a review from hakman August 13, 2026 12:51
@kubernetes-prow
kubernetes-prow Bot requested a review from zetaab August 13, 2026 12:51
@rifelpet

Copy link
Copy Markdown
Member Author

/test pull-kops-e2e-aws-ipv6-karpenter

@rifelpet
rifelpet marked this pull request as ready for review August 13, 2026 13:42
@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 13, 2026
@rifelpet

Copy link
Copy Markdown
Member Author

/cc @hakman

@rifelpet

Copy link
Copy Markdown
Member Author

confirmed the karpenter job passed and the EC2NodeClasses have the expected nodeup config:

https://storage.googleapis.com/kubernetes-ci-logs/pr-logs/pull/kops/18684/pull-kops-e2e-aws-ipv6-karpenter/2087885134400327680/artifacts/cluster-info/karpenter.k8s.aws.ec2nodeclasses.yaml

      cat > conf/kube_env.yaml << '__EOF_KUBE_ENV'
      APIServerIPs:
      - 2600:1f16:284:9002:71e4:b233:3211:5cee
      CloudProvider: aws
      ClusterName: pr18684-kops-e2e-aws-ipv6-karpenter.tests-kops-aws.k8s.io
      ConfigServer:
        servers:
        - https://[2600:1f16:284:9002:71e4:b233:3211:5cee]:3988/
        tlsServerName: kops-controller.internal.pr18684-kops-e2e-aws-ipv6-karpenter.tests-kops-aws.k8s.io
      InstanceGroupName: nodes
      InstanceGroupRole: Node

@hakman

hakman commented Aug 13, 2026

Copy link
Copy Markdown
Member

confirmed the karpenter job passed and the EC2NodeClasses have the expected nodeup config:

https://storage.googleapis.com/kubernetes-ci-logs/pr-logs/pull/kops/18684/pull-kops-e2e-aws-ipv6-karpenter/2087885134400327680/artifacts/cluster-info/karpenter.k8s.aws.ec2nodeclasses.yaml

      cat > conf/kube_env.yaml << '__EOF_KUBE_ENV'
      APIServerIPs:
      - 2600:1f16:284:9002:71e4:b233:3211:5cee
      CloudProvider: aws
      ClusterName: pr18684-kops-e2e-aws-ipv6-karpenter.tests-kops-aws.k8s.io
      ConfigServer:
        servers:
        - https://[2600:1f16:284:9002:71e4:b233:3211:5cee]:3988/
        tlsServerName: kops-controller.internal.pr18684-kops-e2e-aws-ipv6-karpenter.tests-kops-aws.k8s.io
      InstanceGroupName: nodes
      InstanceGroupRole: Node

Nice!

@kubernetes-prow kubernetes-prow Bot added the lgtm "Looks good to me", indicates that a PR is ready to be merged. label Aug 13, 2026
@kubernetes-prow

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: hakman

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow kubernetes-prow Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 13, 2026
@kubernetes-prow
kubernetes-prow Bot merged commit 47827ce into kubernetes:master Aug 13, 2026
28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/nodeup cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. lgtm "Looks good to me", indicates that a PR is ready to be merged. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants