Sitelet https://github.com/kubernetes/kops/issues/18453
Skip to content

kops on azure cannot schedule coredns pod while cluster creation #18453

Description

@Aditya-017

/kind bug

1. What kops version are you running? The command kops version, will display
this information.

I0607 21:40:59.326981 31818 featureflag.go:178] FeatureFlag "Azure"=true
I0607 21:40:59.327640 31818 featureflag.go:187] ParseFlags: parsed 1 flags from "Azure" (unknown=0, registered=25)
Client version: 1.35.1

2. What Kubernetes version are you running? kubectl version will print the
version if a cluster is running or provide the Kubernetes version specified as
a kops flag.

Client Version: v1.36.1
Kustomize Version: v5.8.1

3. What cloud provider are you using?
Azure

4. What commands did you run? What is the simplest way to reproduce this issue?
Follow this official guide: https://kops.sigs.k8s.io/getting_started/azure/
kops create cluster --cloud azure --name tute.k8s --zones northeurope-1 --azure-admin-user ubuntu --yes

5. What happened after the commands executed?
The process of cluster creation started and the logs showed no such persistent errors.
Next step was to validate using kops validate cluster --name tute.k8s --wait=10m
Validation failed, cluster not yet ready even after an hour.

I0607 22:08:16.775522   35008 featureflag.go:178] FeatureFlag "Azure"=true
I0607 22:08:16.776323   35008 featureflag.go:187] ParseFlags: parsed 1 flags from "Azure" (unknown=0, registered=25)
Validating cluster <cluster-name>

INSTANCE GROUPS
NAME				ROLE		MACHINETYPE	MIN	MAX	SUBNETS
control-plane-northeurope-1	ControlPlane	Standard_B2s	1	1	<cluster-name>
nodes-northeurope-1		Node		Standard_B2s	1	1	<cluster-name>

NODE STATUS
NAME					ROLE		READY
control-plane-northeurope-1000000	control-plane	True
nodes-northeurope-1000000		node		True

VALIDATION ERRORS
KIND	NAME					MESSAGE
Pod	kube-system/coredns-5f977fddbc-tnnk7	system-cluster-critical pod "coredns-5f977fddbc-tnnk7" is pending

Validation Failed
W0607 22:08:24.399771   35008 validate_cluster.go:241] (will retry): cluster not yet healthy
I

6. What did you expect to happen?
Cluster to be created and validations to be passing succesfully.

7. Please provide your cluster manifest. Execute
kops get --name my.example.com -o yaml to display your cluster manifest.
You may want to remove your cluster name and other sensitive information.

apiVersion: kops.k8s.io/v1alpha2
kind: Cluster
metadata:
  creationTimestamp: "2026-06-07T16:23:26Z"
  name: tute.k8s
spec:
  api:
    loadBalancer:
      type: Public
  authorization:
    rbac: {}
  channel: stable
  cloudConfig:
    azure:
      adminUser: <admin-user>
      storageAccountID: /subscriptions/<subscription-id>/resourceGroups/<rg-name>/providers/Microsoft.Storage/storageAccounts/<storage-name>
      subscriptionId: <subscription-id>
      tenantId: ""
  cloudProvider: azure
  configBase: azureblob://<storage-name>/<blob-container-name>/tute.k8s
  etcdClusters:
  - cpuRequest: 200m
    etcdMembers:
    - instanceGroup: control-plane-northeurope-1
      name: etcd-1
    manager:
      backupRetentionDays: 90
    memoryRequest: 100Mi
    name: main
  - cpuRequest: 100m
    etcdMembers:
    - instanceGroup: control-plane-northeurope-1
      name: etcd-1
    manager:
      backupRetentionDays: 90
    memoryRequest: 100Mi
    name: events
  iam:
    allowContainerRegistry: true
    legacy: false
  kubeProxy:
    enabled: false
  kubelet:
    anonymousAuth: false
  kubernetesApiAccess:
  - 0.0.0.0/0
  - ::/0
  kubernetesVersion: 1.35.5
  networkCIDR: 10.0.0.0/16
  networking:
    cilium:
      enableNodePort: true
  nonMasqueradeCIDR: 100.64.0.0/10
  sshAccess:
  - 0.0.0.0/0
  - ::/0
  subnets:
  - cidr: 10.0.0.0/16
    name: tute.k8s
    region: northeurope
    type: Public
  topology:
    dns:
      type: None

---

apiVersion: kops.k8s.io/v1alpha2
kind: InstanceGroup
metadata:
  creationTimestamp: "2026-06-07T16:23:26Z"
  labels:
    kops.k8s.io/cluster: tute.k8s
  name: control-plane-northeurope-1
spec:
  image: Canonical:ubuntu-24_04-lts:server:24.04.202507010
  machineType: Standard_B2s
  maxSize: 1
  minSize: 1
  role: Master
  subnets:
  - tute.k8s
  zones:
  - northeurope-1

---

apiVersion: kops.k8s.io/v1alpha2
kind: InstanceGroup
metadata:
  creationTimestamp: "2026-06-07T16:23:26Z"
  labels:
    kops.k8s.io/cluster: tute.k8s
  name: nodes-northeurope-1
spec:
  image: Canonical:ubuntu-24_04-lts:server:24.04.202507010
  machineType: Standard_B2s
  maxSize: 1
  minSize: 1
  role: Node
  subnets:
  - tute.k8s
  zones:
  - northeurope-1

8. Please run the commands with most verbose logging by adding the -v 10 flag.
Paste the logs into this report, or in a gist and provide the gist link here.

Command: k describe pod coredns-5f977fddbc-tnnk7 -n kube-system

Tolerations:                  CriticalAddonsOnly op=Exists
                              node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                              node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
Topology Spread Constraints:  kubernetes.io/hostname:DoNotSchedule when max skew 1 is exceeded for selector k8s-app=kube-dns
                              topology.kubernetes.io/zone:ScheduleAnyway when max skew 1 is exceeded for selector k8s-app=kube-dns
Events:
  Type     Reason            Age    From               Message
  ----     ------            ----   ----               -------
  Warning  FailedScheduling  8m24s  default-scheduler  0/2 nodes are available: 1 node(s) didn't match pod topology spread constraints, 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: 0/2 nodes are available: 1 No preemption victims found for incoming pod, 1 Preemption is not helpful for scheduling.
  Warning  FailedScheduling  3m20s  default-scheduler  0/2 nodes are available: 1 node(s) didn't match pod topology spread constraints, 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: 0/2 nodes are available: 1 No preemption victims found for incoming pod, 1 Preemption is not helpful for scheduling.

Command: k describe node control-plane-northeurope-1000000
We see
Taints: node-role.kubernetes.io/control-plane:NoSchedule

9. Anything else do we need to know?
Fixed it by updating the IG manifest and changing min and max to 2.
kops edit ig nodes-northeurope-1 --name=tute.k8s

Other option I explored was
a) Add toleration in coredns pod
b) remove taint from control plane node.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

kind/bugCategorizes issue or PR as related to a bug.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions