Sitelet https://github.com/kubernetes/kops/issues/18439
Skip to content

Upgrade to cluster-autoscaler 1.35.0 missing new clusterrole entries #18439

Description

@vitaliyf

/kind bug

1. What kops version are you running? The command kops version, will display
this information.

1.35.0 upgraded to 1.35.1

2. What Kubernetes version are you running? kubectl version will print the
version if a cluster is running or provide the Kubernetes version specified as
a kops flag.

1.35.5

3. What cloud provider are you using?

AWS

4. What commands did you run? What is the simplest way to reproduce this issue?

Upgrade kops-1.35.0 to kops-1.35.1, which includes upgrade of cluster-autoscaler 1.34.1 to 1.35.0 (see #18217)

5. What happened after the commands executed?

cluster-autoscaler has these errors:

I0603 15:47:40.036244 1 reflector.go:429] "Data couldn't be fetched in watchlist mode. Falling back to regular list. This is expected if watchlist is not supported or disabled in kube-apiserver." err="resourceclaims.resource.k8s.io is forbidden: User \"system:serviceaccount:kube-system:cluster-autoscaler\" cannot watch resource \"resourceclaims\" in API group \"resource.k8s.io\" at the cluster scope"

6. What did you expect to happen?

In cluster-autoscaler 1.35.0 they added a few more entries to ClusterRole/cluster-autoscaler:

  - resource.k8s.io
  resources:
    - resourceslices
    - deviceclasses
    - resourceclaims
  verbs:
    - watch
    - list
    - get

(see https://github.com/kubernetes/autoscaler/pull/8827/changes#diff-548869636fdb980737339b1e3c701a64248184e045b4305dc96e4c66287bd272 for that change in their helm chart)

Our kops-provided clusterrole does not have these after upgrade, and I don't see them in #18217. Is this something that kops is responsible for upgrading?

7. Please provide your cluster manifest. Execute
kops get --name my.example.com -o yaml to display your cluster manifest.
You may want to remove your cluster name and other sensitive information.

apiVersion: kops.k8s.io/v1alpha2
kind: Cluster
metadata:
  name: zzz
spec:
  ...
  clusterAutoscaler:
    enabled: true
    balanceSimilarNodeGroups: false
    cordonNodeBeforeTerminating: true
    expander: least-waste
    scaleDownDelayAfterAdd: 10m0s
    scaleDownUnneededTime: 5m0s
    scaleDownUnreadyTime: 10m0s
    scaleDownUtilizationThreshold: "0.85"
    skipNodesWithLocalStorage: false
    skipNodesWithSystemPods: true

8. Please run the commands with most verbose logging by adding the -v 10 flag.
Paste the logs into this report, or in a gist and provide the gist link here.

9. Anything else do we need to know?

Appending new resources mentioned above manually to kubectl edit clusterrole cluster-autoscaler solves the problem and makes cluster-autoscaler operable again.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    kind/bugCategorizes issue or PR as related to a bug.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions