/kind bug
1. What kops version are you running? The command kops version, will display
this information.
1.35.0 upgraded to 1.35.1
2. What Kubernetes version are you running? kubectl version will print the
version if a cluster is running or provide the Kubernetes version specified as
a kops flag.
1.35.5
3. What cloud provider are you using?
AWS
4. What commands did you run? What is the simplest way to reproduce this issue?
Upgrade kops-1.35.0 to kops-1.35.1, which includes upgrade of cluster-autoscaler 1.34.1 to 1.35.0 (see #18217)
5. What happened after the commands executed?
cluster-autoscaler has these errors:
I0603 15:47:40.036244 1 reflector.go:429] "Data couldn't be fetched in watchlist mode. Falling back to regular list. This is expected if watchlist is not supported or disabled in kube-apiserver." err="resourceclaims.resource.k8s.io is forbidden: User \"system:serviceaccount:kube-system:cluster-autoscaler\" cannot watch resource \"resourceclaims\" in API group \"resource.k8s.io\" at the cluster scope"
6. What did you expect to happen?
In cluster-autoscaler 1.35.0 they added a few more entries to ClusterRole/cluster-autoscaler:
- resource.k8s.io
resources:
- resourceslices
- deviceclasses
- resourceclaims
verbs:
- watch
- list
- get
(see https://github.com/kubernetes/autoscaler/pull/8827/changes#diff-548869636fdb980737339b1e3c701a64248184e045b4305dc96e4c66287bd272 for that change in their helm chart)
Our kops-provided clusterrole does not have these after upgrade, and I don't see them in #18217. Is this something that kops is responsible for upgrading?
7. Please provide your cluster manifest. Execute
kops get --name my.example.com -o yaml to display your cluster manifest.
You may want to remove your cluster name and other sensitive information.
apiVersion: kops.k8s.io/v1alpha2
kind: Cluster
metadata:
name: zzz
spec:
...
clusterAutoscaler:
enabled: true
balanceSimilarNodeGroups: false
cordonNodeBeforeTerminating: true
expander: least-waste
scaleDownDelayAfterAdd: 10m0s
scaleDownUnneededTime: 5m0s
scaleDownUnreadyTime: 10m0s
scaleDownUtilizationThreshold: "0.85"
skipNodesWithLocalStorage: false
skipNodesWithSystemPods: true
8. Please run the commands with most verbose logging by adding the -v 10 flag.
Paste the logs into this report, or in a gist and provide the gist link here.
9. Anything else do we need to know?
Appending new resources mentioned above manually to kubectl edit clusterrole cluster-autoscaler solves the problem and makes cluster-autoscaler operable again.
/kind bug
1. What
kopsversion are you running? The commandkops version, will displaythis information.
1.35.0 upgraded to 1.35.1
2. What Kubernetes version are you running?
kubectl versionwill print theversion if a cluster is running or provide the Kubernetes version specified as
a
kopsflag.1.35.5
3. What cloud provider are you using?
AWS
4. What commands did you run? What is the simplest way to reproduce this issue?
Upgrade kops-1.35.0 to kops-1.35.1, which includes upgrade of cluster-autoscaler 1.34.1 to 1.35.0 (see #18217)
5. What happened after the commands executed?
cluster-autoscaler has these errors:
I0603 15:47:40.036244 1 reflector.go:429] "Data couldn't be fetched in watchlist mode. Falling back to regular list. This is expected if watchlist is not supported or disabled in kube-apiserver." err="resourceclaims.resource.k8s.io is forbidden: User \"system:serviceaccount:kube-system:cluster-autoscaler\" cannot watch resource \"resourceclaims\" in API group \"resource.k8s.io\" at the cluster scope"6. What did you expect to happen?
In cluster-autoscaler 1.35.0 they added a few more entries to ClusterRole/cluster-autoscaler:
(see https://github.com/kubernetes/autoscaler/pull/8827/changes#diff-548869636fdb980737339b1e3c701a64248184e045b4305dc96e4c66287bd272 for that change in their helm chart)
Our kops-provided clusterrole does not have these after upgrade, and I don't see them in #18217. Is this something that kops is responsible for upgrading?
7. Please provide your cluster manifest. Execute
kops get --name my.example.com -o yamlto display your cluster manifest.You may want to remove your cluster name and other sensitive information.
8. Please run the commands with most verbose logging by adding the
-v 10flag.Paste the logs into this report, or in a gist and provide the gist link here.
9. Anything else do we need to know?
Appending new resources mentioned above manually to
kubectl edit clusterrole cluster-autoscalersolves the problem and makes cluster-autoscaler operable again.