Tolerate unready nodes during scalability test validation - #18616
Conversation
|
Skipping CI for Draft Pull Request. |
| KUBETEST2_ARGS+=("--max-nodes-to-dump=${MAX_NODES_TO_DUMP:-5}") | ||
| KUBETEST2_ARGS+=("--node-dump-timeout=${NODE_DUMP_TIMEOUT:-5m}") | ||
| # Tolerate a handful of worker nodes failing to join at scale, matching kubeup's ALLOWED_NOTREADY_NODES. | ||
| KUBETEST2_ARGS+=("--max-unready-nodes=${MAX_UNREADY_NODES:-5}") |
There was a problem hiding this comment.
Are you sure 5 is not too much as default?
There was a problem hiding this comment.
The old kube-up script had 1% so 50 of 5000 😅. Updated this to only default to 5 for 5k jobs since 5/100 node runs is too much.
I will defer to @serathus for the question of whether 5 makes sense as default.
There was a problem hiding this comment.
Spoke to @serathius offline, we'll be conservative and start with 2 since we've only see 1/5000 fail recently.
58f453e to
1ebd1b5
Compare
|
/assign @serathius |
1ebd1b5 to
808ed02
Compare
|
/lgtm |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: hakman The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
Pass --max-unready-nodes to
kops validate clusterin the scalability scenario so 5k scale runs tolerate a few worker nodes that fail to join.Testing
Verified on the 100-node scale presubmit that the deployer forwards the flag to the real validate call and the job passed: https://prow.k8s.io/view/gs/kubernetes-ci-logs/pr-logs/pull/kops/18616/pull-kops-gce-master-scale-performance-100/2079611887690977280