Sitelet https://github.com/kubernetes/kubernetes/pull/133461
Skip to content

fix using stale pod when evict failed and retry - #133461

Merged
k8s-ci-robot merged 6 commits into
kubernetes:masterfrom
justlorain:fix/drain_refresh_stale_pod
Aug 28, 2025
Merged

k8s-ci-robot merged 6 commits into
kubernetes:masterfrom
justlorain:fix/drain_refresh_stale_pod

Conversation

@justlorain

@justlorain justlorain commented Aug 11, 2025 •

Copy link
Copy Markdown
Contributor

What type of PR is this?

/kind bug

What this PR does / why we need it:

The kubectl drain command still uses stale Pods when it fails and retries after evicting Pods.
It should re-fetch the latest Pods when failing and retrying after evicting Pods.

Which issue(s) this PR is related to:

Fixes kubernetes/kubectl#1767

Special notes for your reviewer:

Does this PR introduce a user-facing change?

NONE

Additional documentation e.g., KEPs (Kubernetes Enhancement Proposals), usage docs, etc.:


@k8s-ci-robot k8s-ci-robot added the release-note-none Denotes a PR that doesn't merit a release note. label Aug 11, 2025
@k8s-ci-robot

Copy link
Copy Markdown
Contributor

Please note that we're already in Test Freeze for the release-1.34 branch. This means every merged PR will be automatically fast-forwarded via the periodic ci-fast-forward job to the release branch of the upcoming v1.34.0 release.

Fast forwards are scheduled to happen every 6 hours, whereas the most recent run was: Mon Aug 11 10:25:10 UTC 2025.

@k8s-ci-robot k8s-ci-robot added kind/bug Categorizes issue or PR as related to a bug. size/XS Denotes a PR that changes 0-9 lines, ignoring generated files. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. labels Aug 11, 2025
@k8s-ci-robot k8s-ci-robot added area/kubectl sig/cli Categorizes an issue or PR as relevant to SIG CLI. do-not-merge/needs-sig Indicates an issue or PR lacks a `sig/foo` label and requires one. needs-triage Indicates an issue or PR lacks a `triage/foo` label and requires one. labels Aug 11, 2025
@k8s-ci-robot

Copy link
Copy Markdown
Contributor

This issue is currently awaiting triage.

If a SIG or subproject determines this is a relevant issue, they will accept it by applying the triage/accepted label and provide further guidance.

The triage/accepted label can be added by org members by writing /triage accepted in a comment.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@k8s-ci-robot

Copy link
Copy Markdown
Contributor

Welcome @justlorain!

It looks like this is your first PR to kubernetes/kubernetes 🎉. Please refer to our pull request process documentation to help your PR have a smooth ride to approval.

You will be prompted by a bot to use commands during the review process. Do not be afraid to follow the prompts! It is okay to experiment. Here is the bot commands documentation.

You can also check if kubernetes/kubernetes has its own contribution guidelines.

You may want to refer to our testing guide if you run into trouble with your tests not passing.

If you are having difficulty getting your pull request seen, please follow the recommended escalation practices. Also, for tips and tricks in the contribution process you may want to read the Kubernetes contributor cheat sheet. We want to make sure your contribution gets all the attention it needs!

Thank you, and welcome to Kubernetes. 😃

@k8s-ci-robot k8s-ci-robot added needs-ok-to-test Indicates a PR that requires an org member to verify it is safe to test. needs-priority Indicates a PR lacks a `priority/foo` label and requires one. and removed do-not-merge/needs-sig Indicates an issue or PR lacks a `sig/foo` label and requires one. labels Aug 11, 2025
@github-project-automation github-project-automation Bot moved this to Needs Triage in SIG CLI Aug 11, 2025
@k8s-ci-robot

Copy link
Copy Markdown
Contributor

Hi @justlorain. Thanks for your PR.

I'm waiting for a kubernetes member to verify that this patch is reasonable to test. If it is, they should reply with /ok-to-test on its own line. Until that is done, I will not automatically test new commits in this PR, but the usual testing commands by org members will still work. Regular contributors should join the org to skip this step.

Once the patch is verified, the new status will be reflected by the ok-to-test label.

I understand the commands that are listed here.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@justlorain

Copy link
Copy Markdown
Contributor Author

/cc @ardaguclu

@k8s-ci-robot
k8s-ci-robot requested a review from ardaguclu August 14, 2025 03:04
@ardaguclu

Copy link
Copy Markdown
Member

/ok-to-test
Thank you for opening this PR. I think, we definitely need new tests to verify the behavior. Test should be failing before the patch and should be passing after the patch.

@k8s-ci-robot k8s-ci-robot added ok-to-test Indicates a non-member PR verified by an org member that is safe to test. and removed needs-ok-to-test Indicates a PR that requires an org member to verify it is safe to test. labels Aug 14, 2025
@justlorain

Copy link
Copy Markdown
Contributor Author

/ok-to-test Thank you for opening this PR. I think, we definitely need new tests to verify the behavior. Test should be failing before the patch and should be passing after the patch.

Sure, I'll add the relevant tests.

@brianpursley

Copy link
Copy Markdown
Member

Thanks for the PR.

Looking more closely at the code in drain.go, I have a few thoughts... Not about your changes, but about the entire evictPods func.

Mainly, I'm unclear why refreshPod is even needed.

The code clearly isn't working as intended right now, since refreshPod defaults to false and is never set to true (even the verify CI job detects it as a problem now it looks like with "ineffectual assignment to refreshPod" on line 309).

But I think we can achieve the desired behavior of refreshing the pod by simply moving the getPodFn() call to the end of the for loop block and then getting rid of the refreshPod variable altogether.

If the code makes it through the error handling without returning or breaking, then (I think) we should always attempt to re-get the pod.

So the first time through the loop, we use the original pod (set activePod := pod before the for loop). And then at the end of the for loop block, if the code execution makes it there, call getPodFn() and update activePod if it succeeds.

What do you think? Does this make sense or am I missing something? I feel like this logic would be more straightforward since it relies on regular code flow instead of having to manage and check a boolean.

@k8s-ci-robot k8s-ci-robot added size/S Denotes a PR that changes 10-29 lines, ignoring generated files. and removed size/XS Denotes a PR that changes 0-9 lines, ignoring generated files. labels Aug 18, 2025
@justlorain

Copy link
Copy Markdown
Contributor Author

Thanks for the PR.

Looking more closely at the code in drain.go, I have a few thoughts... Not about your changes, but about the entire evictPods func.

Mainly, I'm unclear why refreshPod is even needed.

The code clearly isn't working as intended right now, since refreshPod defaults to false and is never set to true (even the verify CI job detects it as a problem now it looks like with "ineffectual assignment to refreshPod" on line 309).

But I think we can achieve the desired behavior of refreshing the pod by simply moving the getPodFn() call to the end of the for loop block and then getting rid of the refreshPod variable altogether.

If the code makes it through the error handling without returning or breaking, then (I think) we should always attempt to re-get the pod.

So the first time through the loop, we use the original pod (set activePod := pod before the for loop). And then at the end of the for loop block, if the code execution makes it there, call getPodFn() and update activePod if it succeeds.

What do you think? Does this make sense or am I missing something? I feel like this logic would be more straightforward since it relies on regular code flow instead of having to manage and check a boolean.

Thanks for your review! I think your approach is completely reasonable and intuitive for me. I've tried to implement a new commit based on it - could you please check if it aligns with your thinking?

@brianpursley

Copy link
Copy Markdown
Member

I think the changes look good.

A few more things:

  1. Can you write a unit test for this? This will help demonstrate it is working as expected and prevent a future regression.
  2. The lint CI job flagged a couple of problems you'll need to handle since you touched those Fprintf lines. You can use _, _ = or //nolint:errcheck on the Fprintf lines to handle it.
  3. I don't know if the other failed CI tests are legit or not. You can re-run the tests but if the failures are legit you'll need to check why.

@k8s-ci-robot k8s-ci-robot added size/M Denotes a PR that changes 30-99 lines, ignoring generated files. and removed size/S Denotes a PR that changes 10-29 lines, ignoring generated files. labels Aug 22, 2025
@k8s-ci-robot k8s-ci-robot added size/L Denotes a PR that changes 100-499 lines, ignoring generated files. and removed size/M Denotes a PR that changes 30-99 lines, ignoring generated files. labels Aug 25, 2025
@justlorain

Copy link
Copy Markdown
Contributor Author

I think the changes look good.

A few more things:

  1. Can you write a unit test for this? This will help demonstrate it is working as expected and prevent a future regression.
  2. The lint CI job flagged a couple of problems you'll need to handle since you touched those Fprintf lines. You can use _, _ = or //nolint:errcheck on the Fprintf lines to handle it.
  3. I don't know if the other failed CI tests are legit or not. You can re-run the tests but if the failures are legit you'll need to check why.

Hi, I've added the relevant unit tests and resolved the CI issues. Could you please review it again?

@brianpursley

Copy link
Copy Markdown
Member

Excellent work.

The changes you made look good and thank you for writing the unit tests.

The only thing I noticed is that the unit test takes 15s to run. Can you make some changes to improve the test speed?

One approach might be to change the hard-coded 5 * time.Second (in drain.go) into a variable so the unit test can change it.

For example:

var drainErrorRetryDelay = 5 * time.Second

And then something like this in the two places (here and here) where it is used. For example:

fmt.Fprintf(d.ErrOut, "error when evicting pods/%q -n %q (will retry after %v): %v\n", activePod.Name, activePod.Namespace, drainErrorRetryDelay, err)
time.Sleep(drainErrorRetryDelay)

And finally, in the unit test, you can then reduce the delay to 1ms and also adjust the globalTimeout variable to make the test run faster:

drainErrorRetryDelay = 1 * time.Millisecond
globalTimeout := drainErrorRetryDelay * 2

What do you think?

@justlorain

Copy link
Copy Markdown
Contributor Author

Excellent work.

The changes you made look good and thank you for writing the unit tests.

The only thing I noticed is that the unit test takes 15s to run. Can you make some changes to improve the test speed?

One approach might be to change the hard-coded 5 * time.Second (in drain.go) into a variable so the unit test can change it.

For example:

var drainErrorRetryDelay = 5 * time.Second

And then something like this in the two places (here and here) where it is used. For example:

fmt.Fprintf(d.ErrOut, "error when evicting pods/%q -n %q (will retry after %v): %v\n", activePod.Name, activePod.Namespace, drainErrorRetryDelay, err)
time.Sleep(drainErrorRetryDelay)

And finally, in the unit test, you can then reduce the delay to 1ms and also adjust the globalTimeout variable to make the test run faster:

drainErrorRetryDelay = 1 * time.Millisecond
globalTimeout := drainErrorRetryDelay * 2

What do you think?

I'm thinking if would be better to introduce the retry delay as a field of the Helper.

// Helper contains the parameters to control the behaviour of drainer
type Helper struct {
According to the comments, Helper is also described as a struct used to uniformly control drainer behavior, and it meets the need for configuration in unit tests.

Another consideration I have is whether it's necessary to make the retry delay setting a new option for the drain command. In that case, having it as a field of Helper would also maintain consistency with the style of other options. However, this consideration seems to have already moved beyond the scope of what this PR is discussing.

What do you think?

@brianpursley

Copy link
Copy Markdown
Member

I'm thinking if would be better to introduce the retry delay as a field of the Helper.

That seems like it should work too. In that case though, it would be up to the caller to set the retry duration. Not a problem for kubectl, but if any other tool uses this drain helper, it will need to know to set that retry duration to 5 seconds or else the behavior would be different. I do think the existing retry behavior is broken right now anyway though, so it's probably not an issue.

Another consideration I have is whether it's necessary to make the retry delay setting a new option for the drain command. In that case, having it as a field of Helper would also maintain consistency with the style of other options. However, this consideration seems to have already moved beyond the scope of what this PR is discussing.

Yeah, I agree and don't think we need to expose retry delay as a command option at this point. If it is something that someone really needs for some reason, they can open an issue for it and we can decide whether to add it then. I think you're right... let's keep this PR focused on fixing the bug.

@justlorain

Copy link
Copy Markdown
Contributor Author

The only thing I noticed is that the unit test takes 15s to run. Can you make some changes to improve the test speed?

Attempted to introduce EvictErrorRetryDelay field in Helper to resolve this. Could you please check again to see if there are any issues?

@brianpursley brianpursley left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks!

/lgtm
/label tide/merge-method-squash

@k8s-ci-robot k8s-ci-robot added the tide/merge-method-squash Denotes a PR that should be squashed by tide when it merges. label Aug 27, 2025
@k8s-ci-robot k8s-ci-robot added the lgtm "Looks good to me", indicates that a PR is ready to be merged. label Aug 27, 2025
@k8s-ci-robot

Copy link
Copy Markdown
Contributor

LGTM label has been added.

DetailsGit tree hash: 8397e0a4cdae6936f6551ced025d9c99b9ee55a1

@k8s-ci-robot

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: brianpursley, justlorain

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@k8s-ci-robot k8s-ci-robot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 27, 2025
@k8s-ci-robot
k8s-ci-robot merged commit 66fdbe1 into kubernetes:master Aug 28, 2025
13 checks passed
@k8s-ci-robot k8s-ci-robot added this to the v1.35 milestone Aug 28, 2025
@github-project-automation github-project-automation Bot moved this from Needs Triage to Done in SIG CLI Aug 28, 2025
fusida pushed a commit to fusida/kubernetes that referenced this pull request Sep 10, 2025
* fix using stale pod when evict failed and retry

* simplify pod refresh process

* use activePod at getPodFn

* fix lint check

* add ut

* introduce EvictErrorRetryDelay
bwsalmon pushed a commit to bwsalmon/kubernetes that referenced this pull request Oct 22, 2025
* fix using stale pod when evict failed and retry

* simplify pod refresh process

* use activePod at getPodFn

* fix lint check

* add ut

* introduce EvictErrorRetryDelay
mhan8796 pushed a commit to mhan8796/kubernetes that referenced this pull request Jun 27, 2026
* fix using stale pod when evict failed and retry

* simplify pod refresh process

* use activePod at getPodFn

* fix lint check

* add ut

* introduce EvictErrorRetryDelay
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/kubectl cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. kind/bug Categorizes issue or PR as related to a bug. lgtm "Looks good to me", indicates that a PR is ready to be merged. needs-priority Indicates a PR lacks a `priority/foo` label and requires one. needs-triage Indicates an issue or PR lacks a `triage/foo` label and requires one. ok-to-test Indicates a non-member PR verified by an org member that is safe to test. release-note-none Denotes a PR that doesn't merit a release note. sig/cli Categorizes an issue or PR as relevant to SIG CLI. size/L Denotes a PR that changes 100-499 lines, ignoring generated files. tide/merge-method-squash Denotes a PR that should be squashed by tide when it merges.

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Using stale Pod when eviction fails and retries

4 participants