tests/e2e: determine the SSH user from the cluster when not configured - #18699
Conversation
kubetest2-kops takes its SSH user from KUBE_SSH_USER, which the prow jobs set from a hand-maintained per-distro table. kops already knows the answer: it registers the SSH key under a username it derives from the image, and reports that as sshUser for every instance in "kops toolbox dump" -- guessSSHUser on AWS, SSHUsernameForImage on GCE, the VMSS admin username on Azure, root on DigitalOcean. Use it, but only as a last resort. The order is now the --ssh-user flag, then the users the deployer assigns itself for azure and digitalocean, then KUBE_SSH_USER, and only then the cluster. Every job that sets KUBE_SSH_USER keeps exactly the user it has today, so this changes nothing on merge; a job opts in by dropping that variable, which keeps the rollout per-job and revertable without touching kops. The lookup needs the instances to exist, so it runs in Up() once the cloud resources are created and before validation, so that a cluster which comes up but fails to validate is still reachable for log collection. Up() is not the only entry point, so DumpClusterLogs() also resolves the user for --down invocations in a fresh process. The environment handed to the tester is built during init(), before any of this is knowable, so extract that export and run it again once the user is known. Where no user can be determined -- an older kops that does not report sshUser, or a cluster that failed before creating instances -- leave it empty and omit --ssh-user entirely, so kops applies its own default rather than being handed an unusable empty value.
|
Skipping CI for Draft Pull Request. |
|
/test pull-kops-e2e-k8s-gce-cilium |
|
@rifelpet: The following test failed, say
Full PR test history. Your PR dashboard. Please help us cut down on flakes by linking to an open issue when you hit one in your PR. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
|
/test pull-kops-kubernetes-e2e-cos-gce-serial |
fdc3db6 to
28e30ce
Compare
The failure here was unrelated, the behavior was confirmed with inferring the |
|
Sorry i disrupted the retry with the temp commit. i can push the temp commit back again if you want. |
No, all good. |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: hakman The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
kubernetes/kops#18699 lets kubetest2-kops ask the cluster which user kops registered the SSH key for, but only when the job does not set KUBE_SSH_USER. Every GCE job sets it to "prow", so nothing exercises the new path yet. Add a derive_ssh_user option that omits the variable, and turn it on for two nftables periodics before touching the rest of GCE: e2e-kops-gce-nftables-cos125 expect "admin" e2e-kops-gce-nftables-u2404 expect "ubuntu" cos125 is the load-bearing case. kops defaults --ssh-user to "ubuntu" when it is not supplied, so a u2404 job cannot distinguish a working lookup from a broken one, while cos125 derives "admin" and fails visibly if discovery does not work. Both distros were confirmed to have that user created with sudo, holding the key kops installs in instance metadata. Only the user moves: KUBE_SSH_KEY_PATH and preset-k8s-ssh stay, so these jobs keep the shared CI key. The other 36 nftables jobs keep KUBE_SSH_USER as controls. Restricted to jobs on the master kops marker, because GCE only started reporting sshUser in kops 1.37 (c6e61af681) and that is not in the release branches. The rest of GCE becomes eligible as those branches age out of the matrix.
kubernetes/kops#18699 lets kubetest2-kops ask the cluster which user kops registered the SSH key for, but only when the job does not set KUBE_SSH_USER. Every GCE job sets it to "prow", so nothing exercises the new path yet. Add a derive_ssh_user option that omits the variable, and turn it on for two nftables periodics before touching the rest of GCE: e2e-kops-gce-nftables-cos125 expect "admin" e2e-kops-gce-nftables-u2404 expect "ubuntu" cos125 is the load-bearing case. kops defaults --ssh-user to "ubuntu" when it is not supplied, so a u2404 job cannot distinguish a working lookup from a broken one, while cos125 derives "admin" and fails visibly if discovery does not work. Both distros were confirmed to have that user created with sudo, holding the key kops installs in instance metadata. Only the user moves: KUBE_SSH_KEY_PATH and preset-k8s-ssh stay, so these jobs keep the shared CI key. The other 36 nftables jobs keep KUBE_SSH_USER as controls. Restricted to jobs on the master kops marker, because GCE only started reporting sshUser in kops 1.37 (c6e61af681) and that is not in the release branches. The rest of GCE becomes eligible as those branches age out of the matrix.
kubetest2-kopstakes its SSH user fromKUBE_SSH_USER, which the prow jobs set from a hand-maintained per-distro table (aws_distros_ssh_user, and a hardcodedprowfor GCE). kops already knows the answer — it registers the SSH key under a username it derives from the image, and reports that assshUserfor every instance inkops toolbox dump:guessSSHUser(pkg/resources/aws/aws.go)gce.SSHUsernameForImage(pkg/resources/gce/dump.go)OSProfile.AdminUsername(pkg/resources/azure/dump.go)root(pkg/resources/digitalocean/resources.go)This uses that, only as a last resort.
Precedence
--ssh-userflagkops, digitaloceanrootKUBE_SSH_USERPutting discovery last is the point. All 2,660 jobs that set
KUBE_SSH_USERkeep exactly the user they have today, so this changes nothing on merge and costs nothing — the lookup early-returns before shelling out whenever a user is already known. A job opts in by deleting itsKUBE_SSH_USER, which keeps the rollout per-job and revertable without touching kops.Where it runs
The lookup needs instances to exist, so it runs in
Up()after the cloud resources are created and before validation — a cluster that comes up but fails to validate is exactly when you want log collection to work.Up()is not the only entry point, soDumpClusterLogs()resolves the user too, for--downinvocations in a fresh process.Without
--dir,kops toolbox dumponly lists cloud resources; it does not SSH anywhere, so it does not need the credentials it is being used to determine.The environment handed to the tester is built during
init(), before any of this is knowable, so that export is extracted intoexportEnvForTester()and run again once the user is known.When nothing can be determined
An older kops that does not report
sshUser, or a cluster that failed before creating instances, leaves the user empty. In that case--ssh-useris omitted entirely rather than passed as"", so kops applies its own default (ubuntu) instead of being handed an unusable empty value. Bothtoolbox dumpcall sites do this.Testing
TestSSHUserFromDump: control plane preferred over workers, fallback to any instance, instances with no user skipped, and the two empty cases.TestResolveSSHUserFromClusterKeepsExistingUser: an already-set user survives untouched, with a deliberately bogusKopsBinaryPathso that any attempt to shell out would fail loudly rather than silently pass.go build ./...,go vet ./...andgo test ./...intests/e2e.Using SSH user: [prow]exactly as today. That is the assertion — nothing should change.Rollout
Opting jobs in is a separate test-infra change, and is gated on the kops version a job runs, because
gce/dump.goonly started reportingsshUserinc6e61af681(2026-07-12), which is not in release-1.34/1.35/1.36. Of 941 GCE periodics, the 361 on themastermarker are eligible; the other 580 become eligible as those branches age out of the matrix. The AWS jobs follow once #18698 (Rocky Linux) reaches the versions they run.Full plan:
docs/plans/gce-ssh-cleanup-step1.md.🤖 Generated with Claude Code