tests/e2e: resolve SSH keys the same way on every cloud - #18686
Conversation
|
Skipping CI for Draft Pull Request. |
|
/test pull-kops-e2e-k8s-aws-amazonvpc |
kubetest2-kops resolved its SSH keypair in four different per-cloud
branches: AWS read AWS_SSH_*, DigitalOcean read DO_SSH_*, GCE read
GCE_SSH_* or KUBE_SSH_KEY_PATH depending on whether a project was
supplied, and Azure ignored all of them and always generated. Only AWS
and Azure had a fallback, so DigitalOcean and GCE-with-a-project failed
outright when their variables were unset.
Replace all four with a single resolveSSHKeys(), which takes the
--ssh-private-key/--ssh-public-key flags first, then the cloud-agnostic
KUBE_SSH_KEY_PATH/KUBE_SSH_PUBLIC_KEY_PATH, then the legacy per-cloud
variables, and otherwise generates an ephemeral keypair for the cluster.
Generating works on every cloud because kops registers whatever public
key it is handed with the cloud provider itself -- an EC2 key pair, GCE
"ssh-keys" instance metadata, the Azure VMSS osProfile, or a DigitalOcean
droplet key.
The legacy variables are still honored, so the existing presets keep
working and this is a no-op for every job today. They can be removed from
the job configs once this has soaked.
Also export KUBE_SSH_KEY_PATH and KUBE_SSH_USER from env() on every cloud
rather than only exporting the key on AWS, so ginkgo and clusterloader2
reach the nodes with the same credentials the deployer resolved. Both are
needed together: the framework skips its SSH tests when GetSigner() finds
no key, so exporting the key alone un-skips them on azure and
digitalocean, where the user is assigned internally ("kops" and "root")
rather than by the job config and would otherwise fall back to $USER.
Two removals fall out of this:
- gce.SetupSSH ran "gcloud compute config-ssh" only when both GCE_SSH_*
were empty. preset-k8s-ssh always sets them, so it never ran in CI --
"config-ssh" appears zero times in GCE job build logs. Its remaining
effect was registering a key in GCE project metadata, which the
instance-metadata key kops installs makes unnecessary.
- lib/common.sh forced AWS_SSH_PRIVATE_KEY_FILE to ~/.ssh/id_rsa when
unset, which defeated the generation fallback by pointing at a file
that does not exist in the prow pod.
746fe7d to
b611345
Compare
|
/test pull-kops-e2e-k8s-aws-amazonvpc |
|
/override pull-kops-azure-dns-none |
|
@hakman: Overrode contexts on behalf of hakman: pull-kops-azure-dns-none DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: hakman The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
kubernetes/kops#18686 gave kubetest2-kops a single SSH key resolution path that falls back to generating a throwaway keypair per cluster, and kops registers that key with the cloud provider itself -- on AWS as an EC2 key pair, which cloud-init installs for the AMI's default user. That fallback has only ever run on azure, where no key was configured; every AWS job still supplies the shared CI key via preset-aws-ssh. Add an ephemeral_ssh_key option that omits preset-aws-ssh, and turn it on for two periodics before switching the whole fleet: e2e-kops-aws-k8s-1-36 8/day, u2404, ssh user "ubuntu" e2e-kops-aws-distro-al2023 3/day, al2023, ssh user "ec2-user" Both already run "[sig-node] SSH should SSH to all nodes and run commands" and pass it today, so a broken key or a key/user mismatch fails a test rather than only degrading log collection. The two cover different default users, and the sibling jobs in each family keep the shared key as a control. KUBE_SSH_USER is unchanged on both; only the key source moves.
kubernetes/kops#18686 gave kubetest2-kops a single SSH key resolution path that falls back to generating a throwaway keypair per cluster, and kops registers that key with the cloud provider itself -- on AWS as an EC2 key pair, which cloud-init installs for the AMI's default user. That fallback has only ever run on azure, where no key was configured; every AWS job still supplies the shared CI key via preset-aws-ssh. Add an ephemeral_ssh_key option that omits preset-aws-ssh, and turn it on for two periodics before switching the whole fleet: e2e-kops-aws-k8s-1-36 8/day, u2404, ssh user "ubuntu" e2e-kops-aws-distro-al2023 3/day, al2023, ssh user "ec2-user" Both already run "[sig-node] SSH should SSH to all nodes and run commands" and pass it today, so a broken key or a key/user mismatch fails a test rather than only degrading log collection. The two cover different default users, and the sibling jobs in each family keep the shared key as a control. KUBE_SSH_USER is unchanged on both; only the key source moves.
First step toward dropping the SSH key env vars and presets from the kops prow jobs entirely. This PR is a no-op for every job today — it unifies the code path and adds the missing fallbacks, but keeps honoring every variable the current presets set.
The problem
deployer/common.goresolved the keypair in four unrelated per-cloud branches:AWS_SSH_PRIVATE_KEY_FILE/AWS_SSH_PUBLIC_KEY_FILEDO_SSH_PRIVATE_KEY_FILE/DO_SSH_PUBLIC_KEY_FILEGCE_SSH_*(boskos) orKUBE_SSH_KEY_PATH(explicit project)gcloud compute config-ssh, or noneSo two clouds had no fallback, Azure silently discarded
--ssh-private-key, and GCE behaved differently depending on whether--gcp-projectwas supplied.The change
One
resolveSSHKeys()for all clouds, in precedence order:--ssh-private-key/--ssh-public-keyKUBE_SSH_KEY_PATH/KUBE_SSH_PUBLIC_KEY_PATH(cloud-agnostic)Generating is safe everywhere because kops registers whatever public key it is given with the cloud provider itself — an EC2 key pair (
awstasks.SSHKey), GCEssh-keysinstance metadata (gcemodel/autoscalinggroup.go), the Azure VMSSosProfile(azuretasks/vmscaleset.go), or a droplet key (dotasks/droplet.go).up.goalready passes--ssh-public-keyon every cloud.env()now exports bothKUBE_SSH_KEY_PATHandKUBE_SSH_USERfor all providers, instead of exporting only the key on AWS, so ginkgo and clusterloader2 use the credentials the deployer resolved.Both are required together, which the first revision of this PR got wrong and
pull-kops-azure-dns-nonecaught. The framework'sSkipUnlessSSHKeyPresent()skips its SSH tests whenGetSigner()finds no key, so exporting the key alone un-skips[sig-node] SSH should SSH to all nodeson azure — which then failed with an empty username (error getting SSH client to @20.39.226.86:22), because azure and digitalocean assign the SSH user internally (kops/root) rather than from the job config, and the framework otherwise falls back to$USER. Note this cannot be fixed with an e2e flag:test/e2e/framework/ssh/ssh.goreadsKUBE_SSH_USERfrom the environment andtest_context.gohas no--ssh-userflag.That azure run also confirms the fix works:
toolbox-dump.yamlreportedsshUser: kopson all five instances, and the deployer had already SSHed into that same cluster as--ssh-user kops --private-key /tmp/kops/pr18686-kops-azure-dns-none/id_ed25519to collect a 1.6 MBjournal.log. The credentials were proven good; the framework simply was not told the user.Two dead things removed
gce.SetupSSHrangcloud compute config-sshonly when bothGCE_SSH_*were empty.preset-k8s-sshalways sets them, so it never ran —config-sshappears zero times in GCE job build logs. Its remaining effect was registering a key in GCE project metadata; the instance-metadata key kops installs makes that unnecessary.lib/common.shforcedAWS_SSH_PRIVATE_KEY_FILE=${HOME}/.ssh/id_rsawhen unset, which defeated the generation fallback by pointing at a file that doesn't exist in the prow pod. It had to go before generation could ever engage.Why this is a no-op today
Every job still sets the legacy variables via its preset, and they resolve to the same paths as before:
preset-aws-ssh→/etc/aws-ssh/aws-ssh-{private,public}preset-k8s-ssh→/etc/ssh-key-secret/ssh-{private,public}preset-do-ssh→/etc/do-ssh/{private,public}-ssh-keyBehavior changes only where there was previously no working path at all (DigitalOcean and GCE-with-project without their variables), or where a flag was being ignored (Azure).
Testing
TestEnvExportsSSHKeyAndUser: assertsenv()exports both variables on azure, digitalocean and aws.TestResolveSSHKeys: 11 cases covering flag precedence, cloud-agnostic vars winning over legacy ones, each preset's variables resolving unchanged, one cloud's variables not leaking into another, the.pubsibling lookup, and generation.go build ./...,go vet ./..., andgo test ./...intests/e2e. Two pre-existing failures are unrelated and reproduce identically on master:kubetest2-kops/aws TestRandomZones(needs AWS credentials) andscenarios/bare-metal(needs a live cluster at 127.0.0.1:6443).--testbut no--up(aws-boskos/run-test.sh,lib/upgrade.sh): both pass--cluster-nameexplicitly, andCreateSSHKeyPairis deterministic per cluster name and skips regeneration when the key already exists, so repeated invocations in one pod share a key.Follow-ups
kops toolbox dump(instance.SSHUser), which is populated on all four clouds, so that it no longer has to come from the job config at all.KUBE_SSH_KEY_PATH,KUBE_SSH_USER,preset-aws-ssh,preset-k8s-ssh, andpreset-do-sshfrom the job configs.Note for the GCE part of that follow-up: GCE jobs currently SSH as
prow, whose key comes from GCP project metadata, not from kops. An ephemeral key lands in instance metadata under the image-derived user (ubuntuon Ubuntu,adminelsewhere), so the GCEKUBE_SSH_USERhas to change in the same step. I confirmed on u2404, cos125, deb12 and rocky10 job artifacts that the guest agent already creates that user and installs kops's key for it, and that the user has sudo.🤖 Generated with Claude Code