Sitelet https://github.com/kubernetes/kops/pull/18686
Skip to content

tests/e2e: resolve SSH keys the same way on every cloud - #18686

Merged
kubernetes-prow[bot] merged 1 commit into
kubernetes:masterfrom
rifelpet:unified-ssh-key-resolution
Aug 14, 2026
Merged

kubernetes-prow[bot] merged 1 commit into
kubernetes:masterfrom
rifelpet:unified-ssh-key-resolution

Conversation

@rifelpet

@rifelpet rifelpet commented Aug 13, 2026 •

Copy link
Copy Markdown
Member

First step toward dropping the SSH key env vars and presets from the kops prow jobs entirely. This PR is a no-op for every job today — it unifies the code path and adds the missing fallbacks, but keeps honoring every variable the current presets set.

The problem

deployer/common.go resolved the keypair in four unrelated per-cloud branches:

Cloud Read Fallback if unset
AWS AWS_SSH_PRIVATE_KEY_FILE / AWS_SSH_PUBLIC_KEY_FILE generate
Azure (ignored flags and env entirely) always generate
DigitalOcean DO_SSH_PRIVATE_KEY_FILE / DO_SSH_PUBLIC_KEY_FILE none
GCE GCE_SSH_* (boskos) or KUBE_SSH_KEY_PATH (explicit project) gcloud compute config-ssh, or none

So two clouds had no fallback, Azure silently discarded --ssh-private-key, and GCE behaved differently depending on whether --gcp-project was supplied.

The change

One resolveSSHKeys() for all clouds, in precedence order:

  1. --ssh-private-key / --ssh-public-key
  2. KUBE_SSH_KEY_PATH / KUBE_SSH_PUBLIC_KEY_PATH (cloud-agnostic)
  3. the legacy per-cloud variables
  4. an ephemeral keypair generated for this cluster

Generating is safe everywhere because kops registers whatever public key it is given with the cloud provider itself — an EC2 key pair (awstasks.SSHKey), GCE ssh-keys instance metadata (gcemodel/autoscalinggroup.go), the Azure VMSS osProfile (azuretasks/vmscaleset.go), or a droplet key (dotasks/droplet.go). up.go already passes --ssh-public-key on every cloud.

env() now exports both KUBE_SSH_KEY_PATH and KUBE_SSH_USER for all providers, instead of exporting only the key on AWS, so ginkgo and clusterloader2 use the credentials the deployer resolved.

Both are required together, which the first revision of this PR got wrong and pull-kops-azure-dns-none caught. The framework's SkipUnlessSSHKeyPresent() skips its SSH tests when GetSigner() finds no key, so exporting the key alone un-skips [sig-node] SSH should SSH to all nodes on azure — which then failed with an empty username (error getting SSH client to @20.39.226.86:22), because azure and digitalocean assign the SSH user internally (kops / root) rather than from the job config, and the framework otherwise falls back to $USER. Note this cannot be fixed with an e2e flag: test/e2e/framework/ssh/ssh.go reads KUBE_SSH_USER from the environment and test_context.go has no --ssh-user flag.

That azure run also confirms the fix works: toolbox-dump.yaml reported sshUser: kops on all five instances, and the deployer had already SSHed into that same cluster as --ssh-user kops --private-key /tmp/kops/pr18686-kops-azure-dns-none/id_ed25519 to collect a 1.6 MB journal.log. The credentials were proven good; the framework simply was not told the user.

Two dead things removed

gce.SetupSSH ran gcloud compute config-ssh only when both GCE_SSH_* were empty. preset-k8s-ssh always sets them, so it never ran — config-ssh appears zero times in GCE job build logs. Its remaining effect was registering a key in GCE project metadata; the instance-metadata key kops installs makes that unnecessary.

lib/common.sh forced AWS_SSH_PRIVATE_KEY_FILE=${HOME}/.ssh/id_rsa when unset, which defeated the generation fallback by pointing at a file that doesn't exist in the prow pod. It had to go before generation could ever engage.

Why this is a no-op today

Every job still sets the legacy variables via its preset, and they resolve to the same paths as before:

  • preset-aws-ssh → /etc/aws-ssh/aws-ssh-{private,public}
  • preset-k8s-ssh → /etc/ssh-key-secret/ssh-{private,public}
  • preset-do-ssh → /etc/do-ssh/{private,public}-ssh-key
  • Azure sets none, so it still generates

Behavior changes only where there was previously no working path at all (DigitalOcean and GCE-with-project without their variables), or where a flag was being ignored (Azure).

Testing

  • New TestEnvExportsSSHKeyAndUser: asserts env() exports both variables on azure, digitalocean and aws.
  • New TestResolveSSHKeys: 11 cases covering flag precedence, cloud-agnostic vars winning over legacy ones, each preset's variables resolving unchanged, one cloud's variables not leaking into another, the .pub sibling lookup, and generation.
  • go build ./..., go vet ./..., and go test ./... in tests/e2e. Two pre-existing failures are unrelated and reproduce identically on master: kubetest2-kops/aws TestRandomZones (needs AWS credentials) and scenarios/bare-metal (needs a live cluster at 127.0.0.1:6443).
  • Checked the two scenarios that invoke kubetest2 with --test but no --up (aws-boskos/run-test.sh, lib/upgrade.sh): both pass --cluster-name explicitly, and CreateSSHKeyPair is deterministic per cluster name and skips regeneration when the key already exists, so repeated invocations in one pod share a key.

Follow-ups

  • Derive the SSH user from kops toolbox dump (instance.SSHUser), which is populated on all four clouds, so that it no longer has to come from the job config at all.
  • Then drop KUBE_SSH_KEY_PATH, KUBE_SSH_USER, preset-aws-ssh, preset-k8s-ssh, and preset-do-ssh from the job configs.

Note for the GCE part of that follow-up: GCE jobs currently SSH as prow, whose key comes from GCP project metadata, not from kops. An ephemeral key lands in instance metadata under the image-derived user (ubuntu on Ubuntu, admin elsewhere), so the GCE KUBE_SSH_USER has to change in the same step. I confirmed on u2404, cos125, deb12 and rocky10 job artifacts that the guest agent already creates that user and installs kops's key for it, and that the user has sudo.

🤖 Generated with Claude Code

@kubernetes-prow

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@kubernetes-prow kubernetes-prow Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. size/L Denotes a PR that changes 100-499 lines, ignoring generated files. labels Aug 13, 2026
@kubernetes-prow
kubernetes-prow Bot requested review from hakman and olemarkus August 13, 2026 22:03
@rifelpet

Copy link
Copy Markdown
Member Author

/test pull-kops-e2e-k8s-aws-amazonvpc
/test pull-kops-e2e-k8s-gce-cilium
/test pull-kops-azure-dns-none
/test pull-kops-do-dns-none

kubetest2-kops resolved its SSH keypair in four different per-cloud
branches: AWS read AWS_SSH_*, DigitalOcean read DO_SSH_*, GCE read
GCE_SSH_* or KUBE_SSH_KEY_PATH depending on whether a project was
supplied, and Azure ignored all of them and always generated. Only AWS
and Azure had a fallback, so DigitalOcean and GCE-with-a-project failed
outright when their variables were unset.

Replace all four with a single resolveSSHKeys(), which takes the
--ssh-private-key/--ssh-public-key flags first, then the cloud-agnostic
KUBE_SSH_KEY_PATH/KUBE_SSH_PUBLIC_KEY_PATH, then the legacy per-cloud
variables, and otherwise generates an ephemeral keypair for the cluster.
Generating works on every cloud because kops registers whatever public
key it is handed with the cloud provider itself -- an EC2 key pair, GCE
"ssh-keys" instance metadata, the Azure VMSS osProfile, or a DigitalOcean
droplet key.

The legacy variables are still honored, so the existing presets keep
working and this is a no-op for every job today. They can be removed from
the job configs once this has soaked.

Also export KUBE_SSH_KEY_PATH and KUBE_SSH_USER from env() on every cloud
rather than only exporting the key on AWS, so ginkgo and clusterloader2
reach the nodes with the same credentials the deployer resolved. Both are
needed together: the framework skips its SSH tests when GetSigner() finds
no key, so exporting the key alone un-skips them on azure and
digitalocean, where the user is assigned internally ("kops" and "root")
rather than by the job config and would otherwise fall back to $USER.

Two removals fall out of this:

  - gce.SetupSSH ran "gcloud compute config-ssh" only when both GCE_SSH_*
    were empty. preset-k8s-ssh always sets them, so it never ran in CI --
    "config-ssh" appears zero times in GCE job build logs. Its remaining
    effect was registering a key in GCE project metadata, which the
    instance-metadata key kops installs makes unnecessary.

  - lib/common.sh forced AWS_SSH_PRIVATE_KEY_FILE to ~/.ssh/id_rsa when
    unset, which defeated the generation fallback by pointing at a file
    that does not exist in the prow pod.
@rifelpet
rifelpet force-pushed the unified-ssh-key-resolution branch from 746fe7d to b611345 Compare August 14, 2026 00:54
@rifelpet

Copy link
Copy Markdown
Member Author

/test pull-kops-e2e-k8s-aws-amazonvpc
/test pull-kops-e2e-k8s-gce-cilium
/test pull-kops-azure-dns-none
/test pull-kops-do-dns-none

@rifelpet
rifelpet marked this pull request as ready for review August 14, 2026 02:09
@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 14, 2026
@kubernetes-prow
kubernetes-prow Bot requested a review from zetaab August 14, 2026 02:09
@hakman

hakman commented Aug 14, 2026

Copy link
Copy Markdown
Member

/override pull-kops-azure-dns-none
/lgtm
/approve

@kubernetes-prow

Copy link
Copy Markdown
Contributor

@hakman: Overrode contexts on behalf of hakman: pull-kops-azure-dns-none

Details

In response to this:

/override pull-kops-azure-dns-none
/lgtm
/approve

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@kubernetes-prow kubernetes-prow Bot added the lgtm "Looks good to me", indicates that a PR is ready to be merged. label Aug 14, 2026
@kubernetes-prow

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: hakman

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow kubernetes-prow Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 14, 2026
@kubernetes-prow
kubernetes-prow Bot merged commit 98e5697 into kubernetes:master Aug 14, 2026
29 checks passed
kubernetes-prow Bot pushed a commit to kubernetes/test-infra that referenced this pull request Aug 14, 2026
kubernetes/kops#18686 gave kubetest2-kops a single SSH key resolution
path that falls back to generating a throwaway keypair per cluster, and
kops registers that key with the cloud provider itself -- on AWS as an
EC2 key pair, which cloud-init installs for the AMI's default user. That
fallback has only ever run on azure, where no key was configured; every
AWS job still supplies the shared CI key via preset-aws-ssh.

Add an ephemeral_ssh_key option that omits preset-aws-ssh, and turn it on
for two periodics before switching the whole fleet:

  e2e-kops-aws-k8s-1-36      8/day, u2404,  ssh user "ubuntu"
  e2e-kops-aws-distro-al2023 3/day, al2023, ssh user "ec2-user"

Both already run "[sig-node] SSH should SSH to all nodes and run
commands" and pass it today, so a broken key or a key/user mismatch fails
a test rather than only degrading log collection. The two cover different
default users, and the sibling jobs in each family keep the shared key as
a control.

KUBE_SSH_USER is unchanged on both; only the key source moves.
alien1403 pushed a commit to alien1403/test-infra that referenced this pull request Aug 19, 2026
kubernetes/kops#18686 gave kubetest2-kops a single SSH key resolution
path that falls back to generating a throwaway keypair per cluster, and
kops registers that key with the cloud provider itself -- on AWS as an
EC2 key pair, which cloud-init installs for the AMI's default user. That
fallback has only ever run on azure, where no key was configured; every
AWS job still supplies the shared CI key via preset-aws-ssh.

Add an ephemeral_ssh_key option that omits preset-aws-ssh, and turn it on
for two periodics before switching the whole fleet:

  e2e-kops-aws-k8s-1-36      8/day, u2404,  ssh user "ubuntu"
  e2e-kops-aws-distro-al2023 3/day, al2023, ssh user "ec2-user"

Both already run "[sig-node] SSH should SSH to all nodes and run
commands" and pass it today, so a broken key or a key/user mismatch fails
a test rather than only degrading log collection. The two cover different
default users, and the sibling jobs in each family keep the shared key as
a control.

KUBE_SSH_USER is unchanged on both; only the key source moves.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. lgtm "Looks good to me", indicates that a PR is ready to be merged. size/L Denotes a PR that changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants