gce: download nodeup from private GCS buckets with curl - #18623
Conversation
|
Skipping CI for Draft Pull Request. |
fa05780 to
e361409
Compare
e361409 to
63b4bcc
Compare
b7735c5 to
bf7d6fe
Compare
bf7d6fe to
8edd197
Compare
|
/test pull-kops-e2e-k8s-gce-ipalias |
8edd197 to
3cdd2c0
Compare
3cdd2c0 to
ccc958f
Compare
|
/test pull-kops-gce-master-scale-performance-100 |
|
/test pull-kops-e2e-aws-load-balancer-controller |
| case "gs": | ||
| // Only GCE instances can authenticate to GCS with their service account. | ||
| if cloudProvider != kops.CloudProviderGCE { | ||
| allErrs = append(allErrs, field.Invalid(fieldPath, s, "gs:// fileRepository is only supported on GCE")) |
There was a problem hiding this comment.
I think it would be nice to include the current cloud provider in the error message. If they think they are using GCE (likely with a gs:// protocol) then knowing what its actually set to may help debugging.
There was a problem hiding this comment.
Done. The message now includes the configured cloud provider.
| func buildVFSPath(target string) (string, error) { | ||
| if !strings.Contains(target, "://") || strings.HasPrefix(target, "memfs://") || strings.HasPrefix(target, "file://") { | ||
| if !strings.Contains(target, "://") || strings.HasPrefix(target, "memfs://") || strings.HasPrefix(target, "file://") || | ||
| strings.HasPrefix(target, "gs://") { |
There was a problem hiding this comment.
If we now exits early on "gs://" which is the expected form, then comment and code from lines 196-206 maybe should reflect they are dealing with the "http(s)://storage.googleapis.com/" edge case?
There was a problem hiding this comment.
Good point. Updated the comment to make it clear that this branch now only handles the https://storage.googleapis.com/ form, and the error hint mentions both forms.
| local token | ||
| echo "== Downloading ${url} ==" | ||
| # Pipe the service account token to curl, so that it stays out of the logs. | ||
| if ! token=$(curl -s -f --noproxy '*' --connect-timeout 2 --max-time 5 -H 'Metadata-Flavor: Google' 'http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token' | grep -o '"access_token" *: *"[^"]*"' | cut -d '"' -f 4); then |
There was a problem hiding this comment.
It seems like creating an entry for 169.254.169.254 to a hostname like LOCAL_METADATA would make this easier to understand. Would also allow us to change the IP to something like its IPv6 equivalent more easily.
There was a problem hiding this comment.
Done, factored into a metadata_server variable. Kept the IP rather than metadata.google.internal so that early boot does not depend on DNS, which matches what the Go metadata client does. GCE does have an IPv6 endpoint these days, http://[fd20:ce::254], though only for IPv6-only VMs, which kOps does not create on GCE yet, as nodes are always dual-stack.
| # Pipe the service account token to curl, so that it stays out of the logs. | ||
| if ! token=$(curl -s -f --noproxy '*' --connect-timeout 2 --max-time 5 -H 'Metadata-Flavor: Google' 'http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token' | grep -o '"access_token" *: *"[^"]*"' | cut -d '"' -f 4); then | ||
| echo "== Failed to get a service account token ==" | ||
| elif ! echo "Authorization: Bearer ${token}" | curl -f -Lo "${file}" --connect-timeout 20 --retry 6 --retry-delay 10 -H @- "https://storage.googleapis.com/${url#gs://}"; then |
There was a problem hiding this comment.
I'm used to using a few extra headers from https://docs.cloud.google.com/apis/docs/system-parameters but if this works for you, lets keep it simple.
There was a problem hiding this comment.
Tested the plain Authorization: Bearer header against a real private bucket and it works, so keeping it simple.
| continue | ||
| } | ||
| for _, location := range asset.Locations { | ||
| if strings.HasPrefix(location, "gs://") { |
There was a problem hiding this comment.
Is there any chance this could be in the unsanitized "http(s)://storage.googleapis.com/" form?
There was a problem hiding this comment.
It can, and that form takes the standard unauthenticated download path, same as before this PR: it works for public buckets and gets 403 for private ones. gs:// is the explicit opt-in for the authenticated download.
| if asset == nil { | ||
| continue | ||
| } | ||
| for _, location := range asset.Locations { |
There was a problem hiding this comment.
Is there any chance that we get a mix of gs:// and not gs:// ? Seems like this should be decided per location?
There was a problem hiding this comment.
In practice no, mirrors are only added for the default artifacts.k8s.io base. A custom base, which is the only way to get gs://, always yields a single location.
PR kubernetes#18466 added support for fetching nodeup from a private GCS bucket by calling "gcloud storage cp", which requires the gcloud CLI to be installed on the node image. Use curl with the instance service account token instead, so that no extra tooling is needed on the node. The source URLs are known when the bootstrap script is generated, so the download method is chosen there rather than by inspecting the URL at boot. For the default https sources, the script goes back to the standard download commands, identical to the other providers. On GCE, spec.assets.fileRepository can now be a gs:// URL, so that the node assets can be hosted in the same private bucket. Nodes elsewhere cannot authenticate to GCS, so validation keeps rejecting it there. "kops get assets --copy" can now also write to a gs:// repository. GCE e2e jobs and dev-build-gce.sh now use the gs:// form of the staged artifacts, so that the download is covered by CI and by dev clusters. Only kops is given this form; the tooling that downloads the kops binary keeps using https.
ccc958f to
be10c9e
Compare
|
/test pull-kops-verify-terraform |
|
/retest |
| // nodes download the staged artifacts with their instance service-account credentials. Any other | ||
| // URL, and any other cloud provider, passes through unchanged. Only kops invocations get the gs:// | ||
| // form; the scripts that download the kops binary run outside the deployer and keep using https. | ||
| func (d *deployer) maybeGSurl(/sitelet?url=https%3A%2F%2Fgithub.com%2Fkubernetes%2Fkops%2Fpull%2FbaseURL%2520string) string { |
There was a problem hiding this comment.
Nice to have. I don't find the outcome of a method called "maybeGSURL" intuitive. I would suggest something like "canonicalizeGSURL" or "standardizeGSURL" as better indicating what the method does.
There was a problem hiding this comment.
canonicalizeBaseURL seems like a good option, thanks for the suggestion.
I will update the name in the followup.
|
This seems cleaner. I like it. It would be nice if we can come up with a better name than "maybeGSURL". |
|
/approve |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: hakman The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
When KOPS_BASE_URL is an s3:// URL, the bootstrap script downloads nodeup with curl, signing the request with the instance profile credentials from IMDS using curl's native AWS SigV4 support. The credentials are piped in via --config -, so they never reach the disk, the process arguments, or the console log. An explicit x-amz-content-sha256 header lowers the required curl version from 8.0 to 7.86. The s3:// URL does not carry the bucket region, which SigV4 needs, so kops resolves it at script generation time through vfs, which already discovered and cached it when reading the nodeup hash from the bucket. spec.assets.fileRepository now accepts an s3:// URL on AWS, so node assets can live in the same private bucket. nodeup reads such assets through vfs, which signs requests with the instance credentials and resolves the bucket region. Also rename the kubetest2 deployer's maybeGSURL to canonicalizeBaseURL, as promised in the review of kubernetes#18623, and teach it to convert the public S3 https forms of staged artifacts to s3:// on AWS.
When KOPS_BASE_URL is an s3:// URL, the bootstrap script downloads nodeup with curl, signing the request with the instance profile credentials from IMDS using curl's native AWS SigV4 support. The credentials are piped in via --config -, so they never reach the disk, the process arguments, or the console log. An explicit x-amz-content-sha256 header lowers the required curl version from 8.0 to 7.86. The s3:// URL does not carry the bucket region, which SigV4 needs, so kops resolves it at script generation time through vfs, which already discovered and cached it when reading the nodeup hash from the bucket. spec.assets.fileRepository now accepts an s3:// URL on AWS, so node assets can live in the same private bucket. nodeup reads such assets through vfs, which signs requests with the instance credentials and resolves the bucket region. Also rename the kubetest2 deployer's maybeGSURL to canonicalizeBaseURL, as promised in the review of kubernetes#18623, and teach it to convert the public S3 https forms of staged artifacts to s3:// on AWS.
When KOPS_BASE_URL is an s3:// URL, the bootstrap script downloads nodeup with curl, signing the request with the instance profile credentials from IMDS using curl's native AWS SigV4 support. The credentials are piped in via --config -, so they never reach the disk, the process arguments, or the console log. An explicit x-amz-content-sha256 header lowers the required curl version from 8.0 to 7.86. The s3:// URL does not carry the bucket region, which SigV4 needs, so kops resolves it at script generation time through vfs, which already discovered and cached it when reading the nodeup hash from the bucket. spec.assets.fileRepository now accepts an s3:// URL on AWS, so node assets can live in the same private bucket. nodeup reads such assets through vfs, which signs requests with the instance credentials and resolves the bucket region. Also rename the kubetest2 deployer's maybeGSURL to canonicalizeBaseURL, as promised in the review of kubernetes#18623, and teach it to convert the public S3 https forms of staged artifacts to s3:// on AWS.
When KOPS_BASE_URL is an s3:// URL, the bootstrap script downloads nodeup with curl, signing the request with the instance profile credentials from IMDS using curl's native AWS SigV4 support. The credentials are piped in via --config -, so they never reach the disk, the process arguments, or the console log. An explicit x-amz-content-sha256 header lowers the required curl version from 8.0 to 7.86. The s3:// URL does not carry the bucket region, which SigV4 needs, so kops resolves it at script generation time through vfs, which already discovered and cached it when reading the nodeup hash from the bucket. spec.assets.fileRepository now accepts an s3:// URL on AWS, so node assets can live in the same private bucket. nodeup reads such assets through vfs, which signs requests with the instance credentials and resolves the bucket region. Also rename the kubetest2 deployer's maybeGSURL to canonicalizeBaseURL, as promised in the review of kubernetes#18623, and teach it to convert the public S3 https forms of staged artifacts to s3:// on AWS.
|
@hakman Looks like the curl version in some distros doesn't support |
|
@rifelpet it looks like bug, curl should not get any 'gs://' url. |
What this PR does / why we need it:
#18466 added support for fetching
nodeupfrom a private GCS bucket by prependinggcloud storage cpto the bootstrap script's download list, which requires the gcloud CLI on the node image. This uses plaincurlinstead, authenticated with the instance service account token from the metadata server. The token is piped in via-H @-, so it never reaches the disk, the process arguments, or the serial console.Since the download URLs are known when the script is generated, the method is chosen there rather than by testing the URL scheme at boot. Scripts using the default
httpssources are unchanged on every provider, and the GCE golden files return to their pre-#18466 form.Two follow-ons make it usable and tested:
spec.assets.fileRepositorynow accepts ags://URL on GCE, so node assets can live in the same private bucket (validation still rejects it elsewhere, as only GCE instances can authenticate to GCS), and GCE e2e jobs now pass kops thegs://form of the staged artifacts so the path is covered by CI./cc @cheftako @justinsb @rifelpet