Sitelet https://github.com/keploy/samples-go/pull/245
Skip to content

fix: retry the Go module fetch in every Go sample's Dockerfile - #245

Open
slayerjain wants to merge 3 commits into
mainfrom
fix/retry-go-module-fetch-in-dockerfiles
Open

slayerjain wants to merge 3 commits into
mainfrom
fix/retry-go-module-fetch-in-dockerfiles

Conversation

@slayerjain

Copy link
Copy Markdown
Member

Problem

proxy.golang.org now and then resets an HTTP/2 stream partway through a module download:

go: golang.org/x/text@v0.31.0: read "https://proxy.golang.org/...": stream error: stream ID 187; INTERNAL_ERROR; received from peer

The next attempt succeeds. But every Go sample's Dockerfile here fetched its modules exactly once, using a bare go mod download, a bare go mod tidy, or a go build that fetches what it needs. One reset therefore fails the image build. It also fails the CI lane of whichever keploy repo was running docker build on the sample: keploy/keploy's workflows and other keploy CI clone this repo at main and build these Dockerfiles. This repo's own build.yml has the same exposure through its bare go build ./....

Changes

  1. fix: retry the Go module fetch in every Go sample's Dockerfile (27 Dockerfiles). Each fetch now runs in a loop: 5 attempts, 5-20s apart, then the build fails.

    RUN n=0; until go mod download; do n=$((n+1)); [ "$n" -lt 5 ] || exit 1; echo "go mod download failed (attempt $n), retrying in $((n*5))s"; sleep $((n*5)); done
    
    • The 19 bare go mod downloads and go-grpc's two go mod tidys run in the loop.
    • echo-sql, gin-redis, mux-mysql, sse-svelte and users-profile fetched inside go build. They now run the looped go mod download first.
    • nethttp-mysql's go mod download || true ignored a failed fetch and left it to an un-retried go mod tidy. Both are looped now.

    A module fetch is idempotent and fails only because of the network, so retrying it is safe. A missing module or a bad go.mod fails every attempt and still fails the build.

  2. fix(gin-redis, gin-pulsar): build them on a Go their go.mod accepts. Neither image built before this PR. gin-redis's go.mod needs Go 1.23 but the image was golang:1.20; gin-pulsar's needs Go 1.25 but the image was golang:1.23.

  3. ci: retry each sample's module fetch before building it. build.yml now runs go list -deps ./... in a loop, the same 5 attempts, before each go build ./.... go list -deps loads the same packages go build loads, so it fetches nothing more.

Verification

  • All 27 changed Dockerfiles were built on a test machine: 25 built. gin-pulsar and gin-redis failed for the same reason the unchanged Dockerfiles fail, and both build after commit 2. docker build --check reports no warnings on any of them.
  • A GOPROXY was set up that cuts the first download of one module partway:
    • The unchanged go-docker-timefreeze and go-grpc server Dockerfiles fail with unexpected EOF (on golang-jwt/jwt v5.3.0 and on grpc v1.69.2). The looped ones recover on the second attempt.
    • The same holds for build.yml's step: the old step fails on go-docker-timefreeze, and the new one recovers.
  • build.yml's step was run as written, with Go 1.27.0 and an empty module cache: all 44 modules build, both with and without this change.
  • keploy/keploy's and keploy's other CI were searched for sed edits to these Dockerfiles. The only one rewrites go-docker-timefreeze's RUN CGO_ENABLED=0 GOOS=linux go build -o /main . line, and this PR leaves that line unchanged.

proxy.golang.org now and then resets an HTTP/2 stream partway through a
download:

  go: golang.org/x/text@v0.31.0: read "https://proxy.golang.org/...":
      stream error: stream ID 187; INTERNAL_ERROR; received from peer

The next attempt succeeds. But each Go sample's Dockerfile fetched its
modules exactly once, with a bare `go mod download`, a bare `go mod
tidy`, or a `go build` that fetches what it needs. So one reset failed
the image build, and with it the CI lane of whichever repo was building
the sample: the CI of keploy/keploy and of other keploy repositories
clones this repo and runs `docker build` on these Dockerfiles.

Each fetch now runs in a loop: 5 attempts, 5-20s apart, and then the
build fails. A module fetch is idempotent and fails because of the
network, so retrying it is safe. A missing module or a bad go.mod fails
every attempt and still fails the build, about 50s later.

- The 19 bare `go mod download`s and go-grpc's two `go mod tidy`s run
  in the loop.
- echo-sql, gin-redis, mux-mysql, sse-svelte and users-profile fetched
  inside `go build`. They now run the looped `go mod download` first.
- nethttp-mysql's `go mod download || true` ignored a failed fetch and
  left it to an un-retried `go mod tidy`. Both run in the loop now, so
  a fetch that fails five times fails the build.

Built all 27 Dockerfiles: 25 build, and gin-pulsar and gin-redis fail
for the same reason as before this change: their image's Go is older
than their go.mod requires, which a separate commit fixes. With a GOPROXY that
cuts the first download of one module partway, the original
go-docker-timefreeze and go-grpc server builds fail ("unexpected EOF"
on golang-jwt/jwt v5.3.0 and on grpc v1.69.2), and the looped ones
recover on the second attempt.

Signed-off-by: slayerjain <shubhamkjain@outlook.com>
Neither image built. gin-redis's go.mod says `go 1.23.0` and `toolchain
go1.23.1`, which the Go in golang:1.20 cannot parse:

  go: errors parsing go.mod:
  /app/go.mod:3: invalid go version '1.23.0': must match format 1.23
  /app/go.mod:5: unknown directive: toolchain

gin-pulsar's go.mod says `go 1.25.0`, which golang:1.23 refuses:

  go: go.mod requires go >= 1.25.0 (running go 1.23.12; GOTOOLCHAIN=local)

Each now builds on the Go minor release its go.mod requires,
golang:1.23 and golang:1.25, and both images build. gin-pulsar's README
asked for Go 1.23+ as well; it now asks for 1.25+.

Signed-off-by: slayerjain <shubhamkjain@outlook.com>
build.yml runs `go build ./...` in every top-level sample with a
go.mod, and that build fetched the sample's modules from
proxy.golang.org with no retry. One reset stream ("stream error: stream
ID N; INTERNAL_ERROR; received from peer") would fail the whole job.

Each sample now runs `go list -deps ./...` first, in a loop: 5
attempts, 5-20s apart, and then the job fails. `go list -deps` loads
the packages `go build` loads, so it fetches the same modules and no
more. `go mod download` would also fetch every module go.mod requires,
whether the build needs it or not.

Run as build.yml runs it, with Go 1.27.0 and an empty module cache, all
44 modules build, both with and without this change. With a GOPROXY
that cuts the first download of golang-jwt/jwt partway, the old step
fails on go-docker-timefreeze ("unexpected EOF"), and the new one
recovers on the second attempt.

Signed-off-by: slayerjain <shubhamkjain@outlook.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant