Background
On a cluster with slow etcd disks, cloudshell-controller-manager intermittently times out while acquiring its leader-election lease. After repeated timeouts the controller becomes stuck and stops reconciling any new CloudShell CRDs, so web terminals can no longer be opened.
When this happened, there was no way to diagnose where the controller was blocked: the controller manager exposes no pprof endpoint, so we could not capture a goroutine dump to see whether the hang was in leader election or elsewhere.
Proposal
Add an optional Go pprof profiling HTTP server to cloudshell-controller-manager.
- New flags:
--enable-pprof (bool, default false)
--profiling-bind-address (string, default :6060)
- Start the pprof server before leader election, so it is reachable even when the instance is not (or fails to become) the leader — that stuck / non-leader state is exactly what we need to profile.
- Register handlers on a dedicated
http.ServeMux (not the global DefaultServeMux), and shut the server down on context cancel.
- Wire it through the Helm chart with
pprof.enabled / pprof.bindAddress / pprof.port values (container args + containerPort).
With this in place, a wedged controller can be inspected, e.g.:
go tool pprof http://<pod>:6060/debug/pprof/goroutine
Scope
This is observability only. Fixing the underlying lease-timeout retry behavior, and adding health/liveness probes so a wedged controller auto-restarts, are worthwhile follow-ups but out of scope here.
Notes
cloudshell-controller-manager does not use controller-runtime's Manager, and the vendored controller-runtime v0.14.6 predates manager.Options.PprofBindAddress, so pprof needs to be wired manually rather than via the manager.
Background
On a cluster with slow etcd disks,
cloudshell-controller-managerintermittently times out while acquiring its leader-election lease. After repeated timeouts the controller becomes stuck and stops reconciling any newCloudShellCRDs, so web terminals can no longer be opened.When this happened, there was no way to diagnose where the controller was blocked: the controller manager exposes no pprof endpoint, so we could not capture a goroutine dump to see whether the hang was in leader election or elsewhere.
Proposal
Add an optional Go pprof profiling HTTP server to
cloudshell-controller-manager.--enable-pprof(bool, defaultfalse)--profiling-bind-address(string, default:6060)http.ServeMux(not the globalDefaultServeMux), and shut the server down on context cancel.pprof.enabled/pprof.bindAddress/pprof.portvalues (container args +containerPort).With this in place, a wedged controller can be inspected, e.g.:
Scope
This is observability only. Fixing the underlying lease-timeout retry behavior, and adding health/liveness probes so a wedged controller auto-restarts, are worthwhile follow-ups but out of scope here.
Notes
cloudshell-controller-managerdoes not use controller-runtime'sManager, and the vendoredcontroller-runtime v0.14.6predatesmanager.Options.PprofBindAddress, so pprof needs to be wired manually rather than via the manager.