Sitelet https://github.com/googleapis/google-cloud-node/issues/9501
Skip to content

firestore: a stale Watch idle timer re-enters initStream and closes the listener with "A backoff operation is already in progress" #9501

Description

@jayenashar

A long-lived onSnapshot listener dies with Error: A backoff operation is already in progress. The error arrives at the user's error callback, so the listener is gone and the application has to resubscribe. On 2026-10-01 four listeners in one process died within 100ms of each other, on a Cloud Run instance that had just been woken by a request.

Environment: @google-cloud/firestore 7.11.6 via firebase-admin 13.10, Node 20, Cloud Functions 2nd gen on Cloud Run with cpu: 0.125, no always-on CPU, maxInstances: 1, and one request every five minutes. The stack below is from that version's build/src, but the code it points at is unchanged on main, so the line numbers in my reading are from handwritten/firestore/dev/src as it stands today.

The stack:

Error: A backoff operation is already in progress.
    at ExponentialBackoff.backoffAndWait (/workspace/node_modules/@google-cloud/firestore/build/src/backoff.js:182:35)
    at QueryWatch.initStream (/workspace/node_modules/@google-cloud/firestore/build/src/watch.js:290:14)
    at QueryWatch.maybeReopenStream (/workspace/node_modules/@google-cloud/firestore/build/src/watch.js:244:18)
    at Timeout.<anonymous> (/workspace/node_modules/@google-cloud/firestore/build/src/watch.js:267:18)
    at listOnTimeout (node:internal/timers:581:17)
    at processTimers (node:internal/timers:519:7)

What I think happens. resetIdleTimeout arms idleTimeoutHandle when a stream opens (watch.ts:428), and only two places clear it: resetIdleTimeout itself on the next data event (watch.ts:424), and shutdown (watch.ts:873). The stream's error and end handlers call maybeReopenStream without clearing it (watch.ts:517 and watch.ts:526), so the idle timer armed by a dead stream stays live. When it fires it ends a stream that has already gone and calls initStream a second time (watch.ts:439), while the backoff from the first call is still pending. backoffAndWait rejects on awaitingBackoffCompletion (backoff.ts:256), initStream's own .catch treats that as fatal and calls closeStream (watch.ts:532), and the listener is handed to the caller as a permanent failure. awaitingBackoffCompletion is cleared only by its own delay timer (backoff.ts:290), so any second initStream during a backoff reaches the same place. The stale idle timer is just the easiest way to get there.

A CPU-throttled instance makes it likely rather than rare. Between requests the instance gets no CPU, so the 120s idle timer and the backoff's own delay both come due during the freeze and then run in the same wake. That also explains four listeners failing at once: each Watch has its own backoff and its own idle timer, and they all came due together.

Is clearing idleTimeoutHandle in maybeReopenStream, or at the top of initStream, the right place to fix it?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    api: firestoreIssues related to the Firestore API.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions