Sitelet https://github.com/googleapis/google-cloud-python/issues/15881
Skip to content

Race condition in MetricsTracer causes AttributeError: 'NoneType' object has no attribute 'status' #15881

Description

@rafi-rr

Environment details

  • OS type and version: Linux (Debian 12 on GKE)
  • Python version: 3.11.1
  • pip version: 24.0
  • google-cloud-spanner version: 3.61.0

Steps to reproduce

  1. Use google-cloud-spanner with SQLAlchemy (sqlalchemy-spanner 1.17.2) in a multi-threaded environment (e.g., FastAPI with uvicorn workers)
  2. Perform concurrent database operations (SELECT queries) under load
  3. The error occurs intermittently

Code example

The bug is a race condition in google/cloud/spanner_v1/metrics/metrics_interceptor.py:

def intercept(self, invoked_method, request_or_iterator, call_details):
    # ...
    SpannerMetricsTracerFactory.current_metrics_tracer.set_method(method_name)
    SpannerMetricsTracerFactory.current_metrics_tracer.record_attempt_start()
    response = invoked_method(request_or_iterator, call_details)
    SpannerMetricsTracerFactory.current_metrics_tracer.record_attempt_completion()  # BUG: current_metrics_tracer may have been replaced by another thread
    # ...

SpannerMetricsTracerFactory.current_metrics_tracer is a class variable shared across all threads. Between record_attempt_start() and record_attempt_completion(), another thread can call MetricsCapture.__enter__() which replaces current_metrics_tracer with a new instance:

# In metrics_capture.py __enter__():
SpannerMetricsTracerFactory.current_metrics_tracer = factory.create_metrics_tracer()

When record_attempt_completion() is called, it operates on a different MetricsTracer instance that hasn't had record_attempt_start() called, so current_op.current_attempt is None.

Suggested fix: Make current_metrics_tracer thread-local instead of a class variable:

import threading

class SpannerMetricsTracerFactory(MetricsTracerFactory):
    _metrics_tracer_factory: "SpannerMetricsTracerFactory" = None
    _thread_local = threading.local()
    
    @property
    def current_metrics_tracer(cls):
        return getattr(cls._thread_local, 'metrics_tracer', None)
    
    @current_metrics_tracer.setter
    def current_metrics_tracer(cls, value):
        cls._thread_local.metrics_tracer = value

Additional bug: There's a typo in spanner_metrics_tracer_factory.py line 93:

cls._metrics_tracer_factory.enabeld = enabled  # Should be "enabled" not "enabeld"

Stack trace

File "/usr/local/lib/python3.11/site-packages/google/cloud/spanner_v1/metrics/metrics_interceptor.py", line 148, in intercept
    SpannerMetricsTracerFactory.current_metrics_tracer.record_attempt_completion()

File "/usr/local/lib/python3.11/site-packages/google/cloud/spanner_v1/metrics/metrics_tracer.py", line 332, in record_attempt_completion
    self.current_op.current_attempt.status = status

AttributeError: 'NoneType' object has no attribute 'status'

Workaround: Disable built-in metrics by calling SpannerMetricsTracerFactory(enabled=False) before any Spanner operations.

Activity

  1. ericzundel commented on Jan 24, 2026

    @ericzundel

    TL;DR The problem is intermittent, but very frequent when you make ample use of threading. So much so as to make it unusable without turning off this value.

    Another workaround is to set the environment variable:

    SPANNER_DISABLE_BUILTIN_METRICS=true
    

    More details on how to reproduce

    I am seeing this error quite frequently when deployed to GCP with google-cloud-spanner==3.62.0 under load using uvicorn as the server with FastAPI and this pattern:

    From the FastAPI endpoint, my code like this:

    query_result = asyncio.to_thread(query_table, key=value)
    

    The stack trace shows the exception occurring when my code calls:

            my_orm = s.query(MyORM).first()
    

    from inside the spawned thread with the stack trace ending up the same as the original report.

    Unfortunately, when I say "under load" I mean we are running less than 10qps in this test (not a production service) and the frequency of the message is high, a large percentage of my calls fail, meaning (Python + Spanner + threading) is unusable without turning off client side metrics.

    Commit that exposed the problem

    commit 64aebe7e3ecfec756435f7d102b36f5a41f7cc52
    Author: Subham Sinha <35077434+sinhasubham@users.noreply.github.com>
    Date:   Thu Dec 4 11:27:30 2025 +0530
    

    Alternate workarounds

    Looks like you could downgrade the library to google_cloud_spanner==3.59.0 (I didn't try it)

    Looking through google/cloud/spanner_v1/client.py I saw that we can also set this environment variable to turn it off:

    SPANNER_DISABLE_BUILTIN_METRICS=true
    
  2. ericzundel commented on Feb 12, 2026

    @ericzundel

    @sinhasubham it looks like you are hard at work on this issue. Thanks for that.

  3. waiho-gumloop commented on Feb 24, 2026

    @waiho-gumloop
    Contributor

    We had a similar issue, as we're using both Uvicorn and Gunicorn workers, but the additional symptom was too many writes to the metrics endpoint in the same timeframe. Had to downgrade from 3.62.0 back to 3.55.0.

    Image
  4. chalmerlowe commented on Mar 2, 2026

    @chalmerlowe
    Contributor

    This issue was moved from the following repo as part of the monorepo migration: googleapis/python-spanner

  5. waiho-gumloop commented on Mar 25, 2026

    @waiho-gumloop
    Contributor

    For anyone arriving here with this error instead of the AttributeError race condition:

    INVALID_ARGUMENT: 400 One or more TimeSeries could not be written:
    timeSeries[...]: the set of resource labels is incomplete, missing (instance_id)
    

    That is a separate bug — a redundant MetricsCapture() inside trace_call that produces orphan metric data points with incomplete resource labels on every Spanner operation. The thread-safety fix in v3.63.0 (contextvars) resolved the AttributeError crash from this issue, but not the INVALID_ARGUMENT error.

    Tracked separately: #16173
    Fix: googleapis/python-spanner#1522

  6. jbbqqf commented on May 9, 2026

    @jbbqqf

    Hi! Triaging older issues this week — I think this one has shipped and can be closed.

    The race condition described here was caused by SpannerMetricsTracerFactory.current_metrics_tracer being a shared class variable. On current main, that storage has been replaced with a contextvars.ContextVar, which gives per-request isolation and removes the cross-thread interference described in this report:

    • packages/google-cloud-spanner/google/cloud/spanner_v1/metrics/spanner_metrics_tracer_factory.py:50 — _current_metrics_tracer_ctx = contextvars.ContextVar("current_metrics_tracer", default=None)
    • packages/google-cloud-spanner/google/cloud/spanner_v1/metrics/spanner_metrics_tracer_factory.py:95 — getter now reads from the ContextVar (_current_metrics_tracer_ctx.get())
    • packages/google-cloud-spanner/google/cloud/spanner_v1/metrics/spanner_metrics_tracer_factory.py:99 — setter writes into the ContextVar (_current_metrics_tracer_ctx.set(tracer))

    The enabeld typo at spanner_metrics_tracer_factory.py:93 mentioned in the original report is also no longer present.

    Would you mind closing this out? If you can still reproduce the AttributeError: 'NoneType' object has no attribute 'status' against a recent google-cloud-spanner release, please share the version and a minimal repro and I'll dig further.

    (Disclosure: I drafted this comment with help from Claude Code while triaging stale issues; the references above were verified manually against the current main branch.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

api: spannerIssues related to the Spanner API.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions