Get in touch

9io.ai / Blog

How we pinned down four intermittent bugs in React 19 and Python

Four bugs in a React 19 and Python stack that failed only some of the time, the condition behind each, and how we made each one show up on demand.

Key takeaways

  • An ordering error equal to a UTC offset usually means naive local timestamps.
  • Run test suites with TZ set to a non-UTC zone to expose naive datetimes.
  • To make a timing race fail reliably, slow down the side that usually wins.
  • Importing a module should send no network requests, because tests import it once per file.
  • Six clean runs barely test a 1-in-6 flake; ruling it out takes about 17.

Four bugs on an AI study platform we build and run showed up only some of the time. The front end is React 19 and the back end is Python. Chat messages sorted out of order, but only on a laptop set to India Standard Time. A React test failed only under load. A network probe that ran at import time leaked into the fetch mocks of 31 test files, and each getByRole query in jsdom cost 100 to 250 ms.

Each bug depended on a condition that changed between machines or between runs. In every case the work was to find that condition and force it, until the failure happened reliably or could be measured. The causes turned out to be naive timestamps, React effects racing a fake clock, an import-time side effect and jsdom discarding its style cache on every DOM change.

Below, each bug gets a section with the number that gave it away and what to do about it. After that come the measurement details and a table of how many clean runs it takes to call a flaky test fixed.

The four bugs at a glance

Bug What triggered it What we measured
Naive chat timestamps A machine clock not set to UTC Questions sorted 5 h 30 min after their own replies, on a laptop set to IST
React 19 effect timing CPU load during the test run Reliable failure once zero-delay timers were stalled by 2 ms
Network probe at import A request sent whenever a module loaded Probe calls in 31 test files’ fetch mocks; 1 failing run in 6, then 0 in 6 without it
Slow getByRole in jsdom Style lookups repeated after every DOM change 100 to 250 ms per query; 0.9 s of queries in one test with a 5 s timeout

Questions sorted five and a half hours after their answers

The chat UI on the platform is OpenAI’s ChatKit. Its Python server stamps each new user message with created_at=datetime.now(), as its API reference shows, and its agents helper does the same for assistant messages. With no argument, datetime.now() returns local time with no timezone attached, which Python calls a naive datetime. The docs leave it to the program to decide whether a naive value means UTC or local time.

A machine set to UTC hides the problem, because local time and UTC are then the same reading. IST is UTC+5:30, so on a laptop set to IST every naive stamp is five and a half hours ahead of UTC. A question typed at 15:30 on that laptop is stamped with the bare value 15:30. A reply stamped a moment later in UTC reads 10:00. Read the bare 15:30 as UTC and the question sits five and a half hours after its own answer, which is the bug we saw.

Python objects to the mix-up in only one case. Ordering a naive and an aware datetime with < raises TypeError. Comparing them with == just returns False, and once the values are serialised to strings or stored as plain numbers, nothing checks at all. JavaScript has the same trap in another form. A date-time string with no offset is parsed as local time, while a date-only string is parsed as UTC.

The fix was timezone-aware UTC timestamps end to end. In Python that comes down to creating timestamps with datetime.now(timezone.utc), converting naive values once where you know what they mean, and refusing naive values at the storage boundary.

from datetime import datetime, timezone
from pydantic import AwareDatetime, BaseModel

def utc_now() -> datetime:
    return datetime.now(timezone.utc)

def local_to_utc(value: datetime) -> datetime:
    # Only for naive values known to be local time, such as datetime.now().
    # astimezone() treats a naive datetime as system local time.
    return value.astimezone(timezone.utc)

class StoredMessage(BaseModel):
    text: str
    created_at: AwareDatetime  # a naive datetime fails validation

datetime.utcnow() doesn’t help here. It returns a naive value and has been deprecated since Python 3.12. astimezone() on a naive value assumes system local time. That is right for values made by datetime.now() on the same machine and wrong for anything else, so keep the conversion in the one place where you know the source. Pydantic’s AwareDatetime rejects naive input with a timezone_aware error. With ChatKit, the natural boundary is your store, whose add_thread_item and save_item methods see every item before it is written.

To make this class of bug show up on any machine, run the suites under a non-UTC zone as well as UTC. Python takes its zone from TZ on Unix and Node.js does the same, so one variable covers both halves of the stack.

TZ=Asia/Kolkata pytest
TZ=Asia/Kolkata npm test

React 19 effects ran after the test’s timer drain

When an update doesn’t come from an interaction such as a click, React generally lets the browser paint before it runs useEffect callbacks, according to the useEffect reference. In a test that means a later task. React’s scheduler posts it with setImmediate when the environment has one, then MessageChannel, then setTimeout(fn, 0), and it keeps the references it finds when it first loads. React 19.1 also changed how passive effects are scheduled, and the pull request notes that effects can now run earlier or later than before. The moment an effect runs can shift with the React version as well as with load.

React Testing Library’s async helpers end with a short drain. After a waitFor callback passes, the wrapper waits for a zero-delay timer so that in-flight promises settle. When it detects fake timers, it advances the fake clock by 0 ms instead of waiting for a real timer. Under fake timers, then, the drain finishes without the real event loop reaching its next task.

That made a race between two clocks. The effect waited for a real task from React’s scheduler, while the drain ran on the fake clock. On a quiet machine the effect usually ran first. Under load it ran late, and the test made its assertion before the effect had run.

We made it fail reliably by stalling zero-delay timers by 2 ms. A setup file that loads before React can do it.

// Test setup, loaded before React. Delays zero-delay timers by 2 ms to mimic
// a loaded machine. Remove it once the race is fixed.
const realSetTimeout = globalThis.setTimeout;
globalThis.setTimeout = (fn, ms = 0, ...args) =>
  realSetTimeout(fn, ms > 0 ? ms : 2, ...args);

The stall reaches React only when its scheduler is on the setTimeout branch, which happens when the environment has neither setImmediate nor MessageChannel. Log typeof setImmediate and typeof MessageChannel in a test to see which branch yours takes, and delay that mechanism if it differs. The same move works for many races. Slow down, on purpose, the side that usually wins. Once the failure is reliable, each candidate fix takes one run to judge.

For this pattern, these are the fixes worth trying first:

  • Run every step that changes React state inside await act(...), fake-timer advances included. React applies the updates and effects queued inside act before it returns (act reference), and Kent C. Dodds covers the fake-timer case.
  • With user-event, pass the fake clock in through the advanceTimers option.
  • Drop fake timers from tests that don’t need them, and wait with findBy queries in real time.

A network probe that ran on import

One module sent a network probe as soon as it was imported. Jest gives each test file its own module registry by default, and Vitest isolates test files by default too. So the probe went out again in every test file that imported the module, directly or through another import. Its calls ended up in the fetch mocks of 31 test files.

A stray call can break a mocked fetch in more than one way. It adds to call counts that a test asserts on. It can also take a response the test queued for its own request with a one-shot mock such as mockResolvedValueOnce. Exactly when the probe’s request lands, relative to the test’s own requests, can change from run to run. That is how a side effect like this fails 1 run in 6 instead of every run.

Removing the side effect, so that importing the module sends nothing, took the suite from 1 failing run in 6 to 0 in 6. The general rule is that importing a module should not do I/O. Export a function and call it from the app’s entry point instead.

To find other offenders, make the default fetch mock in your test setup record a stack trace for any call that arrives before the first test of a file starts, and fail the file if it recorded one. The stack trace names the module that made the call.

Why each getByRole query cost 100 to 250 ms in jsdom

The last bug was a slow test with a 5 s timeout. By default getByRole ignores elements that are hidden from the accessibility tree. To check, Testing Library calls getComputedStyle on each candidate to look for visibility: hidden, then walks up its ancestors calling it again to look for display: none (role-helpers.js). It caches those answers only for the length of one query (role.ts). Matching by name adds an accessible-name computation, which reads more styles.

jsdom has cached computed styles since version 21.1.0. The cache belongs to the document, and the pull request that added it says it is invalidated completely on every DOM change. In a test that renders, types and waits, the DOM changes between most queries, so most queries start with a cold cache.

Each getByRole query cost 100 to 250 ms. One test spent 0.9 s in queries alone against its 5 s timeout, which is also the default in Jest and Vitest. Everything else the test did had to fit in the time left, and on a busy machine it sometimes didn’t.

Testing Library’s ByRole docs suggest the first two options below.

  • Pass hidden: true to skip the visibility walk. The query will then also match elements hidden with CSS, so assert visibility separately where it matters.
  • Use getByLabelText or getByText where the role adds nothing to the test.
  • Search a smaller subtree with within(), so fewer candidates need checking.
  • Upgrade jsdom. Releases 29.0.2, 29.1.1 and 30.1.0 all list getComputedStyle() speed-ups in their release notes.
  • Raise the timeout last, since a bigger budget hides the cost without removing it.

Time the queries before and after any of these changes. Two lines around the call are enough.

const t0 = performance.now();
screen.getByRole('button', { name: 'Send' });
console.log(`getByRole: ${Math.round(performance.now() - t0)} ms`);

How the numbers were measured

The timezone offset was observed on one laptop set to IST. For a given zone that bug happens every time, so there is no failure rate to report. The effect-timing failure became reliable once zero-delay timers were stalled by 2 ms, and this post has no failure rate for the original intermittent version.

The probe figures come from six runs of the suite before the change and six after. Six runs is a small sample. If a 1-in-6 failure rate had survived the change, six clean runs in a row would still happen about a third of the time, since (5/6)^6 ≈ 0.33. The evidence that the fix worked is that the mechanism is gone, and 0 in 6 is consistent with that.

The query timings come from one slow test on one machine. Recent jsdom releases list getComputedStyle() speed-ups, so measure on the version you run before reusing our numbers. We have no figures here for CI against local runs, or for how loaded the machine was.

How many clean runs it takes to call a flake fixed

To rule out a failure rate of 1 in N with 95% confidence, you need n clean runs in a row such that (1 - 1/N)^n is below 0.05. That works out to roughly 3N runs, which is the rule of three from statistics.

Failure rate before the fix Clean runs in a row for 95% confidence
1 in 2 5
1 in 6 17
1 in 10 29
1 in 20 59
1 in 50 149
1 in 100 299

For a 1-in-100 flake that is 299 runs. Forcing the failure is cheaper. When it fails on demand, one failing run before the fix and one passing run after show that the change reached the cause.

A checklist for the next intermittent failure

  • Compare any odd offset or delay with known values, such as UTC offsets or default intervals. waitFor polls every 50 ms by default.
  • Run the back-end and front-end suites under a non-UTC TZ in CI, in addition to UTC.
  • Store timestamps as timezone-aware UTC, and make the storage layer reject naive values.
  • When a failure depends on timing, delay the side that usually wins until it fails reliably.
  • Keep imports free of I/O, and keep a tripwire in the test setup for fetches made during import.
  • Time slow queries before raising timeouts, and use hidden: true or cheaper queries where they fit.
  • Before calling a flake fixed, remove a cause you can explain or collect about 3N clean runs.

Frequently asked questions

Why do chat messages sort out of order on some machines only?

Often because the timestamps are naive, meaning local time with no zone attached. On a machine set to UTC that equals UTC and looks right. Anywhere else it is off by the local UTC offset.

How do I get a timezone-aware UTC datetime in Python?

Call datetime.now(timezone.utc). datetime.utcnow() returns a naive value and has been deprecated since Python 3.12.

Why does a React test with fake timers fail only some of the time?

React runs many effects in a later task. If the test moves a fake clock and asserts before that real task runs, it fails, and machine load decides which happens first.

Why is getByRole slow in jsdom?

By default it calls getComputedStyle on each candidate element and its ancestors, and jsdom discards its computed-style cache whenever the DOM changes. Passing hidden: true skips that check.

How many passing runs show that a flaky test is fixed?

To rule out a 1-in-N failure rate with 95% confidence you need about 3N clean runs in a row, for example 17 for 1 in 6 and 59 for 1 in 20.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.