Get in touch

9io.ai / Blog

Is Claude nerfed? How to measure model drift yourself

Is Claude nerfed, or did something else change? Documented degradations, the changes that only look like one, and a nightly drift check with error bars.

Key takeaways

  • Anthropic traced degraded Claude responses in August and September 2025 to three infrastructure bugs.
  • Opus 5.5 defaults to medium effort where Opus 5 defaulted to high, so set effort explicitly.
  • Safety-classifier fallbacks can answer with an older model, so log the served model on every response.
  • Compare per-task scores with a baseline using paired differences, error bars and a control model.
  • In Livenerf’s validation, low effort cut output tokens by 62% while the accuracy drop wasn’t significant.

When someone asks “is Claude nerfed?”, a handful of bad sessions can’t answer it. Real regressions happen, and Anthropic traced degraded responses in August and September 2025 to three bugs in its serving infrastructure. A model can also seem worse because something around it changed, such as a lower default effort level, a fallback to another model, a client update or a long, cluttered context. The reliable test is to run fixed tasks against a pinned model ID every night, repeat each one, and compare per-task scores with a baseline, using error bars and a control model.

Claude Opus 5.5 launched on 22 September 2026 and Sonnet 5.5 on 28 September,1 and a Hacker News thread on 29 September asked whether Opus 5.5 had been nerfed yet.2 We haven’t measured Opus 5.5 ourselves, and this post doesn’t claim that any model is or isn’t degraded.

Below are the degradations vendors have documented, the changes that only look like one, a measurement method with a nightly script, and the public Livenerf project. The method works for any provider’s model, and everything is as of 3 October 2026.

Is Claude nerfed? What Anthropic’s 2025 postmortem found

The best-documented case of Claude getting worse is Anthropic’s own postmortem of 17 September 2025, which describes three bugs that had degraded some responses since early August.3

In the first, from 5 August, some Claude Sonnet 4 requests went to servers configured for the upcoming 1M-token context window, at first about 0.8% of them. A routine load-balancing change on 29 August made it worse, and in the worst hour, on 31 August, 16% of Sonnet 4 requests were affected. Anthropic estimates that about 30% of Claude Code users who made requests in that period had at least one message misrouted. Routing was sticky, so a user who landed on a bad server tended to stay there.

The second was a TPU server misconfiguration on the Claude API from 25 August. It sometimes produced tokens that should rarely appear, such as Thai or Chinese characters in English answers, on Opus 4.1, Opus 4 and Sonnet 4. The third, a latent compiler bug, made an approximate top-k operation sometimes return wrong results for certain batch sizes and model configurations, and was confirmed on Claude Haiku 3.5.

The account of why detection took weeks is the useful part for anyone building a monitor. Anthropic’s evaluations didn’t capture what users reported, partly because Claude often recovers from isolated mistakes, and privacy controls limit when engineers can read user conversations. Each bug produced different symptoms on different platforms at different rates, so the reports looked like random degradation. Its fixes include running quality evaluations continuously on production systems.

The postmortem also states the company’s position on deliberate downgrades: “We never reduce model quality due to demand, time of day, or server load.” It attributes the problems users reported to infrastructure bugs alone.

Is GPT getting worse? What OpenAI and Google have documented

Other vendors have documented changes too, mostly behind a model name that stayed the same.

When Product What changed Source
25 January 2024 OpenAI API A new GPT-4 Turbo snapshot aimed to reduce “laziness”, meaning unfinished tasks; callers on the unpinned gpt-3.5-turbo alias were to move to a new snapshot automatically OpenAI4
25 to 28 April 2025 ChatGPT, GPT-4o An unannounced update made answers sycophantic, and OpenAI began rolling it back on 28 April OpenAI56
6 May 2025 Gemini API The dated preview ID gemini-2.5-pro-preview-03-25 began pointing to the newer 05-06 model Google7
7 August 2025 ChatGPT, GPT-5 The router that picks a model was down for part of launch day, and Sam Altman said GPT-5 seemed much dumber as a result TechCrunch8

OpenAI’s write-up of the GPT-4o episode is worth reading for anyone who relies on a vendor’s own testing. Its offline evaluations and A/B tests looked good before launch, its first fix was a change to ChatGPT’s system prompt, and the full rollback took about 24 hours.6

Pin the most specific model ID a vendor offers and log the model each response reports, because aliases move, and the Gemini case shows that even a dated preview ID can be repointed. A consumer app can also change the model, the system prompt and the routing behind one product name.

What the 2023 ChatGPT drift study found, and what its critics showed

A study often cited when people ask whether GPT is getting worse is Chen, Zaharia and Zou’s “How is ChatGPT’s behavior changing over time?”, first posted on 18 July 2023.9 It compared the March and June 2023 snapshots of GPT-4 and GPT-3.5 in OpenAI’s API on tasks including maths, sensitive questions and code generation.

Its first version reported that GPT-4 identified primes with 97.6% accuracy in March and 2.4% in June. The next day, Arvind Narayanan and Sayash Kapoor pointed out that the test set held only primes.10 A model whose default guess flips from prime to composite goes from near-perfect to near-zero on such a set with no change in ability. On 500 composite numbers, they found all four model versions equally poor. They also noted that the code task only checked whether answers ran as returned, so text added around the code counted as a failure. They concluded that the paper showed a change in behaviour and did not show degraded capability.

The revised paper of 31 October 2023 used 500 primes and 500 composites. GPT-4’s accuracy fell from 84.0% to 51.1%, and the June version labelled 99.7% of numbers composite. On code, 52.0% of GPT-4’s March answers ran as returned against 10.0% in June, because the June model wrapped its code in Markdown fences. With that text stripped, GPT-4’s pass rate rose from 52.0% to 70.0%. The authors also measured a large drop in instruction-following, and argued that drift like this can break pipelines that depend on exact output.

Changes that look like a worse model

When someone asks why Claude is worse now, these are the changes we’d rule out first, and the same list applies to any other model. Most of them leave traces you can check.

Confounder What you notice How to check
Lower default effort Shorter answers, fewer tool calls, less care on hard problems Log effort and output tokens; set effort explicitly
Safety-classifier fallback An answer, or the rest of a session, comes from an older model Log the model each response reports; look for switch notices
Capacity or usage-limit fallback Weaker answers at busy times or late in a usage window Log the served model; check fallback settings and usage
Alias moved to a new snapshot Behaviour changes overnight with nothing changed on your side Pin the full model ID; alert when the served model changes
Client, SDK or CLI update The change starts on the day a tool updated Log versions with every result; pin them while comparing
Long or cluttered context Quality falls as a session grows and recovers in a fresh one Re-run in a fresh context; track input tokens
Provider incident Errors, timeouts or cut-off answers in a time window Match timestamps against the status feed
Sampling variance The same prompt passes, then fails Repeat each task; report error bars
Your own prompt, tool or grader change The drop starts at one of your deploys Version all three; re-run the old version

Effort defaults move between models. On the Claude API, Opus 5.5 runs at medium effort when a request doesn’t set one, where Opus 5 ran at high, and Anthropic’s migration guide tells developers to set it explicitly.11 In Livenerf’s validation on a 78-question panel, medium effort cut output tokens by 26% and low by 62% against high, while the accuracy differences, −4.2 and −8.3 points, were not significant at 95%.12

Safety classifiers can hand a request to another model. Anthropic’s refusals documentation says a flagged request may fall back to a less capable model or be declined, depending on the policy category.13 A refusal arrives as an HTTP 200, so monitoring built on error rates never sees it. With server-side fallback on, the response’s model field names the model that answered, and later turns of that conversation go straight to the fallback model for about an hour. Livenerf saw the classifier answer some of its questions with Opus 5, and rejects and counts those samples.14

Capacity can change the serving model too. Claude Code can switch to a configured fallback model for one turn when the primary is overloaded or unavailable.15 On the API, a refusal’s fallback is skipped when the fallback model is rate limited or overloaded, so fallbacks turn back into refusals under load.13 In ChatGPT, GPT-5.4 mini became the fallback in March 2026 for paid users who reach the GPT-5.4 Thinking rate limit.16

The same prompt gives different answers. Opus 5.5 rejects non-default temperature, top_p and top_k values, and Anthropic notes that temperature 0 never guaranteed identical outputs on earlier models.11 Thinking Machines got 80 different completions from 1,000 temperature-0 runs of one prompt on Qwen3-235B, because kernel results depended on batch size, which changes with server load.17 In Livenerf’s A/A check, two halves of identical runs differed by 6.4 points, with a standard error of 3.6.12 Our post on pinning down intermittent bugs works through how many clean runs it takes before a pass means something.

Long sessions behave differently from short ones. Chroma tested 18 models in July 2025 and found that performance grows increasingly unreliable as input length grows, even on simple tasks.18 Livenerf found that Claude Code loads a global instruction file by default, which took its calls from about 550 tokens of context to 11,200 until the project switched it off.19

Your own changes count too. Sclar and colleagues measured accuracy swings of up to 76 points on LLaMA-2-13B from prompt formatting alone, in few-shot settings.20 A reworded system prompt, a new tool description or an edited grader can move scores with nothing changed at the vendor.

Status pages give you timing. The Claude and OpenAI status pages publish machine-readable feeds, so a monitor can store open incidents with each result.21 Read the incident text, though. Several Claude incidents in August 2026 titled as degraded performance for a named model describe raised error rates, and their updates don’t mention answer quality.

Claude Code getting worse? Check the client before the model

Claude Code adds its own moving parts. npm lists ten Claude Code releases between 22 September and 3 October 2026, from 2.1.280 to 2.1.289.22 Several recent changes affect which model answers and how hard it thinks.23

  • 2.1.280 (22 September) added Opus 5.5 and made Opus the default on Pro and Team Standard plans, which had defaulted to Sonnet. It also stopped applying effort levels saved before per-model effort existed to new models such as Opus 5.5.
  • The model documentation, as of 2 October, says Opus 5.5 starts at medium effort unless you set a level, while most models start at high.15
  • When a safety classifier flags a request on Opus 5.5, Claude Code re-runs it on Opus 5 for biology flags or Opus 4.8 for cybersecurity flags, with a notice in the transcript. The session stays on the fallback model until you switch back with /model.
  • Until 2.1.287 (1 October), that switch also moved the session to the new model’s default effort.
  • 2.1.286 (30 September) made the fallback notice say when a fallback had cut the context window from 1M to 200K tokens.
  • 2.1.268 (10 September) fixed a bug that could silently switch a running session to the organisation’s default model.

None of these items changes the model behind the API. Before comparing two weeks of Claude Code, check claude --version, the effort level in /model and the transcript for model switches, and pin the CLI when you measure. Claude Code offers a stable release channel and a setting that turns off auto-updates.24

LLM drift monitoring with paired runs and error bars

Evan Miller’s 2024 paper from Anthropic, “Adding Error Bars to Evals”, sets out the statistics, including resampling answers within each question, paired differences between models and power analysis.25 A drift monitor built on it works like this.

  1. Freeze a task set from your own workload. Use tasks a program can grade, such as an exact answer or a test suite. Favour ones the model sometimes gets wrong, since a task it always passes can’t show a drop. Check the grader by hand on a sample. An LLM judge needs pinning as carefully as the model under test and brings biases of its own, which our post on two reviews and a judge covers.
  2. Pin everything you can, from the full model ID and effort level to the system prompt, tools, and SDK and CLI versions, and write it all into every result row.
  3. Repeat each task several times a night and score it by its mean.
  4. Compare paired differences. Build a baseline from the first week or two. Each night, subtract each task’s baseline mean from its new mean and average the differences. The standard error is their standard deviation divided by the square root of the number of tasks, and the 95% interval is 1.96 standard errors either side. Pairing removes the variance that comes from some tasks being harder than others.
  5. Run a control model through the same code. If it moves with the model under test, suspect your harness, your grader or the platform first.
  6. Log what was served. Record the model each response reports, the stop reason, refusals, fallbacks, output tokens and latency. Count refusals and answers from other models separately and keep them out of the score.
  7. Store open incidents from the provider’s status feed with each run.
  8. Set the decision rule before collecting data, such as an interval that excludes zero in two consecutive windows while the control stays flat. Livenerf requires 99% intervals over two consecutive 10-day windows, an effect of at least 3 points and no matching move in its control.14

Next, work out what your monitor can detect. At 95% confidence and 80% power, the smallest difference you can reliably see is about 2.8 standard errors (1.96 + 0.84). Miller’s worked example needs about 969 questions to detect a 3-point difference that way, under assumed variances. Livenerf’s design file, which uses 80% power and a 99% test, puts its own panel of 78 questions, asked once a day, at about 7.5 points per 10-day window. It estimates that seeing 3 points would take about 44 samples of each question a week.26

Watch output tokens alongside accuracy. In the Livenerf effort test above, the 95% interval for low effort’s accuracy change ran from −17.1 to +0.5 points, while the 99% interval for its token change ran from −70% to −51%.12 The project treats output tokens per sample as its main secondary signal.

A nightly drift check in Python

The script below, written for this post, covers steps 2 to 7 for the Claude API. It asks each task three times on the model under test and on a control, keeps refusals and answers from other models out of the score, and appends one row per model to a log. Only ask() and grade() depend on your vendor and your tasks.

import json, math, statistics, time, urllib.request
from collections import Counter
from pathlib import Path

import anthropic

ARMS = {  # the model under test and a control, run through the same code
    "primary": {"model": "claude-opus-5-5", "effort": "high"},
    "control": {"model": "claude-opus-5", "effort": "high"},
}
SAMPLES = 3  # answers per task per night
SYSTEM = "Answer the question. End with a line that contains only the final answer."
STATUS = "https://status.claude.com/api/v2/incidents/unresolved.json"
client = anthropic.Anthropic(max_retries=4)

def ask(arm, prompt):  # the only vendor-specific function
    r = client.messages.create(
        model=arm["model"], max_tokens=16000, system=SYSTEM,
        output_config={"effort": arm["effort"]},  # explicit, never the default
        messages=[{"role": "user", "content": prompt}],
    )
    text = "".join(b.text for b in r.content if b.type == "text")
    return text, r.model, r.stop_reason, r.usage.output_tokens

def grade(text, answer):  # exact match on the last line; use a check that fits your tasks
    lines = [line.strip() for line in text.splitlines() if line.strip()]
    return bool(lines) and lines[-1].lower() == answer.strip().lower()

def run(arm, tasks):
    scores, outcomes, tokens = {}, Counter(), []
    for task in tasks:
        marks = []
        for _ in range(SAMPLES):
            try:
                text, served, stop, out_tokens = ask(arm, task["prompt"])
            except anthropic.APIError as exc:
                outcomes[f"error:{type(exc).__name__}"] += 1
                continue
            if stop == "refusal" or served != arm["model"]:
                outcomes["refusal" if stop == "refusal" else f"served_by:{served}"] += 1
                continue  # counted, never scored
            outcomes[stop] += 1
            tokens.append(out_tokens)
            marks.append(1.0 if grade(text, task["answer"]) else 0.0)
        if marks:
            scores[str(task["id"])] = statistics.fmean(marks)
    return scores, outcomes, tokens

def paired_delta(tonight, baseline):
    ids = sorted(tonight.keys() & baseline.keys())
    diffs = [tonight[i] - baseline[i] for i in ids]
    se = statistics.stdev(diffs) / math.sqrt(len(diffs))
    return statistics.fmean(diffs), 1.96 * se, len(ids)

def build_baseline(days):  # run once, after the baseline nights
    out = {}
    for name in ARMS:
        nights = [json.loads(Path(f"scores/{d}-{name}.json").read_text()) for d in days]
        ids = set().union(*nights)
        out[name] = {i: statistics.fmean(n[i] for n in nights if i in n) for i in ids}
    Path("baseline.json").write_text(json.dumps(out))

def open_incidents():
    try:
        with urllib.request.urlopen(STATUS, timeout=10) as resp:
            return [i["name"] for i in json.load(resp)["incidents"]]
    except Exception as exc:  # a status page outage must not stop the run
        return [f"status check failed: {type(exc).__name__}"]

if __name__ == "__main__":
    tasks = [json.loads(line) for line in Path("tasks.jsonl").read_text().splitlines()]
    base = Path("baseline.json")
    baseline = json.loads(base.read_text()) if base.exists() else {}
    day = time.strftime("%Y-%m-%d", time.gmtime())
    Path("scores").mkdir(exist_ok=True)
    for name, arm in ARMS.items():
        scores, outcomes, tokens = run(arm, tasks)
        Path(f"scores/{day}-{name}.json").write_text(json.dumps(scores))
        row = {"day": day, "arm": name, **arm, "sdk": anthropic.__version__,
               "outcomes": dict(outcomes), "incidents": open_incidents(),
               "median_output_tokens": statistics.median(tokens) if tokens else None}
        if name in baseline:
            delta, half_width, n = paired_delta(scores, baseline[name])
            row.update(delta=round(delta, 4), ci95=round(half_width, 4), tasks=n)
        print(json.dumps(row))
        with open("drift_log.jsonl", "a") as log:
            log.write(json.dumps(row) + "\n")

Run it for a week or two, call build_baseline() with those dates, then keep it running nightly at a fixed hour from cron or a scheduled CI job on the same runner image. Pass the full model ID so the served-model check compares like with like, and pin the SDK version. For a GPT or Gemini model, replace ask() with that vendor’s client. The statistics stay the same.

With no real change, a single night’s 95% interval still excludes zero about one night in twenty, which is why step 8 asks for consecutive windows. The script also measures the API itself. If your users work in Claude Code or ChatGPT, the client’s defaults and fallbacks are part of their experience, and a harness that drives the client, as Livenerf’s does, is closer to it.

What Livenerf measures, and when it expects a verdict

Livenerf is an open-source project, independent of Anthropic, started on 22 September 2026, the day Opus 5.5 launched, to answer the question for that model with data.14 It was discussed on Hacker News from 29 September.2

It measures Opus 5.5 as served through Claude Code on a Max subscription, called headlessly with a pinned CLI (2.1.280), a frozen system prompt, no tools and an empty working directory. It runs on Inspect, the UK AI Security Institute’s open-source evaluation framework,27 and follows Miller’s statistics. The panel holds 78 questions from GPQA Diamond, MMLU-Pro, competition maths and AIME, kept from 2,336 screened questions because Opus 5.5 got them right only some of the time. Each runs once a day at high effort, Opus 5 runs the 12 GPQA questions as a control, and grading is exact-match with no LLM judge.19

As of 3 October, it had collected 10 of 30 planned daily runs, completing its baseline (days 1 to 10, from 24 September). The first results row is due after day 20 and the first possible verdict around 24 October, under the decision rule in step 8.

Its documentation lists the limits. A null result would mean no change detected at this sensitivity. In validation, swapping in Opus 5 for Opus 5.5 was not distinguishable at 99%. The panel doesn’t cover the raw API, newer Claude Code releases, long context, tool use or agentic work. A change during the baseline would become part of it, and launch week, the reference point, could itself have been an unusual week.

A checklist before you decide a model got worse

  • Look at the model each disappointing response reports, and for fallback entries or switch notices.
  • Check the effort or reasoning level those requests ran at, and whether a new model or client changed the default.
  • Compare your client, SDK and CLI versions with the date the change started.
  • Re-run the failing prompt several times in a fresh context.
  • Check the status page for incidents in the same window, and read what they describe.
  • Diff your prompts, tool definitions and graders against the last good run.
  • Run fixed tasks against a pinned model ID with repeats, a baseline and a control, and watch output tokens as well as scores.
  • Report a change with its interval and the number of tasks behind it.

  1. Anthropic, Claude Platform release notes, entries for 22 and 28 September 2026, https://platform.claude.com/docs/en/release-notes/overview ↩

  2. Hacker News, “Livenerf: Has Opus 5.5 been nerfed yet?”, 29 September 2026, https://news.ycombinator.com/item?id=49901736 ↩↩

  3. Anthropic, “A postmortem of three recent issues”, 17 September 2025, https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues ↩

  4. OpenAI, “New embedding models and API updates”, 25 January 2024, https://openai.com/index/new-embedding-models-and-api-updates/ ↩

  5. OpenAI, “Sycophancy in GPT-4o: what happened and what we’re doing about it”, 29 April 2025, https://openai.com/index/sycophancy-in-gpt-4o/ ↩

  6. OpenAI, “Expanding on what we missed with sycophancy”, 2 May 2025, https://openai.com/index/expanding-on-sycophancy/ ↩↩

  7. Google, Gemini API changelog, entry for 6 May 2025, https://ai.google.dev/gemini-api/docs/changelog ↩

  8. Julie Bort, “Sam Altman addresses ‘bumpy’ GPT-5 rollout, bringing 4o back, and the ‘chart crime’”, TechCrunch, 8 August 2025, https://techcrunch.com/2025/08/08/sam-altman-addresses-bumpy-gpt-5-rollout-bringing-4o-back-and-the-chart-crime/ ↩

  9. Lingjiao Chen, Matei Zaharia and James Zou, “How is ChatGPT’s behavior changing over time?”, arXiv:2307.09009, version 1 of 18 July 2023 and version 3 of 31 October 2023, https://arxiv.org/abs/2307.09009 ↩

  10. Arvind Narayanan and Sayash Kapoor, “Is GPT-4 getting worse over time?”, AI Snake Oil, 19 July 2023, https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-time ↩

  11. Anthropic, “Migrating to Claude Opus 5.5”, Claude Platform Docs, as of 3 October 2026, https://platform.claude.com/docs/en/models/opus-5-5/migration-guide ↩↩

  12. Livenerf, instrument validation report, generated 24 September 2026, https://github.com/ninjahawk/livenerf/blob/54145ea7b6/docs/VALIDATION.md ↩↩↩

  13. Anthropic, “Refusals and fallback”, Claude Platform Docs, as of 3 October 2026, https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback ↩↩

  14. Livenerf, README as of 3 October 2026, https://github.com/ninjahawk/livenerf/blob/54145ea7b6/README.md ↩↩↩

  15. Anthropic, “Model configuration”, Claude Code Docs, as of 2 October 2026, https://code.claude.com/docs/en/model-config ↩↩

  16. OpenAI Help Center, “Model Release Notes”, entry for 18 March 2026, https://help.openai.com/en/articles/9624314-model-release-notes ↩

  17. Horace He and Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference”, 10 September 2025, https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ ↩

  18. Kelly Hong, Anton Troynikov and Jeff Huber, “Context Rot: How Increasing Input Tokens Impacts LLM Performance”, Chroma, 14 July 2025, https://research.trychroma.com/context-rot ↩

  19. Livenerf, eval card, as of 3 October 2026, https://github.com/ninjahawk/livenerf/blob/54145ea7b6/docs/EVAL_CARD.md ↩↩

  20. Melanie Sclar, Yejin Choi, Yulia Tsvetkov and Alane Suhr, “Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting”, ICLR 2024, https://arxiv.org/abs/2310.11324 ↩

  21. Claude status page, https://status.claude.com/, and OpenAI status page, https://status.openai.com/ ↩

  22. npm, version history of the @anthropic-ai/claude-code package, https://www.npmjs.com/package/@anthropic-ai/claude-code?activeTab=versions ↩

  23. Anthropic, Claude Code changelog, versions 2.1.268 to 2.1.287, https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md ↩

  24. Anthropic, “Advanced setup”, Claude Code Docs, https://code.claude.com/docs/en/setup ↩

  25. Evan Miller, “Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations”, arXiv:2411.00640, 1 November 2024, https://arxiv.org/abs/2411.00640, and Anthropic’s summary, “A statistical approach to model evaluations”, 19 November 2024, https://www.anthropic.com/research/statistical-approach-to-model-evals ↩

  26. Livenerf, design report, generated 24 September 2026, https://github.com/ninjahawk/livenerf/blob/54145ea7b6/docs/DESIGN.md ↩

  27. UK AI Security Institute, Inspect, https://inspect.aisi.org.uk/ ↩

Frequently asked questions

Is Claude nerfed?

Anecdotes can’t settle it, and as of 3 October 2026 the Livenerf project expects its first possible verdict on Opus 5.5 around 24 October. Anthropic says it never reduces model quality because of demand, time of day or server load, and it traced degraded responses in August and September 2025 to three infrastructure bugs. You can test your own workload with fixed tasks, a pinned model ID, repeats and error bars.

Why is Claude worse now?

Possible causes include a lower default effort level, a safety-classifier fallback to another model, a client update, a long or cluttered context, and ordinary sampling noise. Check the served model, the effort level and the client version for the bad session, then re-run the prompt several times in a fresh context.

Is Claude Code getting worse?

Check the client first. Claude Code shipped ten releases between 22 September and 3 October 2026, Opus 5.5 starts at medium effort there, and a request flagged by a safety classifier can move the session to Opus 5 or Opus 4.8 until you switch back.

Is GPT getting worse?

OpenAI has documented behaviour changes behind unchanged names, including a GPT-4o update in ChatGPT that it rolled back in April 2025 for being sycophantic. A 2023 study found large behaviour shifts between GPT-4 snapshots, though critics showed that part of the measured decline came from the test design.

How do I monitor LLM drift?

Freeze a set of gradable tasks, pin the model ID, effort and client versions, and run each task several times a night. Compare per-task scores with a baseline using paired differences and 95% intervals, run a control model through the same code, and log the served model and any fallback.

Does temperature 0 make LLM output deterministic?

No. Anthropic’s Opus 5.5 migration guide says temperature 0 never guaranteed identical outputs on earlier models, and Opus 5.5 rejects non-default temperature values. Thinking Machines got 80 different completions from 1,000 temperature-0 runs of one prompt on Qwen3-235B.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.