Get in touch

9io.ai / Blog

How we halved an AI stylist’s latency, and when hedging stops paying

A streaming AI stylist went from a 30 s median response to 13 to 15 s. The four changes, how we measured them, and why hedged requests stopped paying.

Key takeaways

  • The median response fell from 30.1 s to 13.4 s on replay and to 15.5 s on live traffic.
  • Reasoning tokens are generated and billed as output, so they cost time even though nobody reads them.
  • A hedge can beat a slow server but cannot shorten an answer that is slow because it is long.
  • Once our tail shrank, the duplicate won 4 of 142 hedges, so log the win rate from day one.
  • A quality suite ran next to the latency work and went from 18 to 47 of 47 passing.

The AI stylist on a fashion commerce platform we build and run used to take a median of 30.1 s to answer. After tuning it takes 13.4 s on a replay set and 15.5 s on four days of live traffic. The p90 went from 40.3 s to 19.1 s on replay and 20.6 s live. On either measure, the median and the p90 roughly halved.

The latency work came down to four changes. We turned reasoning off for the step that composes the answer, and added hedged requests to cut slow outliers. We called OpenAI’s API directly instead of Azure OpenAI, where we had seen a slow tail. Retrieval now stops as soon as it has enough results.

Hedging helped while the tail was long, then stopped paying. In a later sample, after the other changes had shrunk the tail, it fired 142 times and the duplicate finished first only 4 times. The reason fits in one line of algebra and applies to anyone hedging LLM calls. This post covers where the time goes in a chained turn, each change, why hedging stopped paying, how we measured, and a checklist for your own chain.

The numbers before and after

The stylist streams its answer to the shopper and can take actions, such as placing an order. Each turn runs a chain of model calls. One classifies the request, one composes the answer, one validates it and one critiques it. Underneath sits hybrid keyword-plus-vector retrieval in Azure AI Search.

The figures below are response times per turn. The baseline is 394 real turns from before tuning. After tuning there are two measurements, a replay of 12 requests and real traffic over four days in late September 2026.

Measurement Sample p50 p90
Before tuning 394 real turns 30.1 s 40.3 s
After, replay 12 replayed requests 13.4 s 19.1 s
After, live traffic Four days, late September 2026 15.5 s 20.6 s

A replay set sends the same requests every time, which makes it good for comparing versions, but 12 requests can’t pin down a p90. The live figures are what shoppers waited, and they are the like-for-like comparison with the baseline. Against it, the live median is 48.5% lower and the live p90 is 48.9% lower.

A quality pass ran over the same period. Latency work on an LLM product needs one running alongside, because a faster answer that is also a worse one is easy to ship without noticing.

Quality check Before After
Regression suite passing 18 of 47 47 of 47
Answers judged bad in review 21 of 116 1 of 116

Where the time goes in a chained turn

Azure’s latency guide writes the time for one model call as a formula.1

time to last token = time to first token + (time between tokens × tokens generated)

Generated tokens dominate. OpenAI’s latency guide2 puts rough numbers on it. Halving the output can cut latency by about half, while halving the prompt may improve it by only 1 to 5%. Streaming lets the shopper start reading while the model is still writing.3 Azure’s guide adds that streaming leaves the time to receive every token unchanged. A streamed answer also still waits for every step that runs before it.

In a chain, the calls run one after another and their times add up. Every generated token in every step costs time, including tokens the shopper never sees. Tails compound as well. Dean and Barroso’s example is a server that is slow on 1 request in 100. A request that has to collect answers from 100 such servers is slow 63% of the time.4 A sequential chain is a milder case of the same effect. As an illustration, if each of four steps were slow on 1 call in 20, about 18.5% of turns would hit at least one slow step.

That leaves two levers. Generate fewer tokens per turn, and remove the variance that makes the slowest turns slow. Our four changes split across the two.

Change What it mainly targets
Reasoning off for the compose step Tokens generated per turn, and how much that count varies
Hedged requests Slow outliers from the serving side
OpenAI’s API instead of Azure OpenAI A slow tail we saw on Azure OpenAI
Early exit in retrieval Retrieval work per turn

Turning reasoning off for the compose step

The compose step writes the answer the shopper reads. We turned reasoning off for it.

Reasoning models use internal reasoning tokens before they produce a response. The API doesn’t return them, but they take up space in the context window and are billed as output tokens.5 In the formula above they are generated tokens, so they add time to the call. OpenAI’s guide says a model may produce anywhere from a few hundred to tens of thousands of them, and that lower effort favours speed and lower token use. Because the count varies from call to call, reasoning adds spread to latency as well as raising the median.

In OpenAI’s API the setting is reasoning.effort. Supported values depend on the model, and the guide describes none as suited to latency-critical tasks that don’t benefit from reasoning. Some models don’t accept none, so check the model page first.

Whether a step benefits from reasoning is an empirical question. A step that writes prose from material it has already been handed is a natural first candidate for turning it off. The way to settle it is to run the step both ways against an evaluation set, which is what a regression suite and a review set are for.

Hedged requests and where to set the delay

Dean and Barroso described hedged requests in “The Tail at Scale”.4 You send a request, and if it hasn’t come back after a delay, you send a copy and use whichever response arrives first. They suggest waiting until the first request has been outstanding longer than the 95th-percentile latency for its class of requests, which limits the extra load to about 5%. In one of their benchmarks, reading 1,000 keys from a BigTable table spread across 100 servers, a hedge sent after 10 ms cut the 99.9th-percentile latency from 1,800 ms to 74 ms, for 2% more requests.

For an LLM call the shape is the same. Here is a minimal version with Python’s asyncio.

import asyncio

async def hedged(make_call, delay_s):
    """Start make_call(). If it is still running after delay_s, start a copy.
    Return the first result, cancel the other call, and report who won."""
    primary = asyncio.create_task(make_call())
    done, _ = await asyncio.wait({primary}, timeout=delay_s)
    if done:
        return primary.result(), "not_fired"

    duplicate = asyncio.create_task(make_call())
    done, pending = await asyncio.wait(
        {primary, duplicate}, return_when=asyncio.FIRST_COMPLETED
    )
    for task in pending:
        task.cancel()
    winner = done.pop()
    return winner.result(), ("primary" if winner is primary else "duplicate")

A few details decide whether this saves time or only costs money.

  • The copy is a full second request. It sends the whole prompt again and counts against your request and token rate limits like any other call.6
  • Cancel the loser as soon as the winner returns.
  • Handle a copy that fails fast. In the sketch above, an error that arrives first wins the race.
  • If the call streams to the shopper, settle the race on the first token, since only one stream can be shown.
  • Set the delay from that call’s own latency distribution, near its p95, and revisit it when the distribution moves.
  • Log the outcome of every hedge. The section on why hedging stopped paying shows what to look for.

Calling OpenAI directly instead of Azure OpenAI

We switched from Azure OpenAI to calling OpenAI’s API directly, to avoid a slow tail we saw on Azure.

Azure’s guide to provisioned throughput7 describes the trade-off between its deployment types. On standard deployments, inference capacity is shared across customers and throughput can vary with demand, and the standard tier carries no latency SLA. Provisioned throughput and priority processing both come with a defined latency target per model. Azure’s latency guide1 names the overall load on the deployment among the main factors in per-call latency, and notes that content filtering adds latency in return for safety.

If you see a tail, check token counts before you blame the endpoint. A call that got slower because it generated more tokens will be slow anywhere, and Azure’s guide makes the same point when it says to pair every latency metric with a token count. When time to first token or time between tokens rises while token counts stay flat, the cause is on the serving side and a different route can help. For us that route was another endpoint. If you need to stay on Azure, for data residency or an existing agreement, provisioned throughput and priority processing are the documented ways to get steadier latency.

Letting retrieval stop once it has enough

Retrieval uses hybrid search in Azure AI Search. A hybrid query is one request that runs full-text and vector search in parallel and merges the two result lists with Reciprocal Rank Fusion.8 Our retrieval stage used to run to completion on every turn. We changed it to return as soon as it has enough results for the compose step.

A generic version of the pattern looks like this.

def retrieve(query, searches, min_results):
    """Run searches in priority order and stop once there are enough results."""
    results = []
    for search in searches:
        results = merge_unique(results, search(query))  # keep order, drop repeats
        if len(results) >= min_results:
            break
    return results

Two choices decide whether an early exit is safe. The order of the searches matters, because the first ones to fill the list decide what the model sees. The threshold matters too, because “enough” for the compose step is a quality question. Both belong under the regression suite.

Image search on the same platform has a different retrieval problem, matching a shopper’s photo to a catalogue product. We compared embedding models for it in Comparing FashionSigLIP, CLIP and DINOv2 for visual product search.

Why hedging stopped paying

In a later sample, hedging fired 142 times and the duplicate finished first 4 times, in 2.8% of fires. The other 138 duplicates were extra requests that changed nothing for the shopper.

Splitting the time of a call into two parts shows why. W is the work the request carries, mostly the tokens it will generate, and both copies carry the same W. N is noise from the serving side, such as queueing, a busy replica or a slow network path, and each copy draws its own.

The hedge fires when the primary is still running at the delay d. Counting from the original start, the primary finishes at W + N₁ and the duplicate at d + W + N₂. The duplicate wins only when N₁ − N₂ > d.

W cancels out. A long answer gives the duplicate no help, because the duplicate has to generate the same long answer and starts d seconds later. Only a gap in serving-side noise larger than the delay lets it win. The sources of variability Dean and Barroso list are serving-side ones, such as shared resources, background daemons, maintenance work and queueing.4

Our other changes cut into that noise. The slow tail on Azure OpenAI was serving-side, the kind of slowness a hedge can beat, and moving the calls was aimed at it. Reasoning can behave like noise too. Two copies of one request may generate very different numbers of reasoning tokens, so with reasoning on, part of W differs between the copies. With less noise left, the calls still running at the hedge delay are mostly the long ones, with a lot of W. The hedge fires on them and the duplicate loses, which is the pattern in our 142 fires.

On the cost side, each fire is a full extra request, with its prompt sent again and its share of your rate limits. At 4 wins in 142 fires, almost all of that buys nothing.

Hedging pays while the time it saves on wins is worth more than the cost of all the fires. Track three numbers for each hedged call site.

  • Fire rate, the share of calls still running at the delay.
  • Win rate, the share of fires where the duplicate finishes first.
  • Time saved per win. Measuring it means letting a sample of losing primaries run to the end, because cancelling them hides when they would have finished.

When the win rate collapses, as ours did, raise the delay, hedge fewer call sites, or switch hedging off and measure again after the next change to the serving path.

How we measured

The baseline is 394 real turns recorded before tuning. The tuned version was measured twice, on a replay of 12 requests and on real traffic over four days in late September 2026. Every latency figure is a p50 or p90 of response time per turn.

With 12 values, a p90 is in effect the second-slowest request, so the replay p90 of 19.1 s is a rough number. The live p90 of 20.6 s comes from four days of traffic and is the better guide. The hedging counts come from a later sample of live traffic. The quality figures come from a regression suite of 47 cases and a review of 116 answers, each run before and after.

What these numbers don’t show:

  • How much each change contributed. The figures cover the four changes together.
  • Time to first token. This post reports response time per turn only.
  • Cost, either the tokens saved by turning reasoning off or the tokens spent on duplicates.
  • The same requests before and after. The baseline is real traffic and the replay set is a fixed dozen, which is why the live figures matter most.

A checklist for a slow LLM chain

  1. Time each step and log its generated tokens before changing anything. The formula then tells you whether a slow call generated a lot or was served slowly.
  2. Treat reasoning effort as a per-step setting. Turn it down where an evaluation set says the step doesn’t need it.
  3. Add hedging with its outcomes logged from the first day, and set the delay near the call’s own p95.
  4. Recheck the win rate after every change to the serving path. Each fix that removes serving-side noise leaves hedging less to do.
  5. Before moving providers, pair latency with token counts. Move when time to first token or time between tokens is the problem.
  6. Let retrieval stop when it has enough, with the search order and the threshold covered by tests.
  7. Quote live percentiles next to replay numbers, with the size of each sample.

  1. Microsoft Learn, “Azure OpenAI in Microsoft Foundry Models performance & latency”, https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/latency. ↩↩

  2. OpenAI API docs, “Latency optimization”, https://developers.openai.com/api/docs/guides/latency-optimization. ↩

  3. OpenAI API docs, “Streaming API responses”, https://developers.openai.com/api/docs/guides/streaming-responses. ↩

  4. Jeffrey Dean and Luiz André Barroso, “The Tail at Scale”, Communications of the ACM 56(2), February 2013, pages 74 to 80, https://doi.org/10.1145/2408776.2408794. ↩↩↩

  5. OpenAI API docs, “Reasoning models”, https://developers.openai.com/api/docs/guides/reasoning. ↩

  6. OpenAI API docs, “Rate limits”, https://developers.openai.com/api/docs/guides/rate-limits. ↩

  7. Microsoft Learn, “Provisioned throughput for Foundry Models”, https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput. ↩

  8. Microsoft Learn, “Hybrid search overview”, Azure AI Search, https://learn.microsoft.com/en-us/azure/search/hybrid-search-overview. ↩

Frequently asked questions

What is a hedged request?

A copy of a request, sent after a delay if the first hasn’t answered, where you keep whichever response arrives first. Dean and Barroso suggest waiting until the first request has run past the 95th-percentile latency for its kind, which keeps the extra load near 5%.

Do hedged requests help with LLM API latency?

They help when the slowness comes from the serving side, such as a busy replica or a queue. They can’t help a response that is slow because it is long, since the copy has the same tokens to generate and starts later.

Does turning off reasoning make an LLM respond faster?

Usually. Reasoning tokens are generated and billed as output tokens, and OpenAI’s guide says lower effort favours speed and lower token use. Check answer quality on your own test set before and after.

Is Azure OpenAI slower than the OpenAI API?

It depends on the deployment type and the load on it. We saw a slow tail on Azure OpenAI and moved to OpenAI’s API to avoid it. On Azure, provisioned throughput and priority processing come with latency targets and standard deployments don’t.

How should I measure the latency of an LLM assistant?

Record p50 and p90 for whole turns from real traffic, plus per-step timings with token counts. A fixed replay set is good for comparing versions, but its p90 means little with only a dozen requests.

When should I switch hedging off?

When the duplicate rarely finishes first. Log every hedge that fires and which copy wins. In a later sample ours fired 142 times and the duplicate won 4, so nearly every duplicate was wasted work.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.