Key takeaways
- Compare what a pipeline ships with what it planned; failure counters only see failures that reach them.
- A limiter created inside each call limits nothing across calls, so create one per tenant per process.
- Give 429s their own retry budget so throttling can’t use up the work’s deadline.
- In one build, 187,152 of 572,209 output tokens were reasoning and 49,544 were shipped prose.
- Record judge failures as unreachable or unparseable, and report the sections it never reviewed.
An LLM pipeline can report zero failures while most of its output is missing. One of ours did. A course build planned 51 diagrams and shipped 26, planned 44 simulations and shipped 13, and put flashcards and quizzes on 7 of its 36 sections. The build’s own report said “0 dropped”.
The cause was a concurrency limiter created per call. Every module ran its own limiter, so all of them hit one Azure OpenAI deployment at the same time. The deployment answered with 429 rate-limit responses. They arrived when the work had no deadline left to retry in, and the work that was given up wasn’t counted. The rule we took from this is to compare what a pipeline ships with what it planned, because a failure counter only sees the failures that reach it.
The pipeline turns textbooks into courses for an AI study platform used by high-school and college students. The companion post covers how it reads the books and merges topics. This one covers the throttling failure, where the tokens went, an LLM judge that left most sections unreviewed, and the numbers before and after the fixes.
What the build planned and what it shipped
| Output | Planned | Shipped | Missing |
|---|---|---|---|
| Diagrams | 51 | 26 | 25 |
| Simulations | 44 | 13 | 31 |
| Sections with flashcards and quizzes | 36 | 7 | 29 |
The report for that build put the number dropped at 0.
A report like that adds up the failures the code knows about. Work that is abandoned without passing through the counter never appears in it, for example a task that gives up quietly when its deadline passes. Comparing the plan with the output is a different measurement. It finds a gap whatever caused it, and the table above comes from that comparison.
How a per-call limiter let every module hit one deployment
A course build fans out into modules that run in parallel. The concurrency limiter was there to cap how many model calls were in flight at once. It was created per call, so every module ran its own. A limiter only limits the callers that share it, so the cap applied to each module separately and the modules together hit the deployment at once.
Here is the shape of the bug and of the fix, written for this post with asyncio:
import asyncio
# Bug: each module creates its own semaphore, so the cap applies per module.
async def build_module(module):
limit = asyncio.Semaphore(8)
async def run(task):
async with limit:
return await call_model(task)
return await asyncio.gather(*(run(t) for t in module.tasks))
# With 20 modules running at once, that is up to 160 requests in flight.
# Fix: one limiter per tenant, created once and shared by every module in the process.
_limits: dict[str, asyncio.Semaphore] = {}
def tenant_limit(tenant_id: str) -> asyncio.Semaphore:
if tenant_id not in _limits:
_limits[tenant_id] = asyncio.Semaphore(8)
return _limits[tenant_id]
All of those requests went to one Azure OpenAI deployment. Azure assigns quota to each deployment in tokens per minute and sets a requests-per-minute limit in proportion to it 1. Two details in Azure’s documentation explain why bursts are expensive. The request rate is checked over short windows, typically 1 or 10 seconds, so a burst can be throttled while the minute’s total is still under the limit. And the token limit is checked against an estimate made when each request arrives, which counts the prompt and the max_tokens setting instead of the tokens the reply turns out to use 1. Rejected requests still count toward the per-minute limit, so resending without a pause keeps a deployment throttled.
Now each worker process holds one limiter per tenant, created once and shared by every module in the process. Across processes, the worker fleet also shares a Redis-backed rate limiter. Azure returns headers with every response that report the requests and tokens remaining (x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens), and a limiter can use them to slow down before the 429s start 1.
Why throttled work ran out of time
Each piece of work runs against a deadline. gRPC’s documentation defines a deadline as a point in time after which the caller won’t wait for a response 2. Everything inside it draws on the same allowance. Waiting for a limiter slot uses it, the request uses it, and so does every retry.
With every module competing for one deployment, the 429s arrived when the work had no deadline left to retry in. A 429 from Azure carries a retry-after-ms header that says how long to wait 1. The work had no time left to wait, so it was given up, and nothing counted it.
The fix is a separate retry budget reserved for 429s. A throttled call waits from that reserve, so throttling no longer spends the time meant for the work. A sketch of the idea, again written for this post:
import random
import time
def send_with_reserve(send, deadline: float, reserve_s: float = 120.0):
"""Retry 429s from their own time reserve instead of the work's deadline."""
attempt = 0
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
raise DeadlineExceeded() # the caller counts it
resp = send(timeout=remaining)
if resp.status_code != 429:
return resp
attempt += 1
hint = resp.headers.get("retry-after-ms")
wait = int(hint) / 1000 if hint else random.uniform(0, min(30.0, 2.0 ** attempt))
if wait > reserve_s:
raise ThrottledOut(attempt) # the caller counts this too
reserve_s -= wait
time.sleep(wait)
deadline += wait # throttled time doesn't use up the work's deadline
The general rules for retries apply here too. Azure’s documentation recommends exponential backoff with random jitter when there’s no retry-after-ms, and a cap on attempts 1. Marc Brooker’s article in the Amazon Builders’ Library explains why jitter matters and why retries belong at one layer. In his example, three retries at each of five layers multiply the load on the bottom layer 243 times 3. The “full jitter” in the sketch, a random wait between zero and the backoff cap, comes from his earlier analysis of backoff strategies 4.
Retry loops also stack. The OpenAI Python SDK retries 429s twice by default 5, so a retry loop of your own around it turns each of your attempts into as many as three requests. Azure’s documentation says to set max_retries=0 on the SDK when you add your own retries 1. Google’s SRE book caps each request at three attempts, and lets a caller retry only while retries stay under 10% of its requests 6.
Where 572,209 output tokens went
One build’s token usage:
| Output tokens | Tokens | Share |
|---|---|---|
| Reasoning | 187,152 | 32.7% |
| Prose that shipped in the course | 49,544 | 8.7% |
| All other output (not broken down here) | 335,513 | 58.6% |
| Total | 572,209 | 100% |
| Input tokens | Tokens | Share |
|---|---|---|
| Served from the prompt cache | 19,712 | 2.9% |
| Not cached | 648,685 | 97.1% |
| Total | 668,397 | 100% |
Reasoning tokens never appear in the response, but they take up room in the context window and are billed as output tokens 7. In this build the pipeline spent about 3.8 reasoning tokens for every token of prose that shipped.
Not every step needs much reasoning. A step that only reshapes text that already exists, such as filling a fixed schema from a finished explanation, is mechanical. Azure’s guidance pairs low effort with work where speed and cost matter, and notes that higher effort means longer requests and, generally, more reasoning tokens 7. We moved the mechanical steps to low reasoning effort. If you do the same, compare the output before and after, because lower effort trades depth of reasoning for speed by design 8.
Reasoning also changes what failure looks like. A reasoning model can hit its output cap before it writes any visible output. You pay for the input and reasoning tokens and get no answer, which is why Azure’s documentation says to check the status of every response instead of treating it as an empty result 7.
Only 2.9% of input tokens came from the prompt cache. On Azure OpenAI, a prompt is eligible when it is at least 1,024 tokens long, and its first 1,024 tokens have to match an earlier prompt exactly. One character of difference in that prefix is a miss, and in-memory caches are typically cleared after 5 to 10 minutes without use 9. Common reasons for a low hit rate:
- The material shared between calls isn’t at the start of the prompt.
- The start of the prompt changes between calls, even by a timestamp or an ID.
- The prompts are shorter than 1,024 tokens.
- Calls that share a prefix are too far apart in time.
Both OpenAI and Azure advise putting stable content first and variable content last 9 10.
A ledger per model and per kind of work
Totals like these don’t say which step to change. Provider dashboards don’t either. Azure’s usage metrics count billed tokens from requests that succeeded, so throttled attempts never show up there 1.
We now record spend per model and per kind of work, such as writing an explanation, drawing a diagram or judging a section. The simplest version is one row per model call, with the token counts from the usage block of the response. A sketch of such a row:
from dataclasses import dataclass
@dataclass
class LedgerRow:
build: str
model: str
kind: str # e.g. "explanation", "diagram", "judge"
status: str # "ok", "incomplete", "throttled_out", "deadline"
input_tokens: int
cached_tokens: int
output_tokens: int
reasoning_tokens: int
def row_from(resp, **labels) -> LedgerRow:
u = resp.usage
return LedgerRow(
input_tokens=u.prompt_tokens,
cached_tokens=getattr(u.prompt_tokens_details, "cached_tokens", 0) or 0,
output_tokens=u.completion_tokens,
reasoning_tokens=getattr(u.completion_tokens_details, "reasoning_tokens", 0) or 0,
**labels,
)
The field names are the ones Azure’s Chat Completions responses use 9. Grouping the rows by model and by kind shows where the tokens go, and the status column shows which calls produced nothing. On newer model families Azure also reports cache writes, which can be billed 9. If your model reports them, give them a column too.
When the judge cannot reach a verdict
An LLM judge reviews each section. In another run it left 27 of 35 sections unreviewed. The counters for that run showed 774 deadline exhaustions and zero throttle retries, which put the problem in the judge’s time budget.
We made two changes. The judge’s time budget went from 90 s to 420 s. And judge errors are now split into two kinds, unreachable and unparseable.
| Judge outcome | What happened | Where the fix usually is |
|---|---|---|
| Pass or fail | The judge returned a verdict | The content |
| Unreachable | No reply within the budget, or the call was refused | Time budget, limiter or quota |
| Unparseable | A reply arrived but didn’t match the expected format | The judge’s prompt or output format |
A section in either failure state is unreviewed, and the report should say so. For unparseable replies, OpenAI’s Structured Outputs makes the model follow a JSON Schema you supply, though a refusal or a reply cut off at the token limit can still fall outside it 11.
For the budget itself, Brooker suggests picking an acceptable rate of false timeouts and setting the timeout at the matching latency percentile, such as p99.9 for a 0.1% rate 3. Measure the judge’s latency under the concurrency the pipeline really runs at, since waiting for a limiter slot counts against the same budget. Reasoning effort matters here too, because higher effort means longer requests 7.
Before and after the fixes
The later build ran after five fixes:
- One limiter per tenant for the whole process.
- A separate retry budget reserved for 429s.
- Low reasoning effort for mechanical steps.
- Spend recorded per model and per kind of work.
- Judge errors split into unreachable and unparseable, with the judge’s time budget raised from 90 s to 420 s.
| Measure | Before | After |
|---|---|---|
| Sections at full length | 18 of 38 | 33 of 33 |
| Simulations | 12 | 33 |
| Total words | 20,904 | 43,644 |
| Words per section (derived) | 550 | 1,323 |
The two builds had different numbers of sections, 38 before and 33 after, so the per-section row is the fairer comparison for length. Every section in the later build reached full length, and the total word count roughly doubled even though the build had fewer sections.
How we measured
- Planned against shipped. One build with 36 sections. Planned counts are what the build’s plan called for, and shipped counts are what the finished course contained.
- Token mix. One build. Reasoning and cached counts come from the usage data the API returns with each response.
- Judge. A separate run with 35 sections. Deadline exhaustions and throttle retries are that run’s recorded counts.
- Before and after. Two more builds, with 38 and 33 sections. The words-per-section row is derived from the totals.
- Not covered. Cost in currency, latency, the effect of each fix on its own, and quality beyond completeness and length. Prices vary by model and deployment type, so we report tokens.
A checklist for LLM batch pipelines
- Compare planned with shipped for every build and every kind of output, and put that comparison at the top of the build report.
- Create each limiter once, at the scope of the quota it protects, and give every caller the same object. Count requests in flight to confirm it works.
- Read
x-ratelimit-remaining-requestsandx-ratelimit-remaining-tokens, and slow down before the 429s start. - Keep
max_tokensno larger than the task needs. Azure’s rate-limit estimate counts it before any tokens are generated. - Retry at one layer. If you write your own retry loop, set the SDK’s
max_retriesto 0. - Give 429s their own retry budget, and wait for
retry-after-mswhen the response includes it. - Log input, cached, output and reasoning tokens for every call, labelled with the model and the kind of work.
- Run mechanical steps at low reasoning effort, and check the output before and after the change.
- Check the status of every response, so an incomplete reply is never treated as an empty one.
- Record judge failures as unreachable or unparseable, and report unreviewed sections as unreviewed.
-
Microsoft Learn, “Manage Azure OpenAI in Microsoft Foundry Models quota”, https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/quota. ↩↩↩↩↩↩↩
-
gRPC documentation, “Deadlines”, https://grpc.io/docs/guides/deadlines/. ↩
-
Marc Brooker, “Timeouts, retries, and backoff with jitter”, Amazon Builders’ Library, https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/. ↩↩
-
Marc Brooker, “Exponential Backoff And Jitter”, AWS Architecture Blog, 4 March 2015, https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/. ↩
-
OpenAI Python library, README, “Retries”, https://github.com/openai/openai-python#retries. ↩
-
Google, Site Reliability Engineering, chapter 21, “Handling Overload”, https://sre.google/sre-book/handling-overload/. ↩
-
Microsoft Learn, “Azure OpenAI reasoning models”, https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning. ↩↩↩↩
-
OpenAI API documentation, “Reasoning models”, https://developers.openai.com/api/docs/guides/reasoning. ↩
-
Microsoft Learn, “Prompt caching with Azure OpenAI in Microsoft Foundry Models”, https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching. ↩↩↩↩
-
OpenAI API documentation, “Prompt caching”, https://developers.openai.com/api/docs/guides/prompt-caching. ↩
-
OpenAI API documentation, “Structured Outputs”, https://developers.openai.com/api/docs/guides/structured-outputs. ↩
Frequently asked questions
Why does Azure OpenAI return 429 errors when my token usage is below quota?
Azure checks each request against an estimate made when it arrives, which includes the max_tokens setting, and it measures request rate over windows of 1 or 10 seconds. Bursts and generous max_tokens values can trigger 429s that billed-token metrics never show.
Are reasoning tokens billed as output tokens?
Yes. OpenAI and Azure OpenAI both document that reasoning tokens never appear in the response but are billed as output tokens. In one of our builds they were 187,152 of 572,209 output tokens.
How should I retry 429 rate-limit errors from an LLM API?
Wait for the retry-after-ms value when the response includes it. Otherwise back off exponentially with jitter, cap the attempts, and retry at one layer only. Giving throttling its own budget means a 429 that arrives late in a job can still be retried.
How do I check prompt cache hits on Azure OpenAI?
Read cached_tokens in the usage block of each response, under prompt_tokens_details in Chat Completions. A prompt is eligible once it is at least 1,024 tokens long and its first 1,024 tokens match an earlier prompt, so put stable content first.
What should an LLM judge step report when it fails?
Why it failed. Unreachable (no reply in time, or the call was refused) and unparseable (a reply in the wrong format) need different fixes, and a section in either state should be reported as unreviewed.
How do I track LLM spend per feature?
Write one ledger row per model call with the model, the kind of work and the input, cached, output and reasoning token counts from the usage block. Grouping by model and by kind shows where the tokens go.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.