Get in touch

9io.ai / Blog

How a backfill bug ran up a $1,100 LLM vision bill

A job re-classified items it had already classified: 419,057 vision calls and about $1,100. What went wrong, and the guardrails that catch it early.

Key takeaways

  • About $1,100 went on 419,057 vision calls that re-classified items already classified.
  • A variable named last2Days held 730 days. Put units in names, or use a duration type.
  • Rate limits cap throughput and provider spend data lags, so keep your own ledger per job and model.
  • On an embedded widget, bots made 2,380 generations against 84 from real users.
  • Widget traffic now runs on a cheaper model, and generations are capped at 6,000 a day.

About $1,100 of LLM vision spend on a fashion commerce platform we build and run came from 419,057 calls it didn’t need to make. A background job was sending items that had already been classified back to a vision model to be classified again. The variable that set how far back the job looked was named last2Days, and it held 730 days.

Separately, bots using an embeddable widget on a partner’s website made 2,380 generations, against 84 from real users. Neither problem was an error in the usual sense. The calls worked, and the problem showed up as spend.

Below are both causes, why provider-side limits make a late backstop, and six guardrails for any job or endpoint that calls a paid model. Code is included where it fits in a few lines.

Where the $1,100 went

419,057 calls for about $1,100 works out at roughly $0.0026 a call, about a quarter of a cent. Nobody reviews a quarter-cent call, and at that price a job can make hundreds of thousands of them before the total stands out.

The job had two faults, and they compounded.

What the code said What it did
Lookback window A variable named last2Days Held 730 days, which is two years
Items already classified No check for an existing result Sent to the model again

A window of 730 days is 365 times wider than the name says, and with no check for existing results, everything the job selected from those two years went back through the model.

Why provider limits make a late backstop

Provider rate limits cap throughput. OpenAI measures requests and tokens per minute and per day, at organisation and project level (rate limits). Gemini applies requests per minute, tokens per minute and requests per day to each project (rate limits). A background job that runs steadily below those numbers never sees a 429, however much it spends over a month.

Gemini also limits spend over a rolling 10-minute window, from $10 on Tier 1 to $200 on Tier 3 (spend-based rate limits). That slows a runaway job down, but $200 every 10 minutes still allows $28,800 a day.

Budgets and alerts arrive later still. Google’s Cloud Billing docs say an alerts-only budget doesn’t cap usage, and the first notification after you create a budget can take several hours (budgets). The Gemini billing page says its cost graphs can take up to 24 hours to update (billing). OpenAI says enforcement of its hard spend limits isn’t instantaneous (spend limits).

Hard caps exist and are worth setting, but the provider-wide ones are blunt. Gemini’s billing tier cap pauses every project on the billing account until the 1st of the next month. Anthropic pauses an organisation’s API usage until the next month once it reaches its tier’s cap (rate limits). A runaway backfill that reaches one of those caps takes your user-facing features down with it.

As of 8 October 2026, the controls the main providers document look like this:

Provider Hard limit you can set Scope What else the docs say
OpenAI Hard spend limit. Requests over it get a 429 Organisation or project Spend alerts only notify. Enforcement isn’t instantaneous, so spend can run slightly past the limit
Gemini API Monthly project spend cap, set in AI Studio Project Marked experimental, with overages possible for about 10 minutes. The billing tier cap pauses every project on the billing account
Google Cloud Billing None. Budgets send alerts Billing account, or projects, services or labels within it A budget doesn’t cap spending by itself. Pub/Sub notifications can trigger an automated response, such as disabling billing
Anthropic Spend limit below your tier’s cap. Requests over it get a 400 Organisation or workspace Reaching the tier’s own cap pauses API usage until the 1st of the next month

Use the provider’s controls as the outer layer. Put each batch job in its own project or workspace with a lower limit than the features people use, so a runaway job stops itself and leaves those features running.

Jobs that skip work already done

Brandur Leach’s post on idempotency at Stripe describes an idempotent operation as one that is safe to repeat: however many times it is called, its side effects happen once. For a batch job that calls a model, the side effect that matters is the paid call. Run a correct job twice and the second run should find nothing to do.

That takes two things. The selection excludes finished work, and each result is written together with the version of the model and prompt that produced it. A deliberate upgrade can then select old results on purpose.

-- Items with no classification from the current classifier version
SELECT id, image_url
FROM items
WHERE created_at >= now() - interval '2 days'
  AND (classifier_version IS NULL OR classifier_version < :current_version)
ORDER BY id
LIMIT 500;

If several workers share the queue, claim rows with FOR UPDATE SKIP LOCKED before calling the model, so two workers never pay for the same item. In PostgreSQL, rows another transaction has locked are then skipped instead of waited on (SELECT docs).

Then add a pre-flight check. Before the first paid call, the job counts what it selected, multiplies by a measured cost per call and refuses to continue above a threshold unless someone passes an explicit flag.

COST_PER_CALL_USD = 0.003     # measured average from your own ledger
MAX_UNATTENDED_USD = 25.0     # per run

def preflight(selected: int, allow_large: bool = False) -> None:
    estimate = selected * COST_PER_CALL_USD
    log.info("selected %d items, estimated cost %.2f USD", selected, estimate)
    if estimate > MAX_UNATTENDED_USD and not allow_large:
        raise SystemExit(f"estimate {estimate:.2f} USD is over {MAX_UNATTENDED_USD:.2f} USD; "
                         "rerun with --allow-large if this is intended")

A run whose estimate is over the threshold stops before its first call, with the estimate in its log. Retries are the other common way a job repeats work, and we covered which model errors are worth retrying in how AI image pipelines fail quietly.

Units in names, and a bound on the window

The variable was named last2Days and held 730. Nothing in the job compared the two, and the program only ever reads the value.

A name that contains a value goes stale as soon as someone changes the value. Name the unit instead, as in lookback_days, or use a duration type so the unit travels with the number. Python’s timedelta represents a duration (docs), and Go’s time.Duration stores one as a nanosecond count (docs). Then put a bound on it.

from datetime import datetime, timedelta

LOOKBACK = timedelta(days=2)
MAX_LOOKBACK = timedelta(days=7)

def window_start(now: datetime) -> datetime:
    if LOOKBACK > MAX_LOOKBACK:
        raise ValueError(f"lookback {LOOKBACK} is longer than {MAX_LOOKBACK}")
    return now - LOOKBACK

Log the resolved window at the start of every run as a date a person can read. A log line saying the job is classifying everything since a date two years ago is hard to miss.

If you do need to re-classify two years of items, after a model upgrade for example, run it as a separate backfill with its own budget and a named person who approved it. Keep it out of the daily job’s code path.

A spend ledger per job and per model

Write one row for every paid call, at the moment you make it:

Column What goes in it
job The job or endpoint that made the call
model The model ID you called
units Input and output tokens, or images
cost_usd Units multiplied by the price in your price table
item_id What the call was for
request_id The provider’s ID for the call, if it returns one
created_at When the call was made

Token counts come back with the response. Gemini, for example, returns usageMetadata with the prompt and candidate token counts for each request (API reference). For token accounting inside a pipeline, why our LLM pipeline reported 0 dropped shows a ledger row built from each response’s usage block. The columns above add the job, the cost and the provider’s request ID, which is what spend control needs.

Roll the ledger up by day, job and model, so any line on a bill can be traced to the job and model that ran it. Keep the price table in config with the date you last checked each price, because prices change.

Daily caps and alerts that fail small

A ledger shows the damage after the fact, so each job also needs a daily budget that it checks in code. Charge the estimated cost against the budget before each paid call, record the actual cost after it, and refuse the call once the budget is spent:

import datetime as dt
import redis

r = redis.Redis()

class BudgetExceeded(Exception):
    pass

def charge(job: str, estimate_usd: float, daily_budget_usd: float) -> None:
    day = dt.datetime.now(dt.timezone.utc).strftime("%Y-%m-%d")
    key = f"spend:{job}:{day}"
    total = r.incrbyfloat(key, estimate_usd)
    r.expire(key, 2 * 24 * 3600)
    if total > daily_budget_usd:
        raise BudgetExceeded(f"{job}: {total:.2f} USD today, budget {daily_budget_usd:.2f} USD")

Alert well before the cap, for example at half of it, in a channel a person reads. When a job hits its cap it should stop and tell someone. Retrying into a spend limit gets nowhere. Anthropic’s docs make the same point about their own cap: the 429 it returns carries no retry-after header, and retries fail until access resumes.

Bots on the widget, a cheaper model and a daily cap

The second cost came from an embeddable widget that runs on a partner’s website. In the period we looked at, bots made 2,380 generations through it and real users made 84. Bots accounted for about 97% of its generations.

Public pages attract a lot of automated traffic. Imperva’s 2026 Bad Bot Report, a vendor report covering 2025, puts bots at over 53% of all web traffic (announcement). Any public page that starts a paid call on a click or a page load can expect bots to start paid calls too.

The platform made two changes. Widget traffic now goes to a cheaper model, and generations are capped at 6,000 a day. The cap doesn’t separate bots from people. It puts a ceiling on the worst day, so the most a bad day can cost is known in advance.

To reduce the bot share itself, filter before the paid call:

  • Require a challenge token and verify it on the server before generating. Cloudflare’s Turnstile can be embedded on any site without routing traffic through Cloudflare, and its validation docs say the token has to be checked on your server for the challenge to count. Each token is valid for 300 seconds and can be validated once.
  • Generate only after an explicit user action, never on page load.
  • Rate-limit per session and per IP address, with a lower daily allowance for anonymous traffic than for signed-in users.
  • Store a bot signal next to each generation, such as the challenge result or your CDN’s bot score, so you can measure the split later.

Check the provider’s numbers against yours

Once a day, pull the provider’s usage for the previous day and compare it with your ledger, model by model. OpenAI’s Usage API breaks usage down by project and model, its Costs API breaks spend down by project and line item, and both need an admin key (OpenAI cookbook). For Gemini, AI Studio has a usage dashboard, and Cloud Billing can export cost data to BigQuery throughout the day (export).

Compare whole days with a day of slack. Google says Gemini cost details usually reach Cloud Billing within a day, and sometimes take longer (billing FAQ).

If the provider counts more calls than your ledger, some caller is reaching the model without going through the code that writes the ledger. Find it before you trust either number. When your ledger has the higher count, look for failed calls you recorded but weren’t billed for. Google’s billing page says Gemini doesn’t charge for requests that fail with a 400 or 500 error, though they still count against quota.

How these numbers were measured

419,057 is an exact count of the vision calls attributed to the job. The $1,100 is the spend attributed to those calls, rounded, and $0.0026 per call is our division of one by the other. It is an average across whatever mix of images and prompts the job sent.

730 is the value the variable held. The two days come from its name.

2,380 and 84 are counts of generations through the widget, from bot traffic and from real users. They count generations, so they say nothing about how many distinct bots or people were involved.

This post doesn’t give the cost of the bot generations or the dates involved. The provider limits and controls above are as of 8 October 2026, and they change often.

A pre-flight checklist for jobs that call a paid model

  • The selection skips items that already have a result from the current model and prompt version.
  • Durations are typed or carry their unit in the name, and the lookback window has an upper bound.
  • The job logs its resolved window, its selection count and an estimated cost before the first paid call.
  • Runs above a cost threshold need an explicit flag from a person.
  • Every paid call writes a ledger row with job, model, units and cost.
  • Each job has a daily budget enforced in code, with an alert at half of it.
  • Batch jobs run under their own project or workspace, with a lower provider-side limit than user-facing features.
  • Public endpoints verify a bot challenge on the server, generate only on user action and use the cheapest model that does the job.
  • A daily check compares the provider’s usage with the ledger.

Frequently asked questions

Do LLM API rate limits protect you from a large bill?

Only partly. Rate limits cap requests and tokens per minute or per day, so a job that runs steadily below them can still spend a lot. Spend-based limits and hard caps help, but providers note that enforcement lags behind spend.

Does a Google Cloud budget stop spending?

No. Google’s documentation says an alerts-only budget doesn’t cap usage or spending. You can send budget notifications to Pub/Sub and automate a response, such as disabling billing on a project.

Can you set a hard spend limit on the OpenAI API?

Yes. As of 8 October 2026, OpenAI supports hard spend limits for organisations and projects, and requests over the limit get a 429 error. OpenAI notes that enforcement isn’t instantaneous.

How do you make a backfill job idempotent?

Select only items that have no result from the current model version, write the result and the version together, and count the selection before calling any model. A second run then finds nothing to do.

How do you stop bots running up AI costs on an embeddable widget?

Verify a challenge token on the server before each generation, generate only on an explicit user action, serve anonymous traffic with a cheaper model, and cap generations per day so the worst day has a known cost.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.