Key takeaways
- A model refusal arrives as a normal response, so classify it before retry logic sees it.
- Stopping at the first safety block cut a failed try-on from about 110 s to about 20 s.
- Cached failures need an expiry. Ours is 30 minutes.
- Dropping alpha without compositing exposes whatever colour sits under transparent pixels, often black.
- Store processing state on the record instead of inferring it from the host that serves a file.
The failures that hurt our AI image pipeline were quiet ones. On a fashion commerce platform we build and run, a safety refusal from Gemini was retried five times as if it were a transient error. A try-on request ran for about 110 to 115 s before it failed, and the failure was then cached with no expiry. In background removal, transparent cutouts were flattened onto black when they were saved, and a check that decided which products still needed processing was looking at the host that served each image.
A refused try-on now stops at the first block and fails in about 20 s. Renders can run as a request followed by polling, a duplicate request inside 4 minutes doesn’t start a second render, and cached failures expire after 30 minutes. On the background side, a repair run fixed 160 of the 161 products the status check had skipped, with 0 failures.
This post covers how Gemini reports a refusal and how to sort responses into retryable and final. It then covers why failed results need an expiry, where black backgrounds come from in Pillow and rembg, and why processing state belongs on the record. The last section is about temperature, which turned out to matter when we regenerated photos of second-hand items.
A try-on request that took nearly two minutes to fail
Try-on takes a photo of a person and a photo of a garment, and renders the garment on the person with Gemini image generation. The first version did this inside one synchronous request. The server called the model, retried on failure up to five times, and answered the caller only when it had an image or had run out of attempts.
When the model refused an input on safety grounds, the code that parsed the response threw away the reason and reported a generic failure. The retry loop had no way to tell that apart from a timeout, so it spent all five retries before giving up. The caller waited about 110 to 115 s for an error.
The cache made it worse. Failures were stored with no expiry, so a render that failed once, for any reason, kept failing from the cache. A transient fault would have become a permanent failure for that combination of photo and garment.
Retries also multiply across layers, so check what your SDK already does before you add a loop of your own. Google’s troubleshooting guide says the official Python SDK retries transient errors up to four times by default, which makes each call up to five attempts. Put that inside a loop of five retries, which is six attempts, and a fault that persists can turn one request into 30 calls. A safety block arrives as a normal response, so the SDK won’t retry it, but a loop of your own will. And a person who gives up after nearly two minutes and taps again starts the whole sequence a second time.
How Gemini reports a refusal
With the generateContent API, a refusal is data inside a normal response. If the input is blocked, promptFeedback.blockReason is set and no candidates come back (API reference). If the output is blocked, the candidate’s finishReason gives the reason, such as SAFETY, PROHIBITED_CONTENT, IMAGE_SAFETY or IMAGE_PROHIBITED_CONTENT, and the blocked content is left out of the response (safety settings). There is also NO_IMAGE, for a response where an image was expected and none was produced.
Google’s API error reference lists the safety, prohibited-content and recitation codes as blocked generations, and its advice for every one of them is to change the input before trying again. The troubleshooting guide says to retry only transient errors, naming 429, 408 and 5xx, and to leave 400, 402 and 403 alone because they point at the request, the key or billing.
That gives a short table to code against.
| What comes back | Where it shows up | Send the same input again? |
|---|---|---|
| 408, 429, 5xx | HTTP status | Yes, with backoff, jitter and a cap. A daily-quota 429 won’t clear until the quota resets |
| 400, 402, 403 | HTTP status | No. Fix the request, key or billing |
| Input blocked | promptFeedback.blockReason is set and there are no candidates |
No. Return a failure now |
| Output blocked | finishReason is SAFETY, PROHIBITED_CONTENT, IMAGE_SAFETY, IMAGE_PROHIBITED_CONTENT, IMAGE_OTHER or a recitation reason |
No. Return a failure now |
| No image | finishReason is NO_IMAGE, or the response has text and no image part |
At most once more, then fail |
The last row is our judgement. The docs don’t say whether NO_IMAGE is worth a second attempt.
In code, the parser should turn each case into its own exception type, so the retry loop only ever sees the transient ones:
import base64
BLOCKED = {"SAFETY", "PROHIBITED_CONTENT", "BLOCKLIST", "SPII", "RECITATION",
"IMAGE_SAFETY", "IMAGE_PROHIBITED_CONTENT", "IMAGE_RECITATION", "IMAGE_OTHER"}
class Refused(Exception):
"""The model declined this input. Sending it again won't help."""
class NoImage(Exception):
"""The model finished without producing an image."""
def image_from(resp: dict) -> bytes:
if reason := resp.get("promptFeedback", {}).get("blockReason"):
raise Refused(f"input blocked: {reason}")
finish = None
for cand in resp.get("candidates", []):
finish = cand.get("finishReason")
if finish in BLOCKED:
raise Refused(f"output blocked: {finish}")
for part in cand.get("content", {}).get("parts", []):
if "inlineData" in part:
return base64.b64decode(part["inlineData"]["data"])
raise NoImage(f"finishReason={finish}")
The retry loop around it catches transport errors and the retryable status codes. Refused goes straight back to the caller with its reason, and NoImage gets one more attempt at most.
Stop at the first block, and stop holding the request open
A refused try-on now stops at the first block and returns in about 20 s, which is the time of the one attempt it makes.
We also added an async mode. The front end submits a render, gets a job ID back straight away and polls for the result. Microsoft documents this shape as the Asynchronous Request-Reply pattern. The first call returns 202 Accepted with a status URL, and the status endpoint reports progress until the result is ready. A render then no longer holds an HTTP connection and a server worker for its whole duration, and the front end can show progress or let the person move on.
The third fix is a guard against duplicates. A second request for the same try-on within 4 minutes doesn’t start a second render. Microsoft’s write-up of the pattern recommends the same idea through an idempotency key: a repeated submit gets the existing status resource back instead of creating a second work item. A sketch with Redis:
import uuid
import redis
r = redis.Redis()
GUARD_SECONDS = 4 * 60
def submit_render(user_id: str, inputs_hash: str) -> str:
key = f"render:{user_id}:{inputs_hash}"
job_id = uuid.uuid4().hex
if r.set(key, job_id, nx=True, ex=GUARD_SECONDS):
enqueue_render(job_id, user_id, inputs_hash)
return job_id
existing = r.get(key)
return existing.decode() if existing else submit_render(user_id, inputs_hash)
inputs_hash should cover everything that changes the output, including both photos and the prompt version. Set the window longer than your slowest successful render, so a duplicate can’t slip in while the first one is still running.
Cached failures need an expiry
Caching a failure is reasonable on its own. It stops a burst of identical requests from each waiting on the same refusal. The mistake was the missing expiry. Failed results now expire after 30 minutes, so a transient failure clears without anyone touching the cache, and a refused input gets checked again later.
DNS has handled the same problem for decades. RFC 2308 made caching of negative answers mandatory for resolvers and bounded it at the same time. It suggests one to three hours as a sensible default and notes that values over a day have caused problems. A failed render needs the same treatment, with a much shorter clock.
Keep the reason with the cached entry. A refusal and a timeout need different messages for the person waiting, and you may want different expiry times for each:
FAILURE_TTL_SECONDS = 30 * 60
def store_result(cache, key: str, result) -> None:
if result.ok:
cache.set(key, result.to_json())
else:
cache.set(key, result.to_json(), ex=FAILURE_TTL_SECONDS)
Put the model name and prompt version in the cache key too. Otherwise a model upgrade keeps serving answers from the old one.
Where the black backgrounds came from
Background removal on the platform runs rembg alongside a second, self-hosted segmentation model. Both give you a cutout with an alpha channel. When the cutouts were saved, the alpha was dropped and the transparent area came out black. One product image was 92.7% black after saving.
Black is the default in more than one place in a typical Python image stack:
- rembg’s default cutout composites the photo onto a canvas created with
Image.new("RGBA", size, 0)(source). Every transparent pixel is stored as(0, 0, 0, 0), which is black with zero alpha. - In Pillow,
convert("RGB")on an RGBA image keeps the stored colour and discards alpha. It doesn’t composite onto anything. We checked this on Pillow 12.2. - Pillow’s
Image.newfills with black when you don’t pass a colour (docs), so pasting a cutout onto a fresh RGB canvas gives a black background. - JPEG has no alpha channel. Before Pillow 4.2.0, saving an RGBA image as JPEG quietly discarded alpha, and since 4.2.0 it raises an error (release notes). The shortest way past that error is
convert("RGB"), which brings the black back.
Soft edges add a subtler version of the same problem. Edge pixels carry partial alpha, and rembg’s default cutout blends their colours toward its black canvas as well, so even a correct composite onto white leaves a faint dark rim. rembg has a putalpha option that keeps the original colours under the mask. With that option, a dropped alpha channel brings back the photo’s original background instead of black, which is just as wrong and harder to spot in a grid of thumbnails.
The fix is to choose the background on purpose and composite onto it before anything can drop alpha. The alternative is to keep PNG or WebP with alpha all the way to the browser.
from PIL import Image
def flatten(img: Image.Image, background=(255, 255, 255)) -> Image.Image:
"""Composite onto a solid colour, then drop alpha."""
rgba = img.convert("RGBA")
canvas = Image.new("RGBA", rgba.size, background + (255,))
return Image.alpha_composite(canvas, rgba).convert("RGB")
def dark_share(img: Image.Image, threshold: int = 16) -> float:
"""Fraction of pixels whose luma is below the threshold."""
hist = img.convert("L").histogram()
return sum(hist[:threshold]) / (img.width * img.height)
Run dark_share on the output and on the source photo when you write the file, and flag outputs that come out much darker than their input. Some products really are black, so compare against the source instead of using a fixed limit. A test that pushes a fully transparent image through the save path and checks that the corners come out white catches this whole class of bug.
The status check that trusted the URL
A separate check decided whether a product’s image had already been processed by looking at which host served it. An image on the host where processed images live counted as done. That assumption skipped 161 products that still needed work. The repair run that followed fixed 160 of them with 0 failures.
A file’s host only tells you where it was written, and a broken cutout can sit on the same host as a good one. Processing state belongs on the record, written by the step that checked the output. Something like this is enough:
| Field | Example value | What it’s for |
|---|---|---|
bg_model |
segmenter-2 |
Re-running products when you change models |
bg_version |
3 |
Bumped whenever the pipeline changes |
bg_checked_at |
a timestamp | When the output passed its checks |
bg_dark_share |
0.04 |
The pixel check, kept for later audits |
Select work with a query on those fields, such as every product whose bg_version is below the current one. A repair run then picks up the products that need it, and a second run finds nothing to do.
Low temperature for regenerated product shots
Second-hand items are photographed by the people selling them. For those, the platform regenerates a clean studio-style shot with Gemini. Early outputs sometimes rotated the item or changed the colour of metal hardware such as buckles and zips. Setting the temperature to 0.2 stopped both.
Google’s prompting strategies guide describes temperature as the degree of randomness in token selection, with lower values suited to tasks that need a more deterministic response. Regenerating a product photo is that kind of task. A second-hand item is a single physical object, and the buyer relies on the photo, so a buckle that changes colour misdescribes what they will receive.
Check this against your own model before you copy it. The same Google guide recommends keeping sampling parameters at their defaults for Gemini 3.x models, and warns that lowering temperature can cause looping or degraded output, particularly on maths and reasoning tasks. Our result is for one editing task on the model we ran. If you change models, run a fixed set of difficult seller photos through both settings and compare them side by side before you switch.
The prompt matters too. List the properties that must not change, such as orientation, colours, hardware and visible wear. Google’s own editing examples ask the model to keep everything else in the image exactly as it was, and that instruction is worth copying.
How we measured these numbers
The timings come from our build notes. They are approximate wall-clock times for single requests, about 110 to 115 s for the old loop and about 20 s for a refused request now. They are not percentiles from a load test.
The 92.7% figure is the share of black pixels in one saved product image. It is one image, and we are not claiming every affected image was that dark.
161 is the number of products the host-based check skipped. 160 is the number the repair run fixed, with 0 failures reported. Both are exact counts from that work.
The temperature change is a qualitative result from reviewing regenerated images. Rotation and hardware recolouring stopped appearing after the change, and we don’t have a before-and-after rate to publish.
This post doesn’t cover how often try-on inputs trigger the safety filter, or what the wasted retries cost. For a case where wasted calls did add up, see how a backfill bug ran up an LLM vision bill.
A checklist for image pipelines
- Parse every model response into an image, a refusal, an empty result or a transient failure before any retry logic sees it.
- Retry only transient failures, with backoff, jitter and a cap on total time as well as attempts. Find out what your SDK already retries.
- Return a refusal on the first attempt, with its reason attached.
- Run anything that takes more than a few seconds as submit-then-poll.
- Guard duplicate submits with a key over the inputs and a window longer than your slowest successful render.
- Give cached failures an expiry, and store the failure reason with them.
- Composite onto a chosen background before anything drops alpha, and test the save path with a fully transparent image.
- Compare output brightness with the input when you write the file.
- Keep processing state and model version on the record, and select work by state.
- Pin sampling settings for editing tasks, and test them again when the model changes.
Frequently asked questions
Should you retry when Gemini blocks an image for safety reasons?
No. A block comes back as a response with a block reason or a finish reason such as IMAGE_SAFETY, and no image. Google’s error guide says to change the input, so sending the same request again only adds delay.
Why do transparent PNG cutouts turn black?
rembg’s default cutout stores transparent pixels as black with zero alpha. Anything that drops the alpha channel without compositing, such as a plain conversion to RGB in Pillow, exposes that black. Composite onto the background colour you want first.
How long should you cache a failed AI generation?
Long enough to absorb repeated requests, then let it expire. We use 30 minutes for failed try-on renders. DNS resolvers apply the same idea to negative answers, and RFC 2308 suggests one to three hours there.
What temperature should you use for Gemini image edits?
For regenerating product photos, 0.2 stopped the model rotating items and recolouring metal hardware in our pipeline. Google recommends default sampling settings for Gemini 3.x models, so test any change on the model you run.
Should long AI image generations run in a synchronous request?
Avoid it for anything that takes more than a few seconds. Return a job ID at once, let the front end poll, and guard against duplicate submits so a second tap doesn’t start a second render.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.