Key takeaways
- A model with a search tool decides for itself whether to search, so check every answer’s grounding metadata.
- A swallowed quota error turned a research feature into confident answers with no research behind them.
- On an older aiohttp 3.10, the google-genai SDK turned every 429 and 503 into an AttributeError.
- Answering from the course, 8 of 8 fact questions came back right in 3.5 to 6.5 s.
- The old, stricter course prompt refused 6 of 24 on-topic questions.
Giving a model a search tool doesn’t mean it will search. On an AI study platform we build and run, a mode that was supposed to answer with web sources never searched at all. A separate deep research path called a provider whose API was returning 401 with an insufficient_quota error, and the code swallowed the error. Students got fluent answers with no research behind them, and nothing on screen told them so.
The assistant now uses the Gemini API’s Google Search grounding tool. It passes Google’s source links through unchanged and shows Google’s Search Suggestions with each grounded answer, as Google’s terms require. On the way we hit two more problems. The google-genai Python SDK, running on aiohttp 3.10, turned every 429 and 503 into an AttributeError. Gemini Flash-Lite wrote “research” from its own memory 1 time in 4 until the prompt required a search result behind every point.
For questions about the course itself, the whole course goes into the context up to 60k tokens, with BM25 search beyond that. That mode answered 8 of 8 specific-fact questions correctly in 3.5 to 6.5 s, at about $0.018 a question. Below are the failures and their fixes, what Google’s terms ask of the interface, the course design and its numbers, how we measured, and a checklist.
Two failures that looked like normal answers
The first failure was the web-sources mode. It was meant to search and cite, and it never searched. The second was the deep research path. Its provider answered with HTTP 401 and an insufficient_quota error, our code caught the exception and carried on, and the student got an answer written without any research.
To a user, both looked like a working feature. Each produced a confident, well-formatted answer. A test that only checks for a non-empty reply passes both, and an error dashboard stays quiet because the error never reaches it. We wrote about a related failure, an LLM pipeline that reported 0 dropped while most of its output was missing.
These checks catch both kinds of failure early:
- Treat a research answer with no sources as a failure. Retry it, or tell the user the search didn’t run. Don’t present it as research.
- Sort provider errors by whether waiting helps. A 429 or 503 can clear on retry. A 401 or a quota error won’t, so alert on the first one.
- Track, for each mode, the share of answers that came back with sources. A mode that promises sources and shows none has a bug, whatever the error rate says.
How grounding with Google Search works in the Gemini API
You switch grounding on by adding the google_search tool to a request. The model then decides, request by request, whether a search would improve the answer, and writes its own queries if it thinks so.1 An attached tool tells you nothing about whether a given answer used it.
When the model does search, the response from generateContent carries a groundingMetadata object on the candidate.2 These are the parts that matter for a study assistant:
webSearchQueries, the queries the model ran.groundingChunks, the sources, each with a weburiandtitle.groundingSupports, which link spans of the answer to sources. Each has asegmentwith start and end offsets and a list of indices intogroundingChunks.searchEntryPoint.renderedContent, the HTML and CSS for Google’s Search Suggestions.
Google’s newer Interactions API reports the same information as google_search_call and google_search_result steps, with url_citation annotations on the answer text.1 With the Python SDK and generateContent, checking that a search happened takes a few lines:
from google import genai
from google.genai import types
client = genai.Client()
class NotGrounded(Exception):
pass
def grounded_answer(question: str, model: str):
resp = client.models.generate_content(
model=model,
contents=question,
config=types.GenerateContentConfig(
tools=[types.Tool(google_search=types.GoogleSearch())],
),
)
meta = resp.candidates[0].grounding_metadata
chunks = (meta.grounding_chunks or []) if meta else []
sources = [c.web for c in chunks if c.web]
if not meta or not meta.web_search_queries or not sources:
raise NotGrounded("the model answered without searching")
return resp.text, sources, meta
Grounding is billed on top of tokens, and the unit differs between model generations. On Gemini 3 models Google bills each search query the model runs, so one question that triggers three searches counts three times. On Gemini 2.5 and older it bills per grounded prompt.1
| Models | Included | After that |
|---|---|---|
| Gemini 3.x, paid tier | 5,000 searches a month, shared across all 3.x models | $14 per 1,000 searches |
| Gemini 2.5 Flash and Flash-Lite, paid tier | 1,500 grounded prompts a day, shared between the two | $35 per 1,000 grounded prompts |
Prices are from Google’s pricing page, as of 8 October 2026.3
What Google’s terms ask of the interface
The Gemini API terms have a section on Grounding with Google Search, and several of its rules govern the interface.4 Paraphrased, they are these:
- Show the grounded answer together with its Search Suggestions, to the user who asked the question.
- Don’t modify grounded results or Search Suggestions, and don’t mix other content into them.
- Don’t put an interstitial page between a link or suggestion and its destination, and don’t redirect users away from the destination.
- Don’t add click tracking or link tracking to grounded results or suggestions.
- Store grounded results only in the limited cases the terms allow. Chat history is one of them, for up to two years.
The assistant passes the source links through exactly as Google returns them, with no tracking wrapper and no interstitial, and shows the Search Suggestions with each grounded answer. The API reference describes renderedContent as a snippet to embed in a web page or an app webview.2 It is Google’s own markup and styling, so read the terms’ display requirements before you decide how to embed it, and check on a phone that tapping a suggestion opens Google Search.
Inline citations need one more piece of care. The API reference defines a segment’s start and end as byte offsets within the part.2 Python slices strings by code point, so once any non-ASCII text appears before a segment, slicing the string directly puts its citation marker in the wrong place. Convert first:
def segment_text(answer: str, segment) -> str:
"""Return the text a grounding support covers. Offsets are UTF-8 bytes."""
data = answer.encode("utf-8")
return data[segment.start_index or 0 : segment.end_index].decode("utf-8")
How an old aiohttp turned every 429 and 503 into an AttributeError
Our async calls to Gemini failed in a way that hid the real error. Every 429 and every 503 from the API reached our code as an AttributeError. The environment had aiohttp 3.10 installed, and two details of the google-genai SDK combine to produce this.
First, the SDK picks its async transport at runtime. Its docs describe httpx as the default and aiohttp as an optional extra for speed.5 In the code, the switch to aiohttp happens whenever aiohttp can be imported, whoever installed it, unless you pass your own httpx async client or transport.6 A project can end up on aiohttp because some other dependency brought it in.
Second, on the non-streaming async path, the SDK raises its APIError for a bad status inside a try block. That block’s except clause lists aiohttp.ClientConnectorDNSError,7 a class aiohttp added in 3.10.10.8 Python only evaluates the exception classes in an except clause when something is raised. A 200 raises nothing, so ordinary calls work. A 429 or 503 raises APIError inside the try, Python looks up the missing attribute, and an AttributeError replaces the APIError. A public issue against the SDK describes this failure.9 The SDK’s optional aiohttp extra asks for 3.10.11 or later, but that constraint only applies when you install the extra. An older aiohttp pulled in by another package still gets picked up.
The cost shows up in retries. The SDK’s own retry options retry an APIError whose status is in a list that includes 429 and 503.10 An AttributeError matches nothing in that list, so neither the SDK’s retries nor any handler keyed on status codes runs. A brief rate limit becomes a failed request.
Switching the SDK’s async transport to httpx fixed it. In current releases, one way to do that is to hand the SDK your own httpx client:
import httpx
from google import genai
from google.genai import types
client = genai.Client(
http_options=types.HttpOptions(httpx_async_client=httpx.AsyncClient()),
)
Upgrading aiohttp to 3.10.11 or later also removes the AttributeError. Whichever you choose, add a test that makes the API answer 429 and 503 and asserts that your code receives an APIError carrying that status. Run it against the transport you actually ship, since a dependency upgrade can change which one the SDK picks.
Making Flash-Lite search before it writes
With the search tool attached, Gemini Flash-Lite still wrote “research” from its own memory in 1 answer out of 4. Those answers were presented as research and had no search queries or sources behind them. The documented design allows this, because the model decides whether a search would help.1 It kept happening until the prompt required a search result behind every point.
An instruction that asks for this, written fresh for this post:
Search before you write. Every point in your answer must rest on a search
result from this turn, and must cite it. If the results don't cover part
of the question, say that part isn't covered. Don't fill gaps from memory.
A model can still ignore a prompt, and the metadata lets you check. groundingSupports maps spans of the answer to sources, so a few lines of code can find sentences with no support, then drop them or ask again. We’d put that check behind every research mode.
Answering from the course: everything in context up to 60k tokens
The second kind of grounding uses the course itself. When a course fits in 60k tokens, the whole course goes into the context with the question. Above that, BM25 keyword search picks the passages that go in.
Sending everything suits a study assistant. No retrieval step can miss the passage that holds the answer, and a request to summarise the whole course needs the whole course anyway. Anthropic suggests that a knowledge base under 200,000 tokens can go into the prompt whole, with no retrieval step.11 A 2024 study that compared long-context models with retrieval found long context ahead on average when resources allowed, and retrieval much cheaper.12
A cap still matters, for two reasons. Input tokens are billed on every question, so cost grows with course length. Models also read long inputs less reliably. Accuracy drops when the relevant passage sits in the middle of a long context,13 and Chroma found performance shifting with input length across 18 models, even on simple tasks.14 Google’s long-context guide says its models are less accurate when asked for many pieces of information at once, and recommends putting the question after the material.15
Prompt length also affects how soon the first word of a reply appears. A separate post covers that side, for an assistant whose context moved into the prompt.
Above the cap we use BM25, the standard ranking function from the probabilistic relevance framework.16 Course text is dense with exact terms such as names and formulas, and keyword ranking matches those well. BM25 also held up as a strong zero-shot baseline against neural retrievers on the BEIR benchmark.17 It needs no embedding pipeline, and some databases ship it. SQLite’s FTS5 extension, for example, has a bm25() ranking function.18
| Test | Result |
|---|---|
| Questions asking for one specific fact in the course | 8 of 8 correct, each in 3.5 to 6.5 s |
| Requests to summarise the whole course | 15 to 17.5 s |
| Cost per question | About $0.018 |
| On-topic questions under the old, stricter prompt | 6 of 24 refused |
Why the stricter prompt refused a quarter of on-topic questions
An earlier version of the course prompt was strict about staying inside the course. On 24 questions that were all on-topic, it refused 6.
A rule like that asks the model to judge, question by question, whether the answer is “in” the material. That is easy for a definition. It is harder for a question that asks why something works, compares two ideas or applies a method to a new example. The course supports those answers without stating them word for word, and a strict rule gives the model room to refuse.
Two practices follow. Test refusals with questions that should be answered, as well as answers to questions that should be refused, because a test that only checks the second will pass a prompt that refuses too much. And write the instruction so the course comes first, the model says when it goes beyond the course, and it declines only requests that have nothing to do with the course.
How we measured
- Fact questions. 8 questions, each asking for one specific fact from a course. All 8 answers were correct, and each took 3.5 to 6.5 s.
- Summaries. Requests to summarise a whole course took 15 to 17.5 s.
- Cost. About $0.018 per question.
- Refusals. 24 on-topic questions against the old, stricter prompt. It refused 6.
- Research from memory. Before the prompt change, Flash-Lite wrote research from memory in 1 answer out of 4.
The latency figures are the ranges we saw, so there are no percentiles here. This post has no quality score for the summaries and no separate figures for courses over 60k tokens, where BM25 chooses the passages. It also has no accuracy score for web-grounded answers. Eight questions is a small sample, so treat 8 of 8 as evidence that the design works for fact lookups.
What we’d do again
- Check the grounding metadata on every web answer. No search queries means no research, however the answer reads.
- Alert on the first quota or authentication error from any provider. Waiting won’t fix it.
- Pin the HTTP transport your SDK uses, and test the error path with a fake 429 and 503.
- Pass Google’s links through untouched, show the Search Suggestions, and tap one on a phone.
- Ask for a source behind every point in the prompt, then check the supports in code.
- Convert byte offsets before you place inline citations.
- Put the whole document in the context when it fits, cap it, and use BM25 above the cap.
- Count refusals of on-topic questions as failures in your tests.
-
Google, “Grounding with Google Search”, Gemini API documentation, https://ai.google.dev/gemini-api/docs/google-search ↩↩↩↩
-
Google, “GroundingMetadata”, “SearchEntryPoint” and “Segment”, Gemini API reference, https://ai.google.dev/api/generate-content ↩↩↩
-
Google, “Gemini Developer API pricing”, https://ai.google.dev/gemini-api/docs/pricing ↩
-
Google, “Gemini API Additional Terms of Service: Grounding with Google Search”, https://ai.google.dev/gemini-api/terms#grounding-with-google-search ↩
-
Google Gen AI Python SDK documentation, “Faster async client option: Aiohttp”, https://googleapis.github.io/python-genai/ ↩
-
googleapis/python-genai v2.29.0,
_use_aiohttp, https://github.com/googleapis/python-genai/blob/v2.29.0/google/genai/_api_client.py#L1287-L1295 ↩ -
googleapis/python-genai v2.29.0, async request path, https://github.com/googleapis/python-genai/blob/v2.29.0/google/genai/_api_client.py#L1655-L1678 ↩
-
aiohttp changelog, release 3.10.10 (10 October 2024), https://docs.aiohttp.org/en/stable/changes.html ↩
-
googleapis/python-genai issue #2016, “Missing aiohttp version constraint causes AttributeError for ClientConnectorDNSError”, https://github.com/googleapis/python-genai/issues/2016 ↩
-
googleapis/python-genai v2.29.0, retry defaults, https://github.com/googleapis/python-genai/blob/v2.29.0/google/genai/_api_client.py#L545-L581 ↩
-
Anthropic, “Introducing Contextual Retrieval”, 19 September 2024, https://www.anthropic.com/news/contextual-retrieval ↩
-
Li et al., “Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach”, EMNLP 2024 Industry Track, https://arxiv.org/abs/2407.16833 ↩
-
Liu et al., “Lost in the Middle: How Language Models Use Long Contexts”, Transactions of the ACL, https://arxiv.org/abs/2307.03172 ↩
-
Hong, Troynikov and Huber, “Context Rot: How Increasing Input Tokens Impacts LLM Performance”, Chroma, 14 July 2025, https://www.trychroma.com/research/context-rot ↩
-
Google, “Long context”, Gemini API documentation, https://ai.google.dev/gemini-api/docs/long-context ↩
-
Robertson and Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond”, Foundations and Trends in Information Retrieval, 2009, https://doi.org/10.1561/1500000019 ↩
-
Thakur et al., “BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models”, NeurIPS 2021 Datasets and Benchmarks Track, https://arxiv.org/abs/2104.08663 ↩
-
SQLite, “SQLite FTS5 Extension”, section on the bm25() function, https://www.sqlite.org/fts5.html ↩
Frequently asked questions
How can I tell whether Gemini actually searched Google?
Read the grounding metadata on the response. If it lists no web search queries and no grounding chunks, the answer came from the model’s own knowledge. The model decides per request whether to search, so check every response.
What do Google’s terms require when you show grounded answers?
Show the Search Suggestions with the grounded answer, don’t modify either or mix other content into them, and don’t put interstitials, redirects or click tracking between users and the linked pages. The Gemini API terms have the full rules.
Why does the google-genai SDK raise AttributeError on a 429?
If aiohttp is installed, the SDK uses it for async calls. Its error handling refers to aiohttp.ClientConnectorDNSError, which aiohttp releases before 3.10.10 don’t have, so HTTP errors surface as AttributeError. Upgrade aiohttp or give the SDK an httpx async client.
Should I use retrieval or put the whole document in the context window?
If it fits comfortably, put all of it in. We send a whole course up to 60k tokens and use BM25 search above that. The cap bounds what each question costs.
What does it cost to answer from a whole course in context?
About $0.018 per question in our measurements, with courses of up to 60k tokens sent whole.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.