Key takeaways
- Every forced tool call puts a model request and a tool run in front of the first word.
- If the server or browser already has the context, put it straight into the instructions.
- Force tool_choice only on turns that must end in an action, and leave other turns unforced.
- Reasoning tokens come before the first visible token, so check the effort level each request sends.
- Our assistant now shows first text at about 3.2 s, and the first byte arrives at about 0.8 s.
We build and run an AI study platform for high-school and college students, and it has an AI assistant in a side panel. Its first word used to arrive 20 to 30 seconds after a student sent a message. It now shows first text at about 3.2 seconds, and the first byte of the response arrives at about 0.8 seconds.
Every turn used to chain three slow steps before the model could start answering. A forced tool call fetched context. That tool ran in the browser, so the turn waited on a round trip to read the current page and the student’s selection. Then came a read of the course content. We now put the page, the highlighted passage, the skill the student chose and the course text straight into the instructions. The course text didn’t need a second read, because it was already loaded for the access check. Tools are forced only for real actions, and reasoning effort is set to low.
This post explains why each of those steps costs so much before the first token, what goes into the prompt instead, how we handle tools and reasoning effort now, and what our numbers cover. The stack is a self-hosted OpenAI ChatKit server, the OpenAI Agents SDK and GPT-5.4 served from Azure. The same reasoning applies to any assistant that gathers context before it answers.
What a turn waited on before the first word
ChatKit provides the chat UI in the panel, and it talks to a ChatKit server that we host ourselves. In a self-hosted ChatKit server, a respond() method runs your agent and streams events back to the browser as server-sent events1. Ours runs an agent built with the OpenAI Agents SDK on GPT-5.4, deployed on Azure2. Self-hosting is also the route OpenAI now recommends for new ChatKit work, because Agent Builder is scheduled to shut down on 30 November 20263.
Before the change, a turn ran in this order, and each step waited for the one before it.
- The agent ran with
tool_choiceforced, so the model’s first output had to be the tool call that fetched context. - That tool ran in the browser. The run stopped, the browser read the current page and the selection, and the result went back to the server, which started a new run.
- The course content was read.
- Only then could the model write the first word of its answer.
| Before the first word | Before the change | Now |
|---|---|---|
| Model requests | At least two: one that had to call the context tool, one that wrote the answer | One, unless the model decides to call a tool |
| Round trips to the browser mid-turn | One, to read the page and selection | None |
| Course content | Read during the turn | Already loaded for the access check |
| Forced tool calls | Every turn | Only turns that must end in an action |
| Time to first word | 20 to 30 seconds | About 3.2 seconds to first text, about 0.8 seconds to first byte |
A forced tool call costs a whole model request
With tool_choice set to required, or to a named function, the model has to answer with a tool call4. Your code runs the tool and then makes a second request to the model with the result, so the answer can only start in that second request or a later one4. So a forced call puts a complete model request, plus the tool’s own run time, in front of the first word. With a reasoning model, that first request may also reason before it produces the call.
The Agents SDK resets tool_choice to auto after a tool call by default, so a forced choice can’t trap the model in a loop of calls to the same tool5. The reset does nothing about the cost of the first forced call, which in our assistant came on every turn.
OpenAI’s latency guide makes the same point in general terms. Each request adds its own round-trip latency, and steps that depend on each other are better combined into a single prompt6.
The round trip to the browser
The context tool ran in the browser, where the current page and the student’s selection live. ChatKit supports this with client tools. A client tool is registered in the client options and on the agent. When the agent calls it, the server queues the call and the run stops, using a tool-use behaviour such as StopAtTools. The browser runs its handler and sends the output back, and ChatKit calls the server’s respond() again with that output so a new run can carry on1.
That makes one client tool call expensive. The first run ends, the call travels to the browser and the output travels back, and a second run starts. The second run loads the thread again and sends a fresh request to the model.
The page and the selection already existed in the browser when the student pressed send, so the trip bought nothing the request couldn’t have carried. ChatKit’s client lets you override fetch to add headers or auth to its requests, and the server passes a request context of your own design into respond()1. Between them there is room to send page context with the message.
The voice tutor on the same platform has a similar wait in its newest version, where the voice goes silent until each hand-off to its teaching model is answered. That is covered in What we learned building a voice tutor that draws while it talks.
Putting the context in the instructions
The fix was to stop asking for context the system already had. The current page, the passage the student highlighted, the skill they chose and the course text now go straight into the instructions, and none of them needs a tool call or a second read.
The Agents SDK supports this directly. instructions can be a function that receives the run context and the agent and returns the prompt5. OpenAI’s agent guide recommends a dynamic instructions callback once the guidance depends on the current user or runtime context7. A minimal version, written for this post:
from dataclasses import dataclass
from agents import Agent, ModelSettings, RunContextWrapper
from openai.types.shared import Reasoning
RULES = "You help students with their course. Answer from the material below."
@dataclass
class Turn:
course_text: str # already loaded by the access check
skill: str # instructions for the skill the student chose
page: str # the page the student has open
selection: str # the highlighted passage, or ""
def build_instructions(ctx: RunContextWrapper[Turn], agent: Agent[Turn]) -> str:
t = ctx.context
# Stable text first and per-turn text last, so the prefix can be cached.
return "\n\n".join([
RULES,
f"Course material:\n{t.course_text}",
t.skill,
f"Current page: {t.page}",
f"Highlighted passage: {t.selection or '(none)'}",
])
assistant = Agent[Turn](
name="Assistant",
instructions=build_instructions,
model="gpt-5.4",
tools=action_tools, # defined elsewhere: tools that change something
model_settings=ModelSettings(reasoning=Reasoning(effort="low")),
)
The server builds a Turn before it starts the run and passes it in as the run context7, and the reasoning effort travels in ModelSettings8. With that in place, nothing on the model’s path has to fetch context.
Two things change when context moves into the prompt. The first is length. Input is processed before the first token appears, so a longer prompt means a later first token910. For most prompts the effect is small. OpenAI’s latency guide says cutting a prompt by half may improve latency by only 1 to 5 percent, though it makes an exception for very large contexts such as whole documents6. Course text can be document-sized, so measure with your real content.
The second is caching. OpenAI’s prompt caching reuses a prefix only when it matches exactly, and its guide says to put stable instructions and shared reference material first, with per-request details at the end11. Course text that changes only when the student moves to another course belongs near the top. The page and the selection can change on every turn, so they go last.
Forcing tools only for real actions
Tools are now forced only for real actions, the turns that can’t be complete without the side effect a tool produces. A question about the page gets an answer with nothing forced, because the context is already in the prompt.
When a turn has to end in an action, forcing is still useful. tool_choice can name a function, which makes the model call exactly that function, or be set to required, which makes it call at least one4. The extra request costs the same as before, but on these turns the tool call is the reason for the turn. In the Agents SDK, tool_choice lives in ModelSettings, and clone() gives you a copy of an agent with different settings for the turns that need it5.
Reasoning effort and the first visible token
Reasoning models spend reasoning tokens before they produce a response, and those tokens are billed as output even though you don’t see them12. For time to first word, what matters is that they come first. OpenAI’s reasoning guide frames effort as a trade-off. Lower settings answer sooner and use fewer tokens, while higher settings reason more thoroughly12.
GPT-5.4 accepts none, low, medium, high and xhigh, and OpenAI’s model page lists none as its default13. The guide pitches low as some reasoning for a small latency cost, and none at latency-critical work that gains nothing from reasoning12. Our assistant runs at low.
Defaults depend on the model, and GPT-5.5 defaults to medium, for example12. Check the effort each request sends, whatever you believe the default to be. The response’s usage object reports reasoning tokens separately, so you can see how much reasoning each request did12.
The same guide suggests one more technique for a faster first visible token, which is to ask the model for a short preamble before it reasons further12. We don’t measure that in this post. Lower effort can also cost quality on harder questions, and this post doesn’t include a before-and-after quality comparison, so test it against your own evaluation set.
What the numbers measure, and what they leave out
Time to first word, or first text, is the wait between a student sending a message and the first words of the reply appearing. It is close to what model providers call time to first token, the time from sending a prompt to the model generating its first token10. First byte is when the first bytes of the streamed response arrive.
The current figures are from production, with first text at about 3.2 seconds and first byte at about 0.8 seconds. The old figure, 20 to 30 seconds to the first word, is a range from before the change.
The numbers leave several things out.
- What each change saved. The figures are for the assistant with all three changes in place, so this post doesn’t attribute the saving to any one of them.
- Where the old time went. This post doesn’t break the old 20 to 30 seconds down by step.
- Distribution. We don’t give percentiles or sample sizes here.
- Quality. There is no before-and-after comparison of answer quality.
- The end of the answer. Both numbers are about when the answer starts. Time to the last word isn’t covered.
Most of the remaining wait, about 2.4 of the 3.2 seconds, falls between the first byte and the first text. If you need to go further, that is the stretch to instrument first.
A checklist for a slow assistant
- List what a turn waits on. Write down every step between the message and the first token, and mark which ones wait on another.
- Use what the server already has. Data loaded for authentication or an access check can go straight into the prompt.
- Send what the browser already has. The page and the selection can travel with the message. A client-side tool costs a round trip and a second model run.
- Force tools only for actions. Leave
tool_choiceunforced when the model only needs information. - Check reasoning effort per request. Defaults vary by model, and the usage data shows how many reasoning tokens each request spent.
- Order the prompt for caching. Put stable rules and reference text first and per-turn details last.
- Track first byte and first text separately. Nielsen’s limits are about 1 second for keeping a user’s train of thought and about 10 seconds for keeping their attention14. A first byte under a second leaves room to show that something is happening before the first limit.
- Stream the answer. OpenAI’s latency guide rates streaming as the most effective way to cut the time users spend waiting6.
-
OpenAI, Advanced integrations with ChatKit (self-hosted server,
respond(), client tools), https://developers.openai.com/api/docs/guides/custom-chatkit. ↩↩↩ -
Microsoft Foundry, gpt-5.4 model page (version 2026-03-05), https://ai.azure.com/catalog/models/gpt-5.4. ↩
-
OpenAI, ChatKit guide (Agent Builder deprecation), https://developers.openai.com/api/docs/guides/chatkit. ↩
-
OpenAI, Function calling guide (tool choice and the function calling flow), https://developers.openai.com/api/docs/guides/function-calling. ↩↩↩
-
OpenAI Agents SDK for Python, Agents (dynamic instructions, forcing tool use,
reset_tool_choice), https://openai.github.io/openai-agents-python/agents/. ↩↩↩ -
OpenAI, Latency optimization guide, https://developers.openai.com/api/docs/guides/latency-optimization. ↩↩↩
-
OpenAI, Agent definitions guide (dynamic instructions and run context), https://developers.openai.com/api/docs/guides/agents/define-agents. ↩↩
-
OpenAI Agents SDK for Python, Models (
ModelSettingsand reasoning effort), https://openai.github.io/openai-agents-python/models/. ↩ -
Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving”, OSDI 2024, https://arxiv.org/abs/2401.09670. ↩
-
Anthropic, Reducing latency (definition of time to first token), https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency. ↩↩
-
OpenAI, Prompt caching guide, https://developers.openai.com/api/docs/guides/prompt-caching. ↩
-
OpenAI, Reasoning models guide (reasoning effort, reasoning tokens and preambles), https://developers.openai.com/api/docs/guides/reasoning. ↩↩↩↩↩↩
-
OpenAI, GPT-5.4 model page, https://developers.openai.com/api/docs/models/gpt-5.4. ↩
-
Jakob Nielsen, “Response Times: The 3 Important Limits”, Nielsen Norman Group, adapted from Usability Engineering (1993), https://www.nngroup.com/articles/response-times-3-important-limits/. ↩
Frequently asked questions
What is a good time to first token for an in-app AI assistant?
Nielsen’s response-time limits are a useful guide. About 1 second keeps a user’s train of thought and about 10 seconds is the limit of their attention, so show something within a second and stream the answer. Ours shows first text at about 3.2 seconds.
Does forcing tool_choice slow down a model’s response?
It adds a step. A forced call makes the model’s first output a tool call, and your application then needs a second request with the tool’s result before any answer text appears. Force tools only on turns that must end in an action.
Should an assistant’s context go in the instructions or come from a tool call?
If the context is already available when the message arrives, put it in the instructions. Keep tools for actions and for data the model only sometimes needs.
Does lowering reasoning effort reduce time to first token?
Usually. Reasoning models produce reasoning tokens before the visible answer, and OpenAI describes lower effort as favouring speed. Defaults differ by model, so check what each request sends. OpenAI lists none as GPT-5.4’s default.
What are client tools in OpenAI ChatKit?
Tools that run in the browser. When the agent calls one, the server stops the run, the browser runs the handler and sends the output back, and the server starts a new run with it. That round trip happens before the answer can continue.
Does a longer prompt make the first token slower?
Somewhat, because input is processed before the first token. OpenAI’s latency guide says halving a prompt may cut latency by only 1 to 5 percent, except with very large contexts, and prompt caching helps when the stable part of the prompt comes first.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.