Get in touch

9io.ai / Blog

What we learned building a voice tutor that draws while it talks

What broke in three versions of a voice tutor that draws as it talks, built on OpenAI’s Realtime API and GPT-Live, and the API details behind each problem.

Key takeaways

  • Send every response.create through one guarded function, because only one response can write to a conversation at a time.
  • response.done arrived three to four times sooner than our audio finished, so turns end on output_audio_buffer.stopped.
  • A per-lesson cap of 10 visuals also capped lesson length, and our lessons ended after about 3 minutes.
  • In our GPT-Live build, the voice goes silent until each hand-off to the teaching model is answered.
  • OpenAI says GPT-Live is trained to paraphrase text it is given to speak, so test any scripted lines.

We build and run an AI study platform for high-school and college students, and part of it is a voice tutor. Students talk to it, and it answers out loud while it draws on a whiteboard. We have built it three times on real-time speech-to-speech models. Most of what went wrong had to do with timing.

Our code could ask for a second response while one was still running. It decided the tutor had finished speaking while the student could still hear it. The tutor drew 5 visuals in 15 seconds, and the cap we added then ended lessons after about 3 minutes. In the current version, the voice goes silent until each hand-off to the teaching model behind it is answered.

The most useful thing we learned is that the model finishing a reply and the student hearing the end of it are separate events. On OpenAI’s Realtime API they arrive as different server events, and in our tutor the first came three to four times sooner than the second. This post goes through each problem with the numbers from our build notes and the documentation that explains it, and ends with a checklist for anyone starting a voice tutor.

Three versions on two kinds of voice model

Versions 1 and 2 ran on the OpenAI Realtime API over WebRTC. In a Realtime session a single model listens, reasons, calls tools and speaks1. The browser sends microphone audio and plays the model’s voice on WebRTC media tracks, and the JSON events that drive the session travel over a data channel2. The student’s speech was transcribed by whisper-1, and server-side voice activity detection (VAD) decided when the student had finished a turn3.

That transcript comes from a separate model. OpenAI’s reference says to treat it as guidance about the audio, because it isn’t exactly what the speech model heard4. Keep that in mind if you show transcripts to students or store them as a record of the lesson.

An audio-only mode ran on Google’s Gemini Live API, which streams audio over a WebSocket5.

Version 3 runs on GPT-Live 1 (gpt-live-1), a newer OpenAI voice model that listens while it speaks and hands reasoning and tool use to a backend. OpenAI calls the hand-off delegation. GPT-Live has its own API, and its model page lists the Realtime endpoint as unsupported, so code written for one doesn’t move to the other unchanged6. Behind it is a server-side teaching model. It answers each hand-off with a validated JSON scene that describes what to draw, plus images generated with gpt-image-27.

These are the problems we hit, with the numbers from our notes.

Problem What our notes recorded
Overlapping responses Six places in the code sent response.create with no guard
Turns ending early response.done arrived three to four times sooner than the audio finished playing
Drawing too fast The model chained 5 visuals in 15 seconds
Lessons too short A cap of 10 visuals per lesson ended lessons after about 3 minutes
Silence during hand-offs In version 3, the voice goes silent until each hand-off is answered
Duplicate answers In one case in version 3, a single question was answered four times

Only one response can write to the conversation at a time

A Realtime session keeps a default conversation, and the model’s replies are written into it. The reference for response.create says only one response can write to the default conversation at a time. Responses created out of band, with conversation set to none, don’t write to it and can run in parallel4.

A voice tutor has more than one source of responses. With server VAD, the session can start a response by itself when the student stops talking, and the create_response setting controls that4. Your own code is the other source. In ours, six different places called response.create directly, and none of them checked whether a response was already running.

When two requests overlap, the API refuses the second. Developers on OpenAI’s forum report the error as “Conversation already has an active response”8.

The fix was one guarded path. Every part of the tutor that wants the model to speak asks one function, and only that function sends response.create. A minimal version of the pattern, written for this post, looks like this:

// One owner for response.create (a sketch written for this post).
class ResponseGate {
  private active = false;
  private wanted = false;

  constructor(private send: (event: object) => void) {}

  // Everything that wants the tutor to speak calls this.
  request() {
    if (this.active) {
      this.wanted = true; // ask again when the current response is done
      return;
    }
    this.active = true;
    this.send({ type: "response.create" });
  }

  // Feed it every server event from the data channel.
  onServerEvent(event: { type: string }) {
    if (event.type === "response.created") {
      this.active = true; // also catches responses that server VAD starts
    } else if (event.type === "response.done") {
      this.active = false;
      if (this.wanted) {
        this.wanted = false;
        this.request();
      }
    }
  }
}

The gate watches response.created as well as its own requests, because it has to see the responses VAD starts. A production version also needs to clear its flag on errors, and to decide for each caller whether a request that arrives mid-response should wait or be dropped.

If you need a second model call while the tutor is speaking, for example to classify what the student just said, make it an out-of-band response so it stays out of the conversation. The gate should then ignore that response’s events, and the reference suggests the metadata field for telling simultaneous responses apart4.

The audio finishes long after response.done

response.done is sent when a response has finished streaming, whatever its final state9. For a spoken reply, that marks the end of generation and says nothing about playback. The server produces audio faster than real time, as the reference notes in its description of conversation.item.truncate4. Over WebRTC, the server manages the output audio buffer itself, so it knows how much of the reply has been played at any moment10.

In our tutor, response.done arrived three to four times sooner than the audio finished playing. With round numbers as an illustration, a reply that takes 12 seconds to say would report done after 3 or 4 seconds. We had been ending the tutor’s turn on response.done, so everything tied to the end of a turn happened while the tutor was still talking.

Turns now end on output_audio_buffer.stopped. The server sends it, over WebRTC and SIP only, once the output buffer has drained and no more audio is coming9. These are the server events that matter for turn-taking.

Server event What the reference says Use it for
response.created A new response has been created Marking a response as in flight, including ones VAD starts
response.done The response has finished streaming, whatever its final state Releasing the one-response guard and reading the final status
output_audio_buffer.started WebRTC/SIP only. The server has started streaming audio to the client Showing that the tutor is speaking
output_audio_buffer.stopped WebRTC/SIP only. The buffer has drained and no more audio is coming Ending the tutor’s turn
output_audio_buffer.cleared WebRTC/SIP only. The buffer was cleared, after a VAD interruption or a client output_audio_buffer.clear Ending a turn the student interrupted

Over a WebSocket there is no output_audio_buffer.stopped, because the client plays the audio. You track your own playback, and when the student interrupts you stop playback and send conversation.item.truncate so the server’s record matches what was heard10.

GPT-Live over a WebSocket has the same gap. Its session guide says the audio deltas carry no timing fields and there is no output-audio-done event, so tracking playback is the application’s job there too11.

Five visuals in 15 seconds, then lessons that ended at 3 minutes

The model chained 5 visuals in 15 seconds, which is a new visual every 3 seconds. We added a cap of 10 visuals per lesson. The cap then ended lessons after about 3 minutes, and teachers told us lessons ended too early.

A cap on the total has two problems here. Because reaching it ended the lesson, it set how long a lesson could run as well as how many visuals it could show. And a budget of 10 still allows 5 in 15 seconds, so a total on its own doesn’t stop bursts.

Pacing and lesson length need separate limits.

  • Pacing. Set a minimum gap between visuals, or a rule that a new visual waits until the tutor’s current turn has ended. That rule only works if you know when the audio has finished playing, which is what the previous section is about.
  • Length. Let the lesson plan, the teacher or the student decide when a lesson ends. Running out of a visual budget is a poor reason to stop teaching.
  • Behaviour at a limit. Decide in advance what the tutor does when a limit is reached. Carrying on with what is already on the board keeps the lesson going.

Version 3 splits the voice from the teaching

GPT-Live handles the spoken conversation and delegates the rest. With Responses delegation, it calls a Responses model you configure and brings the result back into the conversation. Client delegation hands each request to your application instead, which runs its own agent or workflow and sends the result back12.

In our build, a server-side teaching model answers every hand-off with a scene. A scene is a JSON document that says what to draw. Because it is data, it is validated before anything is drawn. The images that go with it come from gpt-image-2. The example below shows the general shape. It is simplified and made up for this post.

{
  "elements": [
    { "id": "a", "type": "box", "at": [40, 80], "label": "Input" },
    { "id": "b", "type": "box", "at": [260, 80], "label": "Output" },
    { "id": "ab", "type": "arrow", "from": "a", "to": "b", "label": "f(x)" }
  ]
}

Two points in OpenAI’s delegation guide matter for a tutor. The first is that speech and delegated work continue independently, and a finished backend response doesn’t mean the user has heard the answer12. That is the response.done problem again, in a different API. The second is about wording. The guide says GPT-Live is trained to paraphrase text you send it to say aloud12. If your tutor has lines it must say exactly, test whether it says them as written.

Silence during hand-offs, and one question answered four times

In version 3, the voice goes silent until each hand-off to the teaching model is answered. For the student, that is dead air in the middle of a lesson.

OpenAI’s guides describe the voice and the backend running side by side, as the previous section noted. The prompting guide says you can prompt GPT-Live to acknowledge a request while delegated work runs, and suggests telling it not to guess the result while it waits13. Our build doesn’t behave that way today. If you are starting out, decide early what the voice should say while it waits, and how long a wait can run before the student needs to hear something.

In one case, a single question was answered four times. The delegation guide asks applications to track each action with its own operation ID, alongside the delegation ID, and to check whether an action already happened before retrying a failed tool call12. The same bookkeeping is worth applying to answers. Something in your system has to know that a question has already been answered, so that a second answer can be dropped before the student hears it.

The text assistant on the same platform had a wait of its own, with every turn chaining three slow steps before the first word. We cut that from 20 to 30 seconds to about 3 seconds, and wrote it up in How we cut an AI assistant’s time to first word to 3 seconds.

How these numbers were gathered

All the numbers in this post come from our build notes for this tutor. They describe our prompts, lessons and set-up, so read them as observations of one system.

  • Six call sites is a count from the code.
  • Three to four times compares when response.done arrived with when the audio finished playing. Our notes give it as a range.
  • 5 visuals in 15 seconds is a burst we observed.
  • About 3 minutes is how long lessons ran with the 10-visual cap.
  • Four answers to one question happened in one case.

We have not measured whether the tutor speaks its cue lines word for word. Since GPT-Live is trained to paraphrase, we make no claim either way. This post also gives no figures for how long the hand-off silences last, and it doesn’t compare answer quality between versions.

A checklist before you build a voice tutor

  • Give response.create one owner. It should know when a response is in flight, including the ones server VAD starts.
  • Decide which “done” each piece of code needs. response.done frees the conversation for the next response. output_audio_buffer.stopped means the audio has run out, over WebRTC and SIP. Over a WebSocket, track playback yourself.
  • Treat the input transcript as a guide. It comes from a separate model, and OpenAI says it isn’t exactly what the speech model heard.
  • Limit drawing speed and lesson length separately. Test what happens when each limit is reached.
  • Plan the wait in a split design. If the voice model delegates, decide what it says while the backend works, and track each hand-off so a second answer to the same question can be caught.
  • Check session limits against lesson length. OpenAI documents a 60-minute maximum for a Realtime session10. Gemini Live limits audio-only sessions to 15 minutes without context window compression, and a connection lasts about 10 minutes, so a long lesson needs both compression and session resumption14.
  • Measure scripted speech. If the tutor must say particular words, check whether it does.

  1. OpenAI, Voice agents guide, https://developers.openai.com/api/docs/guides/voice-agents. ↩

  2. OpenAI, Realtime API with WebRTC, https://developers.openai.com/api/docs/guides/realtime-webrtc. MDN, RTCDataChannel, https://developer.mozilla.org/en-US/docs/Web/API/RTCDataChannel. ↩

  3. OpenAI, Voice activity detection in the Realtime API, https://developers.openai.com/api/docs/guides/realtime-vad. ↩

  4. OpenAI API reference, Realtime client events (response.create, conversation.item.truncate, and the session’s transcription and turn detection fields), https://developers.openai.com/api/reference/resources/realtime/client-events. ↩↩↩↩↩

  5. Google, Gemini Live API overview, https://ai.google.dev/gemini-api/docs/live. ↩

  6. OpenAI, GPT-Live 1 model page, https://developers.openai.com/api/docs/models/gpt-live-1. OpenAI, Getting started with GPT-Live, https://developers.openai.com/api/docs/guides/live. ↩

  7. OpenAI, gpt-image-2 model page, https://developers.openai.com/api/docs/models/gpt-image-2. ↩

  8. OpenAI Developer Community, “Conversation already has an active response” thread, https://community.openai.com/t/realtime-api-server-response-error-message-conversation-already-has-an-active-response/1005582. ↩

  9. OpenAI API reference, Realtime server events (response.created, response.done and the output_audio_buffer events), https://developers.openai.com/api/reference/resources/realtime/server-events. ↩↩

  10. OpenAI, Realtime conversations guide (interruption, truncation and session length), https://developers.openai.com/api/docs/guides/realtime-conversations. ↩↩↩

  11. OpenAI, Managing GPT-Live sessions, https://developers.openai.com/api/docs/guides/live-conversations. ↩

  12. OpenAI, GPT-Live delegation guide, https://developers.openai.com/api/docs/guides/live-delegation. ↩↩↩↩

  13. OpenAI, GPT-Live prompting guide, https://developers.openai.com/api/docs/guides/live-prompting. ↩

  14. Google, Gemini Live API session management, https://ai.google.dev/gemini-api/docs/live-session. ↩

Frequently asked questions

Why does the OpenAI Realtime API say ‘Conversation already has an active response’?

Only one response can write to the default conversation at a time, so a second response.create sent while one is running is refused. Send every request through one guarded function, and remember that server-side VAD can start responses on its own.

How do I know when the Realtime API has finished speaking over WebRTC?

Listen for output_audio_buffer.stopped, which the server sends when its output audio buffer has drained and no more audio is coming. response.done only means generation has finished, and the audio is produced faster than real time. Over a WebSocket you track playback yourself.

What is the difference between GPT-Live and the OpenAI Realtime API?

GPT-Live 1 is a full-duplex voice model on its own API. It handles the conversation and delegates reasoning and tool use to a backend model or to your own agent. A Realtime session runs speech, reasoning and tools in one model.

Does GPT-Live read text word for word?

OpenAI’s delegation guide says GPT-Live is trained to paraphrase text it is given to say aloud. We have not measured how closely our tutor follows its cue lines, so test this if exact wording matters.

How long can a real-time voice session last?

OpenAI documents a 60-minute maximum for a Realtime session. Gemini Live limits audio-only sessions to 15 minutes without context window compression, and a connection lasts about 10 minutes, so long lessons need both compression and session resumption.

How do you stop an AI tutor from drawing too fast?

Limit how often it can draw as well as how much. A per-lesson total still allows bursts, and if reaching it ends the lesson it also sets the lesson length. Ours ended lessons after about 3 minutes.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.