← Archive

The Toolkit·No. 27·September 18, 2026·56,076 reads

Google's $1.38 voice agent has a catch

Gemini 3.8 Live is free to try and less than half of OpenAI's price. The version that tops the leaderboard is not the cheap one. Here's what to build on it anyway.

Welcome to the Toolkit, the Friday half of The AI Debrief. Tuesday is what happened. Friday is what to do about it: one tool worth your time, one play you can run today, and nothing you have to buy.

Here's what we're getting into this week. On Tuesday we said the only hard fact on OpenAI's voice launch page was the price: five cents a minute. Five days later Google answered with a voice model that costs less than half of that, is free in AI Studio, and carries an asterisk worth reading before you wire it into anything.

In this email

» Gemini 3.8 Live: what it costs, what is free, and which of the two models the headline price belongs to

» A 15-minute voice rehearsal partner for the call you are dreading, no card and no key

» The October 14 date to write down if anything you run still selects GPT-5.5

🎙 The tool: two models, one headline price

Gemini 3.8 Live

On September 15 Google released two speech-to-speech models into the Gemini API and AI Studio. gemini-3.8-live is the default for fast voice agents. gemini-3.8-live-extended-thinking reasons in the background while it talks, with a thinking level you set to low, medium or high. Both take live audio, text and images at up to one frame a second, call your functions or Google Search mid-conversation, can be interrupted, and hand back a transcript of both sides. Google says 97 languages, switching mid-conversation, and all audio output carries a SynthID watermark.

The price is the reason this is in the Toolkit. On the pricing page the free tier is free of charge in both directions. On the paid tier, audio in is half a cent a minute and audio out is 1.8 cents a minute. If both sides talked for a full hour without pausing, that is $1.38. OpenAI's GPT-Live-1 is five cents a minute, $3.00 for the same hour, for the voice layer alone with the text model underneath billed on top.

Quality index · cost per hour of input audio
Gemini 3.8 Live
Index 76.0  ·  $0.84/hr
  
Gemini 3.8 Live Extended Thinking
Index 82.6  ·  $3.50/hr
  
OpenAI GPT-Live-1 (Astra)
Index 81.5  ·  $5.83/hr
 
Index scores and hourly costs measured by Artificial Analysis, not by the vendors.

The catch

Google's launch post leads with first place on the Artificial Analysis speech-to-speech index, and that part is independent: 82.6, ahead of GPT-Live-1 on Astra at 81.5 and Grok Voice Think Fast 2.0 at 81.3. But 82.6 belongs to Extended Thinking at its highest setting, which Artificial Analysis measures at $3.50 per hour of input audio. The standard model, the one at the headline price, scores 76.0, ranks fifth and measures $0.84. The thinking tokens are where the bill goes.

The cheap one and the good one are different models.

Even so, the top configuration undercuts OpenAI's, which Artificial Analysis puts at $5.83 an hour, and time to first audio is a wash: 1.35 seconds against 1.34. So the honest summary is not "a third of the price". It is "about 40 percent cheaper at the top, and a genuinely cheap second tier that OpenAI does not offer".

The parts that are not on the launch page. It is turn-based with interruption, where GPT-Live-1 is full-duplex and keeps listening while it speaks. If the feel of the conversation is your product, test both before the price decides for you. Sessions are short by default: the session docs cap audio-only at 15 minutes and audio plus video at 2 minutes unless you turn on context compression, and a connection lasts about 10 minutes before you resume it with a token. Fine for a rehearsal. An engineering task for a support line. And on the free tier the pricing page says your data is used to improve Google's products, so it is for practice and prototypes, not customer calls or anything under an NDA.

Who this is not for. Anyone shipping a regulated call flow this quarter, and anyone whose product depends on the model listening while it talks. Everyone else just got a free voice agent to learn on.

Open it in AI Studio →  ·  the Live API guide

⚡ Steal this: a rehearsal partner, ~15 minutes

Practise the hard call out loud before you make it

What it replaces: rehearsing in your head, which does not work, or borrowing an hour from a colleague who will go easy on you. What it costs after: nothing on the free tier. The 15-minute session cap is a feature here. Nobody needs a longer rehearsal than that.

The ten-minute rehearsal

15 minutes · no purchase required

1
Open AI Studio's Live page with any Google account and choose Gemini 3.8 Live in the model picker. No card, no API key.
2
Paste the prompt below into System instructions and change the first two lines to your call.
3
Turn on the microphone and make the call for ten minutes. Do not restart when it goes badly. That is the part you came for.
4
Say "debrief". Copy what comes back into your notes, next to the transcript.
5
Run it once more with the character made harder, and fix only the one thing the debrief named.

YOU ARE: a skeptical procurement lead at a 200-person company.
I AM: a founder pitching a $2,000-a-month tool.

Stay in character until I say "debrief". Push back on price at least twice. Ask one question I will not have prepared for. If I talk for more than 30 seconds without making a point, interrupt me. Never help me, coach me or soften.

When I say "debrief", drop the character and give me, in this order: the three weakest things I said, quoted word for word; the question I dodged; the sentence I should have opened with; a score out of 10, and the single change that would raise it most.

The interrupt rule is what makes it a rehearsal. Left alone, a voice model is the most patient listener you will ever meet, and a patient listener teaches you nothing about a real buyer. Telling it to cut in when you ramble is the difference between talking at a mirror and being in the room. It also happens to exercise the one thing turn-based models do well, which is take the floor back.

"Quoted word for word" is what makes the feedback usable. Ask any model how a conversation went and you get "consider being more concise". Ask it to quote your three weakest lines and you get three lines you can check against the transcript and rewrite. It is the same move as the timestamps two weeks ago: the model earns your trust by pointing at the evidence, not by summarising it.

The play outlasts the tool. Swap the first two lines and it is a salary conversation, a hard client call, an investor update or an interview, in any of the languages on Google's list. If the character feels slow, run the same prompt on Extended Thinking and you will hear what the extra $2.66 an hour buys. And the one rule that comes with the price: keep real names and numbers you would not put in a public doc out of the free tier.


One date to write down. OpenAI retires GPT-5.5 from ChatGPT, ChatGPT Work and Codex on October 14, on every plan. The API is not affected. Anything that still selects it by name - a saved model setting, a workspace default, a custom agent, a scheduled task, a Codex script - needs to point at GPT-5.6 Sol before then, or it breaks on a Wednesday morning with no warning beyond this one.

If you do only one thing from this issue, run a single ten-minute rehearsal before your next real call. Honest worst case: you feel slightly ridiculous talking to a browser tab. Likely case: you hear yourself dodge the pricing question before a customer does.

Every play this newsletter has shipped lives with the people testing them in the free AI Academy.

Join the free AI Academy →

Free, no card. The tools and people behind this issue live here.


That's the Toolkit.

See you Tuesday for the Debrief.

From Drew and The AI Debrief team


P.S. Which conversation would you rehearse first - a pitch, a raise, a hard client call? Hit reply and tell us. We read every one.

P.P.S. New here, or skipped one? Every past edition is in the archive.