The Debrief · No. 19 · Tue Aug 18 Google halved the price. Until January. The intro rate on Gemini 3.7 Flash doubles on January 1, and almost nobody covering it said so. |
Welcome to the Debrief, the Tuesday half of The AI Debrief. This is what happened: the few AI moves that change what you build, and the ones you can safely skip. Friday is the Toolkit, where we hand you something to use.
Here's what we're getting into this week. Two price changes, one of which already hit your bill on Saturday, and a set of endpoints that stopped working this morning.
In this email ยป The half-price model that isn't, and what doubles on January 1 ยป The image endpoints that died today, and the one-line fix ยป Why your DeepSeek bill went up on Saturday without an email |
| ๐ธ The discount has an expiration date |
Gemini 3.7 Flash is cheap on a timer
Google shipped Gemini 3.7 Flash on August 13, three weeks after 3.6 Flash and, oddly, before Gemini 3.5 Pro. Input runs $0.75 per million tokens, output $3.75. Most of the coverage called it a price cut and moved on.
Today $0.75 | | January 1 $1.50 | | Change +100% |
Read the pricing page and the word "introductory" is doing a lot of work. Every number on it is dated through December 31, 2026. On January 1 input goes to $1.50 and output to $7.50. Batch pricing doubles too, from $0.375 and $1.875 to $0.75 and $3.75. Context caching doubles: reads from $0.075 to $0.15, and storage from $0.50 to $1.00 per million tokens per hour. Nothing is grandfathered, and the increase is automatic.
The second thing worth noticing is the order Google shipped in. A Flash model before its Pro model, with the launch pointed squarely at coding and agents. On Google's own AutomationBench, which measures enterprise workflow automation, 3.7 Flash scores 30.4% against 17.0% for its predecessor. That is close to a doubling. It is also a vendor-reported number, so treat it as a claim rather than a fact until somebody independent replicates it.
I keep coming back to the calendar rather than the benchmark. Four and a half months at half price is real money if you are running multi-step agents, and it is also exactly long enough to build a cost model, forget how you built it, and get surprised in January. The pattern underneath is the one to actually watch: Flash-tier models absorbing work that used to require a frontier tier. If that AutomationBench jump survives independent testing, it reprices a whole category of automation, and the interesting question stops being which model is smartest.
Why you care If you are running agent or coding workloads on a frontier model, there is a real arbitrage here until December 31, and you should take it. Then put a reminder on your calendar for December 15 to re-run your numbers at $1.50 and $7.50, because that is the price you will actually be paying. A cost model with an undated assumption in it is not a cost model. |
See the pricing page โ
โ Google shut off the Imagen 4 endpoints today. All three variants - generate, ultra and fast - hit end of life on August 17. The replacement is gemini-3.1-flash-image. If you have an image pipeline calling those endpoints, it is not breaking soon. It broke this morning. One string swap. The robotics preview endpoint goes on August 31, so check that too while you are in there.
โ Your DeepSeek bill went up on Saturday. V4-Pro launched August 13 and the restructured API pricing took effect August 16 at 16:00 UTC. Outlets reported that date at least three different ways, so if you read somewhere that prices "change this week," they already changed. The detail nobody led with is the useful one: off-peak is exactly half of peak, and peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC. Batch work is now schedulable for a straight 50% discount. Caixin reports some rates rising as much as 1,100%; I could not reproduce that from the published rate card, so treat it as reported rather than verified.
โ Six companies that agree on nothing agreed on a plugin format. Agent Plugins 1.0 shipped with AWS, Microsoft, OpenAI, Google, Vercel and Anysphere as co-maintainers. It bundles agent skills and MCP servers into a single installable unit, with the vendor-specific parts namespaced so a plugin stays portable across tools. If you build internal agent tooling, this is the first credible sign you get to build it once instead of once per platform.
Nvidia backstopping $105 billion for OpenAI's 10-gigawatt data center in Ohio. Biggest number of the week and the least relevant. You cannot buy it, use it, or build on it, and the capacity lands years out. The money is real; the relevance is zero. One in the same bucket that deserves a caveat rather than a shrug: the Stripe and OpenRouter $7 billion acquisition headline. Bloomberg reported it. Neither Stripe's newsroom nor OpenRouter's own blog has said a word, and OpenRouter published twice yesterday about something else entirely. Several outlets have already quietly upgraded "reportedly" to "closes." If OpenRouter is your routing layer, change nothing yet. |
The one thing actually worth doing this week: check whether anything you own calls an Imagen 4 endpoint. That is a live break, not a warning. Everything else here is a calendar reminder. Honest worst case if you ignore all of it - you overpay a bit in January, and you find out about the broken image job when a customer tells you.
One more thing before you go: we just opened the free 30-Day AI Academy Challenge โ one short lesson, one 10-minute worksheet, and one check-in a day, for a month. By Day 16 you've built your own personal AI OS. Day 1 is live now, and it costs nothing.
Start Day 1 โ
That's the Debrief.
See you Friday for the Toolkit.
From Drew and The AI Debrief team
P.S. What's the oldest API endpoint still running inside something you own? Hit reply and tell us. We read every one.
P.P.S. New here, or skipped one? Every past edition is in the archive.