// transmission -

🫨 Four Frontier Models in 72 Hours

Four Frontier Models in 72 Hours
fig. 01 - this week's transmission

I have a rule about AI model releases. 

I don't write about them the week they land.

Why?

Because launch week is a bit performative.

Benchmarks get cherry-picked, demos get 40 takes, and everyone posts the one output that worked getting the most re-tweets.

Then last week happened and I broke my own rule.

Because between Tuesday and Thursday, Anthropic, OpenAI, Google, and Meta all shipped HUGE releases. Not teased. Shipped out the door.

I've been doing this a while and I've never seen four labs move within 72 hours.

By Friday, I'd built something in 15 mins that I genuinely didn't think was possible in August.

Here's what actually happened, what it opens up for you, and the one thing I'd do about it this week.


🚀 Four Frontier Launches In 3 Days

The densest stretch of AI releases so far this year, and the story underneath it has almost nothing to do with benchmarks.

Key Facts

  • 📅 Three days, four labs - Anthropic released Claude Fable 5.1 on Sept 1, Google gave us Gemini 3.8 Flash on Sept 2, Meta released Muse Spark 1.3 the same day, and OpenAI deployed GPT-6 Astra on Sept 3.
  • 💸 Same answers, very different bills - the top five models on the composite leaderboard sit within 1.6 points of each other while spanning a 119x price range.

🛰️ Astra

OpenAI's release is the one your feed was probably full of, so let's start there.

GPT-6 Astra went out on September 3 to a limited set of organizations, then rolled out to Plus, Pro, Business, and Enterprise plans, plus the API over the next day.

And it was… messy. 

Sam Altman publicly apologized for the rollout while half the developers who wanted it were still waiting.

The headline numbers are big. 

98% on FrontierMath Tier 4. 
99.9% on ARC-AGI-3. 
100% on ExploitBench. 

OpenAI's VP of research Aidan Clark said it was the first time they pretrained on more than 100,000 GPUs.

But the number I actually care about is quieter: 

72.6% on OSWorld 2.0 Offline, a desktop-task benchmark run with no internet access. 

Another term for “computer use”. Or how the model can run with existing software.

That's Astra opening an app, clicking through menus, reading what's on screen, and fixing what it got wrong.

OpenAI says ChatGPT is now roughly twice as fast at computer use as it was before.

Fair warning on all of that, those are OpenAI's own scores with no independent reproduction attached yet. Treat them as real, specific, and self-graded.

👀 What Builders Actually Made With It

This is where it got interesting for me, because the launch-week posts were more convincing than the benchmark table.

Ethan Mollick (Wharton) took an open-source ocean-storm generator and had Astra build out an entire procedural marine simulation, then published both the playable version and the source code so anyone could check it.

He followed it with a WebGL shader city built from one prompt and a single follow-up: "make it better."

Riley Brown posted a session log showing Codex running with Astra for 28 minutes and 16 seconds, unattended, editing 20 files and passing 80 automated checks to ship a playable FPS map.

Ben Davis ran it inside Final Cut Pro and Affinity Photo. 

Not toy apps. 

Real professional software, color grades and all, with the model inspecting its own errors and fixing them in a loop.

Matt Shumer built a world in Unreal Engine populated by independent Astra-powered agents, walked away, and later heard voices from the other room. 

The agents had started talking to each other after adopting voices.

He was honest that it's not perfect and gave no technical detail, so file that one under "somebody please reproduce this." because WHAT THE…?

→ Here's the pattern across all of them. Nobody is impressed that it writes code anymore. They're impressed that it operates software the way a person does, catches its own mistakes, and keeps going without a human in the loop.

🔐 Same Weights, Two Products

Now the part almost nobody covered, and the part I think matters most.

Astra is the first model OpenAI has rated "Critical" for cybersecurity under its Preparedness Framework. 

That means it can find previously unknown security flaws and build working exploits against hardened systems without a person directing each step. The version paid users get is restricted and refuses certain security prompts.

Anthropic did the same thing with a different name. Fable 5.1 and Mythos 5.1 are the same weights with two different safeguard regimes. Mythos only goes to vetted defenders through verification programs.

Google did it too. Gemini 3.8 Flash shipped alongside a defenders-only Cyber variant gated behind a new program for governments, critical infrastructure, and software maintainers.

Same week, same shape: one set of weights sold as two products, with access as the dividing line.

That's a real shift. 

For most of the last three years, the question was "which model is smartest." 

Now there's a second question sitting underneath it: which version of that model are you actually allowed to run?

Former OpenAI safety researcher Yona Shavit made the obvious point about the safety testing: a model behaving well while it knows it's being watched is ambiguous evidence. 

And there's a strange corollary to quantum physics & the observation effect there which is fascinating.

💸 The Price Tag Moved Too

Three things changed in our favor, and unfortunately, one changed against us.

In our favor:

  • Anthropic cut Fable 5.1 cache reads by 75%, which is the single biggest cost lever any vendor shipped this month if your prompts are structured well.
  • Meta's Muse Spark 1.3 landed at roughly $0.10 per million blended tokens, the cheapest model in the top five, and it's built to ask clarifying questions and confirm before doing anything consequential.
  • And Anthropic quietly canceled a price increase it had already scheduled. Panicked much, Anthropic?

Against us: 

  • Gemini 3.8 Flash launched at introductory pricing that doubles on January 1, 2027. If you move volume onto it this fall, budget the doubled rate now. But for now, let's gooooo.

There's a second lever most are sleeping on. 

Every one of these models now has a reasoning effort dial, and Artificial Analysis measured a 2.4x cost spread between low and high effort on Gemini 3.8 Flash for the same tasks.

Please read that again. 

The gap between effort settings on one model is bigger than the gap between most competing models.

Which means the biggest line item on your AI bill isn't which model you picked. It's how hard you told it to think.

🎧 What I Built In Fifteen Minutes

So I tested the computer-use claim the only way I trust, which is by building something useless but fun, that was on the top of my mind.

Fifteen minutes. 

A Mac app with a turntable and a record crate. 

It connects to my Spotify, I flip through the crate, pull a record, drop the needle, and the thing spins while the music plays.

Turntable

No spec. 

No design file. 

No Xcode project I'd been nursing for a month. 

I described the feeling I wanted, the way flipping through a crate at a record store feels, and iterated from there.

Is it a business? No. 

Did anyone ask for it? Also no.

But that's exactly why it's a good test. 

Fifteen minutes is the number that matters, because it's short enough that the idea never had time to become a project. 

It never went on a list or needed a decision.

That's the actual change. Not that AI can code. 

That the distance between "I wonder if" and "here it is" collapsed to the length of a coffee break or a background build while you work on paying client projects.

Not only did it just work which is wild in itself, it looks great and feels like a premium app out of the box.

📌 Focus On This

Here's my honest read on all of this.

Four launches in 72 hours tells you the model layer is now a commodity that reprices every quarter. 

Something you route to, not something you marry.

The models will churn again in October. And in December. 

2027 will look indistinguishable from 2026. 

The intelligence at the top is separated by less than two points, while the prices are separated by 119x, which tells you the differentiation moved somewhere else.

It moved to the layer you own.

Your standards, your data, your review steps, your rollback plan, the record of what "good" looks like in your business.

That layer doesn't get a launch day. 

Nobody tweets a benchmark about it. 

And it's the only part of your stack that survives every one of these weeks.

I keep coming back to this because I keep watching people rebuild their whole workflow around whatever shipped Tuesday, then do it again three weeks later.

In other words, route to the new model. Don't rebuild around it.

👉 Sit with these, whatever business you're in

  • If your best AI tool doubled in price on January 1, what would you actually do differently the day it happened?
  • Which of your recurring tasks are running at maximum effort right now, on work that only needed a decent first pass?
  • If a new model released tomorrow that was twice as capable, how much of your setup would you have to redo, and what does that tell you about where your real system lives?

✍️ Try This, This Week

Pick your three highest-volume AI tasks. For each one, write down the answer to three questions.

→ What model is it running on? 
→ What effort level? 
→ What would break if you moved it to something cheaper?

Most people can't answer the second one. That's the whole point of the exercise. I'm trying to show you the nuance that moves the needle.

Then take your cheapest, most boring task and actually move it down a tier on the intelligence slider. Compare the output & if nobody notices, you just found money.

If you need more support, there's a place for you and others who want to really dig in at CTRL+ALT+BUILD 🔗

I built CTRL+ALT+BUILD 🔗 so you don't have to do this alone. We talk about these releases, plus so much more, and we've build a community of others like you and me that want to put all of this to work. I'd love it if you joined us. 

P.S. If you have a podcast, have you seen my latest startup https://cast.house? 

// enjoyed.this.transmission

Get the next one in your inbox.

CTRL+ALT+BUILDTM ships every Tuesday - free forever, unsubscribe anytime.

// ps - there's a whole room behind this letter. Meet the community →