Liberty Workbench is an occasional series where I put one idea on the bench and think it through with you.
One underrated advantage of building the best AI models is that you get to use them before anyone else.
For a while, the leading labs are basically living in the future. 🔮
Thibault “Tibo” Sottiaux, who leads Codex and is master of the Reset Button, put it this way:
This is already a big deal when models improve at their normal rate. But every once in a while there’s a bigger jump, like Anthropic with Fable or OpenAI now with Astra. When the lab itself gets access to that way-ahead model before everyone else, that can change how fast it moves. 🏎️
Just a few days ago, OpenAI said they’ve now reached their “automated research intern” milestone: give the AI a well-defined research task that would take a skilled researcher a few days, point it in the right direction, and steer it as needed while it works. By mid-August, OpenAI’s research org was using 3.1 agent-workdays for every human workday. The ‘median researcher’ was burning through more than $600/day of inference at API-equivalent prices, and they were running more experiments per ‘active experimenter’ than at any point since OpenAI started tracking this in January 2025.
That obviously doesn’t mean research is moving 3.1x faster, but that’s a LOT of extra agent work to throw at your own problems. And OpenAI says agent usage is still growing rapidly. Of course, as more of the work gets automated, the bottleneck probably just moves.
If you’re at Meta or SpaceX AI or even Google and you’re trying to compete, you may have some pretty good unreleased internal models too. And you may be paying one of the frontier labs for access to their best public model (eg. Meta is reported to be spending billions annually on Claude).
But if the models available to you aren’t competitive with your rivals’ unreleased ones, you may be falling further behind unless you can compensate in other ways.
The key question is whether that private lead compounds faster than it diffuses. If each unreleased model helps the lab build the next one before competitors catch up to the last one, the gap doesn’t necessarily reset when the model is released publicly. It may reproduce itself. 🤔
Astra is barely out, and already OpenAI is announcing what could be a major math breakthrough made by what they call “an OpenAI next-generation model significantly more capable than GPT-6 Astra…” with on the order of 10,000 concurrent agents… (Yes, I know there’s controversy about how OpenAI got there and who deserves credit. I have no idea who is right. If the proof holds up, though, it’s a historic breakthrough for both mathematics and AI.)
That doesn’t mean Astra helped build this next model (what will they call the next size class? Galaxy?). The timing may not line up. But it still shows the problem for competitors: by the time they get access to one generation, the frontier lab may already be using the next one.
This reminded me of the Boris Cherny interview I wrote about (he’s the creator of Claude Code).
Anthropic is already running hundreds of agents every day, sometimes thousands, across its codebases doing maintenance work that Boris says used to require dozens or hundreds of engineers.
And in that crazy experiment where he gave Claude the job of rewriting Anthropic’s Electron desktop app in native Swift and let it keep working for more than two weeks, Boris estimated the workflow had spawned thousands, perhaps tens of thousands of agents. At the time, I assumed these were Mythos/Fable, but who knows what unreleased model he may have been using 🤔
Again: it must be nice to work at the token factory! 😅
Having the better model isn’t enough. You also have to get good at using it.
OpenAI and Anthropic are building workflows and infrastructure that can turn thousands of agents into useful work. So the advantage may be something like:
Model quality × how well the organization can put it to work
A slightly weaker model that is deeply integrated into how a company works could create more real-world leverage than a stronger model if the org is slower to absorb it.
The familiar recursive self-improvement (RSI) loop looks like this:
better model → better AI research → better model
But the model can also improve everything around the model:
better model → better software/chips/infrastructure/company → better model
The company itself becomes one of the things the model is improving. 🔄
It’s a broader version of what I’ve been calling the Escher Hands era.
OpenAI says its models helped design and bring up its Jalapeño chip, including optimizing the chip’s arithmetic circuits. More recently, Codex with Astra has been helping program it. On selected GPT-OSS attention and mixture-of-experts blocks, the AI-generated implementations ran 1.5–1.8x faster than the existing human-expert-written versions.
What else are the unreleased models working on? Improving infrastructure utilization, optimizing kernels and networking, designing datacenters, maybe even helping with parts of CapEx planning…? 🏗️🚧
When you’re dealing with hundreds of billions of dollars of infrastructure, even a small improvement in how it’s designed or used can be worth a lot. 💰
At some point, this starts to resemble the leading edge of semiconductors (ASML, TSMC, etc): each generation costs more to develop, requires more specialized infrastructure and accumulated know-how, and fewer players can stay at the frontier.
The big difference is that AI capabilities and techniques can spread much faster than semiconductor manufacturing know-how. And once everyone else gets those models, they can use them to get better faster too. The catch-up process may compound too.
So can the frontier labs keep extending their lead faster than everyone else catches up? What I’ll watch is whether the same labs keep pulling away, or whether each big release changes who’s chasing who. Early access can be a huge advantage without any one company managing to stay ahead. 🤔
🔁 Déjà vu:
Liberty Workbench is where I put one idea on the bench and think it through. Every Workbench edition is collected here.












