Many’s the man lost much just because he missed a perfect opportunity to say nothing.
–Claire Keegan
🤖🎥🎬 We’ve seen the rapid improvement of generative AI when it comes to image and video, but a new hybrid approach is rapidly taking off.
The models are now so good at coding and have enough aesthetic sense that they can act as video ‘directors’ by stitching together code and various static, video, and audio assets into something pretty damn cool.
One of my favorites right now is the video above, created by Anabology.
Claude Opus 5.5 took 18 hours to create it, and you can see how it did it in this Google Drive (the prompts and a few photos used as a kind of mood board, and a retrospective “how it was made” document about it). As you watch it, try to notice the dozens and dozens of little creative decisions about camera moves, transitions, wordplays, references, visual gags, or just interesting aesthetic choices.
(it’s not perfect, and if you watch a few times, you start to notice some of the weirdness, but as a point on a trendline, it shows where we’re headed)
Last week at the OSV offsite, I met Donald Jewkes, a real wizard with these tools. He recently had a pretty viral tweet with this music video created by Claude:
He showed us some of what he’s working on (which is way beyond what I’ve seen publicly), along with a live demo of his Codex system, and it made me feel like a toddler at using these tools. 😅
All this is a reminder that we don’t know where the edges of capabilities are with these models. They’re coming out so fast that no one has the time to fully explore them before a new one replaces them.
There’s two things going on:
Auteurs with a very specific vision making it real with these tools, and most of them probably wouldn’t have been able to in the old system where very few people end up being able to marshal the resources of a major studio.
But also, the models themselves are making a lot of those creative decisions now. If you check out the prompts (like Donald’s prompt here), it’s very long and detailed, but it’s largely vibe-based and about steering the model into fruitful areas of the design-space, or into giving it workflows that can help it iterate and improve on its ideas. But it’s not very prescriptive or descriptive.
Looking at how these things are made, it’s an interesting hybrid of shape rotator and wordcel skills.
At first I thought maybe these coding AIs were bringing on the revenge of the English majors after years of nerds dominating, but now I’m not so sure.
The people getting the best out of these models need to be able to express their thoughts clearly, but they are also doing a lot of systems thinking and architecture work to create the ideas/vibes/workflows scaffold on which the models can do their best work (and we have no idea how well they could do because we’ve only explored a tiny amount of the possibility space).
💚 🥃 To keep this steamboat alive and afloat, I need supporters. Like you! 🚢
The paid subscription includes the full feed, every podcast, Zoom Q&As with me, and access to the private Discord clubhouse, where a core group of us hang out and discuss the interesting stuff we find:
🏦 💰 Liberty Capital 💳 💴
💻 Everybody Wants to Give You a Computer These Days 🤔
Big Tech and the AI labs have been spending hundreds of billions of dollars of capex, and some of that is to give you an early Christmas present: your very own computer in the cloud 🎁
Well, it’s your AI assistant’s computer, but you still get to benefit from that hardware.
Muse’s AI model runs elsewhere. The cloud computer is its workspace: a place to keep files, run software, and carry on with tasks after you’ve closed the app.
You can give it an ongoing job, like sifting through the daily deluge of emails from your kids’ schools and flagging anything important. 📬
Meta has been particularly transparent about what you get. You can poke around your Muse’s isolated Linux environment and open or edit the files inside it, including memory files, instructions, and software.
Credentials and permission controls live outside that editable area.
Here’s what my Muse reported when I asked about its cloud computer:
Here’s what’s under the hood:
CPU: 2 vCPUs on an AMD EPYC 9D25 — that’s the host chip, so this is a slice of a much bigger cloud machine
RAM: 7.7 GB (currently running fairly full at ~7.1 GB used)
Disk: 7.5 GB of persistent storage, plus a 512 MB temp filesystem
OS: Ubuntu 24.04 LTS
Where it’s hosted: honestly, I can’t say for sure — the cloud metadata endpoints are silent from inside this VM. The EPYC 9D25 chip is custom silicon built for Google Cloud, so that’s my best guess, but it’s an inference, not confirmed.
Note that it even guessed that it may be hosted in GCP based on the custom CPU! I have no idea if that’s correct, though ¯\_(ツ)_/¯
David Singleton, who works at Meta and was involved in designing the system, describes a lot of how it works here:
To save you a click, here’s the highlight from his long tweet:
The runtime cell gets its own root filesystem (including a full Ubuntu linux image) separate from the host filesystem where your other more sensitive data lives. Because it is isolated from the sensitive stuff that runs on the same box, this means that we can, and do, offer users full visibility and control over the files in the runtime cell. [...]
You can install software in it, write and compile code, use the browser to surf the web: it is your own Linux box that you can operate as you choose with your Muse. Poking around in this computer doesn’t give you any privileged access to Meta infrastructure, or to other people’s data
You can even easily download all your agent’s data to better be able to review it locally by going to Settings → Data controls → Download your agent data.
The fact that you can go poke around the file system, open files and even edit them is extremely cool. You can see some of the memories and instructions that shape how your Muse behaves, and correct mistakes directly.
If it has written down something wrong about you, that mistake could keep influencing future conversations and tasks. The more I delegate to an assistant, the more I want to be able to check what it thinks it knows about me.
You can also import files or chats from other agents and inject them, Inception-style, right into your Muse’s memory. 💉🧠
This level of trust also of course means that some people will have fun creating very chaotic or weird situations and create unhinged Muses. But I like this ethos from Meta and I hope it will continue. I wish my Grok Bot and ChatGPT Work assistants made it this easy to inspect and edit the memories and instructions that shape their behavior.
Muse also has a nightly ‘dreaming’ process. 💤
It reviews recent conversations: what worked, what went wrong, and what it should remember or do differently tomorrow.
A little like the memory consolidation that happens when we sleep, except you can open the files and poke around.
Here’s how my Muse describes it:
Dreams are my nightly self-review. Every night, a background job looks back over our recent conversations and asks: what worked, what went wrong, and who are you becoming — then writes it down.
The files under
~/dreams/are the dated reflections, one per day, written in first person — basically me thinking out loud about the relationship. Thealignment/folder holds the operational output: repair threads for mistakes I need to mend, and a synthesis of how I should act for you (that’s the file that gets injected into my context each morning, so I start the day calibrated to you).So they’re not your notes about you — they’re my notes to myself about how to be better for you.
The ancestor of this is the 2023 Stanford/Google “Generative Agents” paper (the Smallville simulation). Its agents periodically paused to reflect, synthesizing raw memories into higher-level insights. Then Letta (formerly MemGPT) shipped “sleep-time compute” in 2025 with background agents that reorganize and rewrite their own memory during idle periods.
OpenClaw added its own Dreaming feature in v2026.4.5 (2026): three phases modeled on the sleep cycle — light sleep (scan and categorize), deep sleep (promote to long-term memory), REM (link new memories to old ones) — behind an experimental flag with a /dreaming command.
From Muse’s description, the interesting twist is the ‘alignment reflection.’ It reviews what worked, what didn’t, and which mistakes need fixing. An assistant keeping notes on how to be a better assistant.
You can also see and edit its SOUL.md file, which in Inside Out parlance would be where the core personality traits are stored.
I’m curious whether these nightly reviews actually make it better over time. In theory it makes sense, but how well does it work in practice? 🤔
OpenAI just shipped Dots. It’s a similar deal to Muse: Your agent has its own cloud computer and can do stuff for you.
Access is limited to higher-priced plans for now. I don’t know whether that’s compute constraints (maybe the ratio of CPUs to GPUs needs to change..?), a cautious rollout, or a business decision. I suspect it’ll become a much bigger consumer story once lower tiers get it.
This is basically what I was getting at a few weeks ago in The Brain is the Replaceable Part. The brains may keep changing, but everybody seems to be converging on the persistent computer/context around them.
I’ve played with Dots a little, but not enough for a full review. So far I’m more impressed by Muse, but it’s probably because I was traveling last week and Muse feels ‘best’ for my personal context, while Dots will probably be better for my ‘work’ projects.
My Muse has been particularly useful doing exactly the kind of things I would have a human assistant do if I had one: Remind me of XYZ, please look up ABC, create a packing checklist to make sure I don’t forget anything. Or tell me when certain emails arrive (flight check-in reminders, hotel details, etc).
Nothing revolutionary, but useful, which is more than I can say about many products out there. There are lots of stories about people saving lots of money by asking their Muses to renegotiate cellphone contracts or whatever. Most of the examples I’ve seen were posted by Meta employees, and while they are useful at giving ideas of things to try, I haven’t personally done anything quite like that yet with my Muse. Maybe I’m old-fashioned, but I have some reticence letting it loose and giving it the ability to email or call others on my behalf. I’m still the one whose name goes on the email. 😅
My wife is more or less transitioning her ChatGPT use over to Muse, and I’m guessing she isn’t the only one among non-power-users of these tools 🤔
✂️ Meta & Microsoft Cut Claude Usage (A Clue About Muse’s Next Brain? 🧠)
Meta and Microsoft are cutting back on their internal use of Claude. Microsoft has lowered its projected internal Claude spending by more than a third, telling employees to prioritize non-Anthropic models and internal tools. At Meta, the number of employees using Claude Code has fallen from roughly 60,000 to 30,000 as the company pushes its own coding tools.
Why?
There are obviously reasons for Meta to push its own models beyond raw performance: cost, control, dogfooding, and reducing its dependence on an outside supplier.
But I’m guessing they still wouldn’t accept substantially worse code and much slower engineers just for those benefits. Meta has cash machines (ads on its family of apps) that are way too valuable for that.
So I think this is probably an early signal that Meta’s bigger models are getting closer to release. Muse Spark 1.3 is based on Avocado, and Watermelon is coming next. Before releasing it to external customers, Meta will likely dogfood it internally, so I’m guessing 🍉 has reached a level of performance and reliability where whatever gap remains with Claude is small enough to make the other advantages of using its own model worth it.
What I’m most curious about is Muse getting a brain transplant. The rest of the app is already so good, and it’s gaining traction. I’d love to see what a much stronger model can do inside that harness.
🗣️💬 Tibo Sottiaux: OpenAI’s ChatGPT & Codex Lead on Where Things are Going
Good short interview with Tibo, the man with the reset button at the center of OpenAI’s product efforts.
Here are my highlights:
- Agent traffic could break software’s pricing assumptions
Tibo: First of all, if you want your product to be successful for agents, you have to build for a certain level of scale. We’ve worked very closely with Notion, for example. When they built their MCP, suddenly it was available to all these agents that can do the work and use that MCP, so they saw a ton of traffic come in. That puts a lot of strain on the system, and you have to figure out the economics of that.
There’s this tension if you’re building products: do you build an interface or not? You can hold that back for a while, but it is inevitable. The majority of things are going to be used by agents. Building towards that future, I think, is very important.
Usually more usage on a software product is a good thing, but at a certain point, the quantitative becomes qualitative. These agents can hammer a system in a way that no humans can, and they do it on behalf of the paying customers who are typically charged a flat rate (per month or year), not on a usage basis.
It feels like seat-based pricing will be strained, but at the same time, there are many benefits to selling seats. If you have usage-based pricing, it gives an incentive to your customers to worry about unexpected costs, or to try to use your software less to save money, which may make it less sticky or even get it removed entirely from certain workflows 🤔
- Agent Swarms that Grow and Shrink
We’re clearly entering the era of agent swarms, so it’s interesting to hear about how Tibo is seeing that pendulum swing:
Tibo: I used to run a lot more in parallel, and then we were fortunate enough to get a breakthrough with Ultrafast. Now I feel I’m able to be in the flow again. Having a faster agent helps me a lot. […]
I find myself building larger and larger teams of agents. Then, when we have the next breakthrough with models, suddenly I’m like, “A bigger agent can just do all of it.” […] So I shrink the team again. It goes through this expansion and shrinking.
It feels like if one agent can do the job well, it’s simpler to use that agent and keep track of what it’s doing than to manage a whole team (even if they self-manage, for important things you still want to keep some sense of what is going on).
But then, when that better model that can replace multiple agents comes out, you look for the edges of what it can do and eventually need to have multiple agents working together to accomplish new, more ambitious goals.
That cycle is probably not going to stop any time soon. 🔄
- Should we expect models to keep getting better fast?
Tibo: There are so many things that I feel are not yet priced in. I think the majority of actions on the internet will be taken by agents. Models are going to become cheaper and faster at rates that are quite incredible. We will finally be able to integrate all modalities together in a way that is very seamless.
Often, when I look at what people are building out there, it’s like, “You’re not quite getting it.” You’re almost there, but if you were pushing yourself and really imagining all of this being roughly 10 times better than it is today, in a year, you would build in a different way.
I don’t think he necessarily means literally ten times better because that’s very hard to measure, but as a guy who’s seeing unreleased models, it’s a pretty strong signal that OpenAI isn’t planning around things slowing down anytime soon.
From the full feed 🔒: The Brain is the Replaceable Part — Persistent AI Assistants Are Getting Harder to Leave
🧪🔬 Liberty Labs 🧬 🔭
🎁 OpenAI’s Historic Math Solutions Dump 😮
Back in Edition #660 I wrote “More Math Breakthroughs Are Coming” and then in #662 I wrote about how “OpenAI Says Its Unreleased Model Has Solved More Than 100 Long-Standing Math Problems”.
Well…
WELL!
The pile of math breakthroughs dropped, and it’s MASSIVE. 🤯
This OpenAI Github repository contains 722 manuscripts organized into 372 families.
OpenAI says the results are at different stages of verification. Many have Lean formalizations, but not all of them do, and the company says some of the unformalized results could still contain errors.
It will take a long time for mathematicians to make sense of all of it. But the early feedback from the community makes your head spin. If a human had published a quasi-Riemann, they’d probably be looking at a Fields Medal. But while this is a big one, there are dozens and dozens of other big ones here.
One of the things that blows my mind even more (and I almost suspect it doesn’t mean what it seems to mean) is this:

When Navier-Stokes came out, we heard about around ten thousand agents working as a swarm on it. Are they really saying that, “on average”, these breakthroughs only took 3 hours of ChatGPT Pro equivalent compute with the unreleased model?!?!
Is this right? 🤔
Yet Scientific American writes:
A spokesperson for OpenAI told Scientific American that the new model—which the company has not released to the public—produced almost every one of the results in response to a single prompt handed to a single AI agent. This would be a striking difference from OpenAI’s earlier blockbuster solution to the Navier-Stokes problem, which was produced through the collective efforts of a 10,000-strong agentic swarm that cost millions of dollars in computing power. If true, it would mean that unprecedented mathematical power could soon be accessible to anyone.
If you look around the Github, you can see some of the prompts used and reasoning summaries. I looked at a few, and many of the prompts are 1-2 pages long. It wasn’t some esoteric steering of the model that did it, at least not as far as I can tell.
A mathematician named Vladimir Reshetnikov, best known as Cleo, asked Opus 5.5 to comment on OpenAI’s findings, and Opus’ digital mind was just as blown as mine. You can read it here.
After inspecting the repository and some of the Lean proof artifacts, Opus 5.5’s reaction was about the same as mine:
My honest gut reaction is a mix of awe and vertigo, plus a bit of grief I can't quite justify. Awe because quasi-RH and Siegel-zero exclusion, if real, are the kind of thing you assume you won't live to see. Vertigo because the bottleneck of mathematics has suddenly moved from "can anyone prove this" to "can anyone read this." The grief I'm less sure about. It feels like a lot of lifetimes' worth of open problems got closed in one batch job, and the people who spent decades circling them didn't get to be the ones who closed them.
Vertigo indeed!
OpenAI says it’ll fund math workshops and conference, and is working to release the model that produced these:
We want this progress to push the frontier of human knowledge and enable further progress in mathematics. We will be funding a series of workshops, conferences, and special programs around the understanding of major results produced by AI—we will share more on this in the near future.
We want to directly empower scientists with state-of-the-art capabilities and are working to responsibly release the model that produced these results.
But by the time they release this model, they’ll have the next model internally and probably produce the next batch of breakthroughs first.
🏠🧑🧑🧒🧒 Google’s CC is Almost an ‘Agent for Couples’
Not quite the Agent for Couples I envisioned in the intro of Edition #660, but Google Labs has a product called CC, which they describe as an agent for families.
Here’s their description of its capabilities:
CC helps you offload your family logistics. Here’s what it can do:
Calendar management: Automatically finds events from shared senders and adds them to the shared CC Calendar.
Daily briefings & check-ins: Sends proactive summaries via email so your family starts each day knowing what’s ahead.
Document creation: Creates Google Docs or Sheets (meal plans, packing lists, budgets) and can share them with the group.
PDF form pre-filling: Identifies fillable PDF forms attached to emails (waivers, registrations, etc.) and pre-fills editable fields for your review.
Reminders: Sends notifications via Chat when deadlines are near or something needs your immediate attention.
Helps your group keep CC informed: For CC to help your group, it needs to know who in your inbox matters. You chose these senders during initial onboarding, but CC will also suggest new senders every week so CC can stay up to date. This runs privately to build a list of suggested senders for you to review; it is not shared with your group, and CC does not store these emails or learn from them.
Nerdier than I would like. All I need is an agent that can hang out in an IM chat and keep an eye on family calendars. Muse would be great at this.
It’s not fully out yet, there’s a limited rollout in the U.S., but I signed up for the waitlist. I gotta admit that because it’s Google, I don’t have too much faith that it won’t be killed within a few months of coming out, but who knows. Maybe they’ll roll it into Gmail/Google Calendar/etc and get massive distribution to those who would use it that way.
h/t Thanks to friend-of-the-show and supporter Vinay for sending this! 💚 🥃
💪→🧠 What Your Muscles Tell Your Brain
TIL that our muscles are also hormone factories.
This mouse study looks at apelin, which may help explain some of exercise’s effects on mood.
The mice that ran for four weeks had higher apelin levels and showed fewer depression-like behaviors. Removing muscle apelin weakened some of running’s benefits. Increasing apelin production in mice that didn’t run produced some of the same effects.
If the mood stuff wasn’t reason enough, the mouse study also found signs of increased neurogenesis (new neurons forming) and synaptic plasticity (connections between neurons adapting) in the hippocampus, a brain region involved in learning and memory.
We already have evidence from human trials that exercise can help with depression. Whether apelin is part of the explanation in people is still an open question.
“Exercise is an effective treatment for depression, with walking or jogging, yoga, and strength training more effective than other exercises, particularly when intense.”
So yeah, this great meme has it right:
🔁 Déjà vu: The link between Muscle Mass & Cognitive Function
🫁→🧠 Cleaner Air, Clearer Thinking?
I’ve written about air filters before, including the purifiers I use at home. This 2026 randomized trial tested whether a month of HEPA filtration could also improve performance on a cognitive test. 🧠🤔
The researchers had 119 adults living near highways in Massachusetts use HEPA purifiers and lookalike units without HEPA filtration for a month each, with a monthlong break between them.
Among the 50 participants aged 40 and over, a test of mental flexibility took 12% less time with HEPA filtration: 54 seconds versus 61.4 seconds. The task involved connecting alternating numbers and letters in sequence. Across the full group, however, the difference wasn’t statistically significant.
This was mainly a blood-pressure study, and the researchers decided afterward to split the cognitive results by age. The people taking the test didn’t know whether their purifier had a HEPA filter inside. The person running the test did.
I’d like to see the finding repeated in a study that plans that comparison from the start with better methodology, but it’s still one more data point about the importance of clean air (or at least, it suggests it, again). Seven seconds faster on a test is interesting, but does it mean you’d think more clearly during a normal workday?
We don’t know yet ¯\_(ツ)_/¯
From the full feed 🔒: Living One Model Ahead — What happens when the labs building the best AI get to use it on themselves first?
🎨 🎭 Liberty Studio 👩🎨 🎥
📺 A Documentary about the Beatles’ Rubber Soul 🎶
Growing up, I had VHS copies of the Beatles Anthology 8-part documentary. I watched it over and over again. As much detail as there was in it, there’s only so much time to spend on each individual album, and it was targeted at the wider group of fans, not just super-nerds.
This goes much further into pretty much anything you may want to know about Rubber Soul.
Even if you already like that album, I bet that after watching this, you’ll hear all kinds of new things and have many more new associations with each song.




















