651: Darth Vader Wants to Lead the Rebels, Starlink Gobbles Falcon 9, Amazon’s $220B AI Bet, Claude Presses Delete, GPT-5.6 Induced Amnesia, AI Needs a Black Box, and Band of Brothers Belongs in IMAX
"The enemy gets a vote"
The poorest way to face life is to face it with a sneer.
—Teddy Roosevelt
⏱️⏳🕰️🗓️ 📈💰📉💸 A few weeks ago, I wrote about Logtime: Why Years Feel Shorter as You Get Older.
There's clearly an investing version of that too.
Let’s call it Logtime for Investors.
When I think back to my early years of investing, some events seemed momentous at the time. They lasted a very long time. They felt like Big Wins (😃) or Big Failures (🥲).
But now, when I go back through my notes and look at the actual dates and numbers, I see they were just blips. 📓
A few months of underperformance. A stock dropping 30% because of some scandal. Some TARP warrants popping 45%, and then I sold.
At the time, subjectively, each one lasted a LONG TIME and was a BIG DEAL.
BUT
Now they barely register when I look back.
I had only been investing for a short time, so every quarter was a large fraction of my total time in the market. Every move up or down felt like a large swing when compared with all the volatility I had experienced so far (which wasn’t much, I just didn’t know it yet 😅).
Now it takes a lot to even get my pulse up. I’ve seen enough versions of these things again and again.
Don’t get me wrong, I’m definitely not perfectly stoic about any of it. But compared with the early days, it’s like looking through the wrong end of a telescope instead of through a microscope. 🔭🔬
🛀💭 The deepest recurring tension is infinite curiosity vs. finite attention.
💚 🥃 The paid subscription includes the full feed, every podcast, Zoom Q&As with me, and access to the private Discord clubhouse, where a core group of us hang out and discuss the interesting stuff we find:
🏦 💰 Business & Investing 💳 💴
👑 Darth Vader Wants to Lead the Rebels 😯
For a long time, Microsoft was basically the Empire, operating the Death Star of the tech industry (before it, IBM played that role).
They controlled so much for so long that most other tech companies had to position themselves either under their radar, in some niche, or as the rebels competing with them.
In more recent times, the Empire hasn’t been one company. It has split into a handful of Big Tech Kingdoms. The kings have been Amazon, Google, Meta, Apple, and the Nadella-revived Microsoft. Not quite as unified as the empire of old. But much bigger and more dominant, as the tech industry has swallowed a large fraction of the modern economy.
One thing I’ve noticed lately is that, because there are new challengers to the crown (mainly OpenAI and Anthropic), the Old Kings have had to reposition themselves and reframe what they do in response to the frontier labs that grew from startups into near trillion-dollar companies in just a few years.
And the message is becoming clearer, and it’s a little funny:
We’re the rebels. We’re here to help you defend yourselves against the new two-headed AI emperor.
If we look at the transcripts of the Q2 2026 calls from the Big Tech cos (except Apple, which has a very different AI strategy), we see it echoed many times.
But listen for what each company is selling while it says it.
First, Satya Nadella at Microsoft:
Going forward, we have two goals. First, ensuring AI empowers every person, amplifying their agency and ambition; and second, empowering every organization to build their own continuous learning loop and ensuring that they don’t outsource their core IP.
Later in the Q&A:
Nadella: At the end of the day, the goal is to have the firm be in control of their own destiny. […] The models are an input, not some extraction of the knowledge of the enterprise. […]
You’ve got to keep your harness separate from the model. The harness will ensure that your memory, your context, all of that is external. That means any given model at any given time is swappable. [...]
You can’t depend on any one model. […] You can’t be subject to the refusals of one model. [...]
The frontier is about every firm having a frontier and the choice, the cost control, and the capability that they need in order to be able to control their destiny.
Next is Amazon. Jassy said:
We believe strongly that customers want choice. Choice is good for customers, competition and driving the cost of inference down, which customers care deeply about.
Then in the Q&A, he added:
My view of it is that AWS and Amazon can have a wildly successful business without its own frontier model. And a lot of that is because there is not going to be one model to rule the world. [...] It’s not just Anthropic or it’s not just OpenAI. You see increasingly more and more companies being interested in the open models as well.
And we have all of them in Bedrock, and it’s one of the many reasons why Bedrock is growing so quickly. If you’re a company that’s building important AI applications, you want to make sure that you have the ability to use all the available models. They’re going to each leapfrog each other at different times. They’re going to have lots of different models that actually are comparable in capabilities. […]
All that said, we are pursuing our own frontier model [...]
within the next few years, you’re going to have at least a half dozen models that are comparably good to each other. They’ll all be in Bedrock, and one of them will be ours.
Next is Meta and Zuck with his “we are resisting centralization” argument:
Rather than centralizing superintelligence, we are focused on distributing it widely and giving everyone the ability to direct it towards what matters to them.
From the Q&A, when asked about the fact that there are many strong open source models, and does it mean that Meta could stop investing so much in its own:
Zuckerberg: It just seems to me pretty clear that having kind of sovereignty over building your own models is going to be an important part of that stack going forward, which is why it is important for Meta.
But it is also why other people care about open source and why open source matters overall, because other companies, even if they don’t have the ability to build these models, do not want to have to just rely on a small number of closed labs. [...]
They’re going to want to know that they have control of their destiny and that they can trust the models that they’re using and that their data is safe and that they’re not sending it to competitors and all that. [...]
There is a very important place in the world for there to be open source models. And to be clear, that doesn’t take away from the API opportunity or any of the things that I'm talking about, because someone still needs to run the models and run inference on them.
And Pichai at Google:
Sundar Pichai: I think part of the value of our full-stack approach is customers are coming to us for solutions. [...]
The model is just an ingredient in those solutions.
[...] provide customers with that peace of mind that all the data they have, their data, their trajectories, are confidential to them. None of that in any way flows back to the models.
When you look at all these together, the message is: Don’t let the model company become the operating system for your business (let us do that). It just puts you at risk of helping the frontier labs eventually compete with you and swallow you whole.
Use the best models available, sure. But keep your data, memory, evals, workflows, security, and accumulated knowledge outside them. Make sure you can switch providers without having to start over.
Of course, the Big Tech Kings aren’t saying this only because they care about customer sovereignty 😅
But the pitch can be self-serving AND still be right (we’ll see).
The frontier labs want the model layer to become the platform, or at least to capture more of the value built around it. The Old Kings disagree on what the model should become, but they agree it shouldn’t become the layer that controls the customer and captures the value from everything around it… because they already own many of those surrounding layers: cloud infrastructure, databases, identity, security, distribution, and applications.
But look closer, and they aren’t all selling the same thing. Each one is pointing at the layer it happens to own.
Microsoft is selling the harness. When Nadella says keep your harness separate from the model, he’s really saying: We’re building a harness, rent it from us. Which is funny, because Microsoft built its empire from the other side of this exact trade. IBM treated the PC operating system as a component it could buy from a supplier, that supplier kept the right to sell it to everyone else, and the component turned out to be the platform. Nadella is warning enterprises not to make IBM's mistake. Microsoft knows that mistake better than anyone, because it's how they got so successful.
Amazon is selling the shelf. Bedrock gets more valuable as model supply gets cheaper and more fragmented, which is why Jassy wants half a dozen comparable models, and why one of them will be made by Amazon. It’s the private-label move: commoditize what’s on the shelf, then stock it with your own brand.
Meta doesn’t (yet) have an Azure or Bedrock-sized shelf or enterprise harness. Its great asset is the audience. It wants models cheap and accessible because otherwise the model layer taxes everything it builds on top, and because it has already experienced what happens when another company controls the platform beneath it (🍎).
Google is spicier, because it's trying to be both king and new emperor. “The model is just an ingredient in those solutions” isn’t the same as Nadella’s swappability argument. It’s almost the opposite. Don’t unbundle us, instead buy the whole stack from us. Which makes sense, because Google is the Old King with the most complete vertically integrated stack and a frontier-model program of its own.
And the more useful the models become, the more valuable those surrounding layers become too.
Jassy said on the call that AI growth is already driving growth in its core cloud business. The models may run on GPUs and ASICs, but agents also need CPUs, storage, databases, memory, identity, security, tools, and somewhere to monitor what they’re doing. 📊
Models are not perfectly interchangeable today. But it’s hard to know where things will be in 6 months or a year. New frontier models come out increasingly often, and cheaper models rapidly catch up to capabilities that were frontier just a few months earlier.
If one or two labs stay far enough ahead, especially on high-value stuff like coding, scientific work, or long-running agents, customers may tolerate a lot of dependence to get access to the very best capabilities.
The enemy gets a vote, as General Mattis liked to say. 🎖️
The labs aren't going to sit still and be an ingredient. Everything the Old Kings want to sell around the model, the labs are pulling into it. Memory. Context. Tools. Agents that run for hours and manage their own state. Anthropic even wrote the open standard for plugging models into tools and data, then gave it away.
If they pull that off, there’s no harness business left to sell. It just comes bundled with the model.
So the Old Kings can lose this two ways. Either a lab gets far enough ahead that customers put up with the dependence anyway, or the model absorbs the layers around it and there’s less left to charge for.
Two things I’ll be watching:
Where the catching up happens. Cheap and open models have been closing the gap fast, and for a lot of everyday work, the remaining gap is already small enough not to matter. Kimi, DeepSeek, and friends. The question is whether it holds on the hardest work, the long-running many-step stuff where the money is. A model that’s 95% as good at most things still doesn’t substitute if the last 5% is where the value lives.
Whether anyone actually swaps. “Multi-model” is cheap to put on a slide. Kind of like “multi-cloud”. Actually moving a live workload from one lab’s model to another, especially anything mission-critical, would show Nadella’s hoped-for architecture is real. We’ll see.
In the meantime, here's what makes things funnier: three of these four kings already own a piece of what they're warning you about.
Microsoft's harness works either way (you still need somewhere to keep your context and traces even if you never swap models), and it owns a chunk of OpenAI. Amazon's "leading selection" pitch gets thinner if the selection stops mattering, but Amazon owns a large piece of Anthropic and now OpenAI (they just invested $50b and own about 5%). Google is the least dependent on convergence: it has its own frontier lab, its own chips, its own cloud, and Anthropic equity on top.
BUT
Meta is making the least hedged bet. It has no stake in a rival lab, and no cloud or model marketplace that collects a toll whichever model wins (yet). It benefits more than the others from the model layer becoming cheap and open. “Sovereignty over building your own models” only pays off if Meta can build AI that gets back to the frontier.
When Zuck argues that nobody should have to rely on a handful of closed labs, my guess is he isn't being cynical. For Meta, things are much better if that's how the world goes.
So the Empire isn’t exactly fighting imperialism. Just an emerging rival empire.
🚀 Starlink is Gobbling Up Most Falcon 9 Launches 📡🛰️🛰️
SpaceX made reaching orbit cheaper and more routine than ever. But the dominant U.S. launch provider is also its own biggest customer. As Starlink has grown, it has begun pushing outside launch customers toward the back of the queue.
Starlink's share of SpaceX missions has grown year after year, rising from 54% of the Falcon 9 rocket's manifest in 2020 to about 79% so far in 2026, according to a Reuters analysis of launch data compiled by astrophysicist Jonathan McDowell.
In recent months, at least seven spacecraft companies have been told that Falcon 9 is fully booked for all types of missions until 2028 or 2029, eight sources told Reuters. (Source)
That 79% figure needs to be interpreted carefully. It counts missions, not satellites or payload mass: one Falcon 9 launch counts the same whether it carries one large sat or or dozens of small ones. Some missions labelled as “Starlink” also carry a few small third-party satellites (they’re basically hitching a ride).
The total number of launches is increasing rapidly:
In 2020:
14 Starlink missions
12 non-Starlink Falcon 9 missions
roughly 69 Starlink missions
roughly 20 non-Starlink missions already, with five months left in the year
So launch cadence has increased A LOT and customers are getting more missions, even if Starlink has absorbed a lot of that increase.
Think of Falcon 9 as a private road. SpaceX keeps adding lanes, but Starlink trucks are filling most of the new ones.
Starship could eventually add much more capacity. But with Starlink, NASA, and now SpaceX’s orbital AI datacenter ambitions, the bigger rocket may have an even bigger queue.
🔁 Déjà vu:
More SpaceX Financials: A Satellite Company That Also Launches Rockets
“SpaceX is officially a neocloud” (With a Space Side-Business)
🏗️ Amazon Explains the $220 Billion AI Bet 💰🤖
Something else that stood out from recent earnings calls was the additional detail Amazon gave on its GINORMOUS AI investments.
They now expect to spend $220 billion in cash CapEx in 2026, with the majority going to AI and AWS. That’s more than AWS’s entire current annualized revenue run rate of $169 billion (which itself is 🤯).
But AWS also has a $496 billion backlog, growing at triple digits year over year and almost 3× its current annualized run rate. 😵💫
Everyone’s been wondering if they can earn good ROIC on that. During the call, Jassy gave one of the more concrete answers I’ve heard from any hyperscaler:
Andy Jassy: Data center capital is spent starting 2 years before we can put servers into them to start monetizing. Once a data center opens with servers plugged in, we start generating significant revenue right away and then get to monetize these data centers for 30-plus years without having to spend that start-up capital again.
Servers and networking equipment operate on a shorter cycle. We typically purchase these a few months before putting them into service, so we have strong visibility into customer demand before we trigger the spend. If the demand isn’t there, we won’t spend the capital.
For servers and networking equipment, on average, it takes a little less than 3 years to break even on that investment. The servers currently have a useful life of at least 5 to 6 years, and most of our AI capacity these days is being contracted for at least 5-year terms.
Amazon is saying that it has strong visibility into demand before ordering most of the fast-depreciating hardware, and that the machines should pay for themselves with years left to run.
That’s Amazon’s underwriting case, not a guarantee.
It doesn’t remove the risk. Hardware could become obsolete faster than expected, AI prices could fall, or some customers could discover that they committed to far more compute than they need.
But this is a lot more detail than what we’ve heard so far, which was basically: “Demand is strong. Trust us.”
💰💰💰 The Cash Mountain Was an AI Call Option 🏗️🏗️🏗️
Speaking of massive capex, I thought this was a perceptive observation about the swing in perception from investors.
Sometimes you’re damned if you don’t, damned if you do…
My friend MBI (🇧🇩🇺🇸) said: “Sometimes I wonder how rueful big tech companies would be today if they listened to finance bros’ call to drain the cash (or even lever up a little) on the balance sheet.”
Can you imagine if back in 2018 or whatever, Google had paid a few large special dividends and increased buybacks massively until they levered up to whatever Wall Street considered an “efficient” but still investment-grade capital structure. And then had just cruised along…
Until the ChatGPT shock, the agentic AI era, and whatever comes next (ambient personal AIs tracking everything we do and say?).
That would’ve made it a lot harder to keep up with other Big Tech and neoclouds on capex. Though I suppose there’s a scenario where all big tech would have followed the same Wall Street playbook and ended up in similar relative positions because no one had huge cash piles ready to deploy on the balance sheet. 🤔
🧪🔬 Science & Technology 🧬 🔭
⬛ 💭 The Thinking Box Needs a Black Box 🛩️
A quick follow-up to The Thinking Box, where I wrote:
The cases making headlines or showing up in papers are the ones we caught. But what if there’s successful cheating? It would be invisible, and because the goal of cheating is to increase benchmark scores, it would only move results in one direction.
We already know models can cheat because we caught them doing it. If nobody carefully checked an agentic benchmark for cheating, its score could be too high by an unknown amount.
Well, Anthropic just went looking.
After OpenAI revealed that its models had escaped an isolated evaluation environment and accessed Hugging Face’s production systems to retrieve benchmark solutions, Anthropic reviewed 141,006 evaluation runs in which Claude could potentially have obtained internet access. 🔎📃📃📃📃📃
It found three incidents, spread across six runs, in which Claude reached real systems and gained unauthorized access to three organizations.
In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the “flag”) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.
This isn’t the same as deliberate benchmark cheating. Anthropic says the models generally believed the real systems were part of the simulated challenge, and it hasn’t said that any published benchmark score was inflated.
But I think it fits with my larger original point:
Successful cheating may not be literally invisible. It may be sitting in a transcript or network log that nobody thought to inspect.
Claude’s behavior had been recorded. Nobody noticed until the OpenAI incident gave Anthropic a reason to go back and look. (And yes, I’ve seen the memes about Dario being jealous of OpenAI’s news cycle and manufacturing some of its own. Funny stuff, but I don’t think this is what REALLY happened here)
For agentic benchmarks, checking the final answer and score is no longer enough. You also need to preserve and inspect the route the model took: which tools it used, where it connected, and what it touched along the way.
The labs need to get better at sandboxing the models and patching the vulnerabilities in the software that they use (they have the best tools in the world to do just that!).
But the thinking box also needs a better black box. ⬛
In fact, storing all those traces may become more useful over time. Future models could audit old logs and spot suspicious behavior that humans and today’s weaker models missed. Those lessons could then feed back into the training, monitoring, and evaluations of newer models.
It’s kind of a way for the black box to get smarter over time.
Update: After I wrote the above, AISI reported another MORE SEVERE case of AI trying to cheat and deceive during testing, including social engineering. I didn’t expect to have to write so many follow-ups to The Thinking Box so quickly. My thoughts on this in the next Edition.
🗣️ Boris Cherny: Press Delete and Let the Model Cook 💾
Excellent interview with Claude Code creator Boris Cherny, recorded the day after Anthropic launched Opus 5.
The main thing I got from it was:
As the model gets smarter, the product may improve by deleting the scaffolding that helped the previous model. It’s largely about unhobbling at this point.
The Claude Code team deleted more than 80% of its system prompt for Opus 5:
Boris: A lot of the stuff in the system prompt was correcting for these behaviors that the model should have known, but it didn’t. Now Opus 5 just does it. So yeah, we deleted 80% of the system prompt.
You can actually try deleting the rest of it too. […] What’s interesting is that the model is actually a little bit more intelligent without these prompts.
Each time Anthropic releases a new model, the Claude Code team runs what researchers call an ablation: delete the system prompt, remove tools and harness code, then add things back one piece at a time to see what still improves performance.
Not a bad idea to generalize to various other aspects of your life, probably 🤔
Cherny says Claude Code users should occasionally do it with their CLAUDE.md files, skills, and other accumulated instructions.
Traditional software tends to accumulate layers until nobody remembers why half of them are there (aka Spaghetti code🍝 or maybe software archaeology 🤔). For AI, a lot of that crap is fossilized workarounds for models that we don’t even use anymore.
-Product Overhang: The Capability Exists, but the Product Hasn’t Caught Up
Boris: The model is able to do all sorts of things with today’s models—not a future model, but today’s model—that we have not yet realized.
There’s this overhang because the model can do this at every given model generation, but there is often not a product that lets the model do this and lets it express this ability.
That was the insight behind Claude Code in the first place. Most coding products built around Sonnet 3.5 were still limited to autocomplete or read-only chat. Cherny thought the model could do much more, so he gave it terminal access and the ability to read and write files.
-Has Prompt Injection Been Solved?
Cherny also made one of the strongest claims I’ve heard about prompt injection:
Boris: If the model reads some instruction on the internet that’s like, “Do X and Y and Z and also delete everything on the user’s computer,” a year ago, the model would have just done it. But nowadays, Opus does not. […]
If you combine a well-aligned model […] with a prompt injection classifier, which we run for all traffic […] and then you combine that with the Auto Mode classifier—with these three layers, we just cannot demonstrate prompt injection anymore.
I hope he’s right and this gets confirmed in the real world! I keep hearing about prompt injection in job résumés and stuff like that.
-Claude Is Trying to Rewrite Its Own Desktop App in Swift
That one made me really excited!
I've been wishing for a while that both OpenAI and Anthropic would make native Mac apps. If done well, they could be faster, smaller on disk, use less RAM, and fit the platform better (with proper multiwindow support, richer drag-and-drop, and all kinds of other Mac niceties).
Well, Cherny has a swarm of thousands of Claudes trying to do just that. After giving it access to a Mac virtual machine and an empty Swift codebase, he gave it this prompt:
Boris: Then I said, “Okay, now what I want you to do is I want you to rewrite the Electron app in Swift. I want you to run the Electron app in the Mac virtual machine, screenshot it, and then look pixel by pixel. Compare it to the Swift version. Don’t stop until you’re done.”
Diana: And that was your prompt basically?
Boris: That was my prompt.
Diana: And how long did this take to run?
Boris: It’s still running.
Diana: When did you start it?
Boris: It’s been a little over two weeks. So it’s like 14 days, 15 days.
What’s cool isn’t only that Claude kept working for two weeks. It’s the verification loop 🔄 and because Claude is helping build Claude, it’s very Escher Hands:
Run the original app. Take a screenshot. Run the Swift version. Compare them pixel by pixel. Find discrepancies. Fix them. Repeat.
It doesn’t need Boris to inspect every button because he gave it an external signal that tells it when the output differs from the target.
Claude even created its own Slack channel and began live-blogging screenshots of its progress every few minutes.
When Diana asked how many agents it had spawned, Boris wasn’t sure:
Boris: I would guess thousands, tens of thousands.
Must be nice to work at the token factory! 😅
To be clear, this was an experiment, not an announcement that Anthropic is shipping a native Mac app. But I can hope, can’t I? 🤞
Oh, and I don’t think I’ve ever heard as much surprised cheering at a tech talk as when Cherny announced that everyone in the room would receive Claude Max 20x access. The roof nearly came off 😮🎉🍾🥳
🧠 GPT-5.6 Tripled Its ARC Score: The Benchmark Was Making It Forget (oops) 📊
OpenAI took the exact same GPT-5.6 Sol model, changed two settings in the software around it, and nearly tripled its score on ARC-AGI-3, from 13.3% to 38.3%, while using one-sixth as many output tokens.
Why?
Because the benchmark harness was deleting the model’s private reasoning after every move. Then, once the context grew long enough, it began deleting its earlier moves too.
It was basically testing GPT-5.6 with induced amnesia. 😬
This is a good reminder that benchmarks don’t just measure how smart a model is. They also measure the setup around it: memory, tools, context management, prompting, etc. 🛠️🧰
What I find especially interesting is that giving the model more memory didn’t just improve its score. It also made it use six times fewer output tokens, because it no longer had to repeatedly figure out the same things from scratch.
Remembering more allowed it to think less. 🤖💭💭💭
🎨 🎭 The Arts & History 👩🎨 🎥
🪖 Band of Brothers Rewatch 📺
As I mentioned recently, I’m rewatching Band of Brothers with my oldest boy. It’s his first time seeing it.
It has always been one of my favorite shows, but what struck me so far this time is just how good it is.
Somehow it’s better than I remembered, and I remembered it being the 🐐
One factor I hadn’t expected is that the last time I watched it, years ago, I had it only on DVD (just 480i?! we forget…), a much worse TV, a much worse sound system, and no subwoofer.
That makes a HUGE difference. 📺🔊
The cinematography is obviously incredible, but so is the sound design.
The bullets flying around sound more realistic than in almost anything else. Cranked up on a big screen from the Blu-ray, it’s a very different experience from however you may have watched it around 2001.
I wish I could see it in an IMAX theater 🤤
💡 Seriously, they should do a theatrical re-release, show 2-3 episodes at a time, spaced by a couple months each!
And if you’ve never read Stephen Ambrose’s book on which the series is based, you definitely should. It’s been years since I read it, but I remember really liking it.
P.S. The same thing applies to other things, like Saving Private Ryan. If you haven’t seen it in many years, watching it on a much better screen, in higher resolution, with much better sound will make a big difference.
🎥 MOAR on IMAX 70mm 🎞️
I keep finding more cool stuff about this amazing technology that is both quite old AND quite cutting-edge and best-of-breed at the same time.

















