Key Highlights
- AI has collapsed the software-security disclosure model. Theo (t3.gg) argues that CopyFail (and its CopyFail2/Dirty Frag descendants), an unprivileged Linux LPE, a single-
git pushGitHub.com RCE found by Wiz, and 84 compromised Tanstack npm packages in a single week prove that frontier models can now read patch diffs, infer the vulnerability, and write the exploit before distros ship the fix — the 90-day embargo is effectively dead. Jeff Kaufman’s experiment of handing the CopyFail2 fix-diff to Gemini 31 Pro, GPT-5.5 Thinking and Claude Opus 4.7 (all three flagged it as a security patch from the diff alone) is the empirical kill-shot for “patch-to-exploit is hard.” - OpenAI launches Daybreak. Buried in the same Theo segment: OpenAI announced Daybreak, a request-based vulnerability scanning service that runs your codebase through
5.5-cyber(a non-public hardened variant) to find issues “before they ship” — the first major lab move to put a frontier model behind a defender-only API to rebalance the cat-and-mouse asymmetry against attackers who already have open-weight options like Kimmy K26. - El Niño 2026 is a measurable food-security event, not a forecast. All-In Podcast’s David Friedberg warns ocean temperatures running 4°C above normal hold ~11 million terawatt-hours of excess energy heading into the Northern-Hemisphere summer — putting Indian, Brazilian, Australian and Southeast Asian monsoon crops at risk and threatening caloric deficit for ~1.5 billion people who depend on those rains, with 150M Indian farmers on the front line.
- Benioff: “Not my first SaaSpocalypse.” The Salesforce CEO frames the AI-driven SaaS rerating as a market mood swing, not an existential one, while disclosing Salesforce expects to spend ~$300M/year on Anthropic tokens — a useful data point for how big “we use a frontier model” looks at hyperscaler-customer scale.
- A serious case for AlphaGo as the right scale to study reasoning. Eric Jang (ex-1x, ex-DeepMind Robotics) rebuilt AlphaGo from scratch on sabbatical and tells Dwarkesh Patel that the open question — how a 10-layer net amortises a search tree previously thought intractable — is the cleanest small-budget proxy for what LLM “thinking” actually is.
Interviews & Conversations
Everything is pwn’d now — Theo - t3.gg (34 min)
Theo Browne argues that the past week of disclosures — CopyFail (a 732-byte Python script that root-escalates on every major Linux distro running kernel 6.x or 7.x), CopyFail2, Dirty Frag, a Slab-memory breakout, a Mythos-discovered curl bug, a Wiz-disclosed RCE on github.com via a single git push, and the 84-package Tanstack npm supply-chain compromise (with 121 additional compromised packages found across those names) — is not a bad news cycle but evidence that three assumptions underwriting open-source security have collapsed simultaneously: that only well-paid experts find exploits, that the 90-day embargo gives defenders a usable lead, and that going from a silent patch to a working exploit is hard. The empirical kill-shot is Jeff Kaufman’s test where Gemini 31 Pro, GPT-5.5 Thinking and Claude Opus 4.7 all flagged the CopyFail2 fix-commit as a security patch — two of three did so even with the commit message stripped — meaning any bot can now monitor kernel commits and produce working exploits in the window between merge and distro shipment. The proposed response is structural: a new “trusted actors” disclosure tier (paid certification for distro maintainers and large IT shops to receive embargoed details earlier), a rethink of open-source publishing that allows staged-private patch windows on platforms like GitHub, and at the personal level, treating every system as already compromised and reorienting backup strategy from “prevent leaks” to “survive ransomware-style destruction” — including offline air-gapped Synologys, drives mailed to family, and explicit safe-words to defeat voice-cloned social-engineering calls. The piece also flags OpenAI’s Daybreak announcement as the first frontier-lab move to put a defender-only model (5.5-cyber, non-public) behind a vulnerability-scanning API. Theo is open about the falsifier — if CVE volume drops sharply over the next month, this was just five years of lowhanging fruit found in three weeks — but notes that current trajectory points the other way. Pairs directly with the ArXiv enforcement story from yesterday’s digest: both are institutions trying to re-establish accountability in a world where frontier models removed the cost barrier to certain kinds of bad output.
Building AlphaGo from scratch — Eric Jang — Dwarkesh Patel (2h37m)
Eric Jang (formerly VP of AI at 1x Technologies and previously Google DeepMind Robotics) spent a sabbatical reimplementing AlphaGo from scratch and uses the conversation to argue that Go remains an unusually clean substrate for studying what “reasoning” means inside neural networks. The core mystery he keeps returning to: a 10-layer policy/value net amortises a tree search whose raw branching factor was considered computationally hopeless less than a decade ago, and we still lack a good theoretical account of how. He walks through the scaling curve across AlphaGo / AlphaGo Zero / MuZero and the open-source KataGo (David Wu) line, and proposes that the MCTS-as-thinking primitive may be a more direct analogue to current LLM chain-of-thought and tool-use loops than the community treats it as. The discussion also touches on whether DeepMind’s game-playing heritage helped or hindered its LLM trajectory, the difficulty of building outer verification loops for self-improving AI, and Jang’s pitch for “minimal-budget” research: serious work on Go, MCTS, and reasoning is still possible with very small compute, and the transferability of those insights to harder targets like drug-discovery automation is underrated by labs chasing frontier scale.
Trump-Xi Summit, Benioff on the SaaSpocalypse, OpenAI vs Apple, Multi-Sensory AI, El Niño — All-In Podcast (1h17m)
The Trump-Xi summit segment lands on a cautiously optimistic framing — the panel argues that selling more advanced semiconductors into China actually reduces conflict risk by deepening technological interdependence, and that Taiwan’s strategic premium will erode as both the U.S. and China build out domestic leading-edge fabrication. Mark Benioff drops in to dismiss the “SaaSpocalypse” thesis as a normal market rerating, noting Salesforce expects to spend roughly $300M/year on Anthropic tokens while cash flow and customer-success metrics remain strong — a concrete number for what hyperscaler-customer AI spend now looks like at the application layer. The OpenAI-Apple segment details a deteriorating partnership: poor Siri integration, missed revenue targets, growing privacy friction, and Apple watching OpenAI’s hardware ambitions encroach directly on its turf; the panel reads this as the opening of a multi-sensory model era where pure-text LLMs lose share to systems that observe the user’s screen and physical environment directly. The most consequential segment is David Friedberg’s climate warning: a record-breaking El Niño has loaded ocean temperatures ~4°C above normal — about 11 million terawatt-hours of excess energy — heading into Northern-Hemisphere summer, threatening simultaneous monsoon disruption across India, Brazil, Australia and Southeast Asia, with 150M Indian farmers and 1.5B monsoon-dependent people facing measurable caloric-deficit risk through late 2026 and into 2027. Pairs with the ongoing “AI haves and have-nots” framing from yesterday’s digest — global inequality stories and frontier-AI capital concentration are converging onto the same 12-month horizon.
References
- Everything is pwn’d now — Theo - t3.gg, 2026-05-15 [video]
- Building AlphaGo from scratch – Eric Jang — Dwarkesh Patel, 2026-05-15 [video]
- Trump-Xi Summit, Benioff: Not My First SaaSpocalypse, OpenAI vs Apple, Multi-Sensory AI, El Niño — All-In Podcast, 2026-05-15 [video]