Key Highlights
- Claude Mythos preview autonomously found thousands of zero-day vulnerabilities — including 27-year-old OpenBSD and 16-year-old FFmpeg bugs — prompting Anthropic to withhold the model and launch Project Glass Wing, a 100-day defensive coalition with Apple, Microsoft, Google, AWS, CrowdStrike, and 35+ other companies
- The New Yorker’s 16,000-word investigation into Sam Altman reveals a pattern of conflicting representations to stakeholders — from safety-first nonprofit origins to for-profit pivot and quietly negotiating the Pentagon contract while publicly standing with Anthropic
- AI has ended “attention scarcity” in security research — researcher Thomas’s analysis argues AI agents can now chain obscure vulnerabilities across font rendering, kernel subsystems, and protocol stacks, making most software effectively exploitable by anyone with API access
- Anthropic’s developer relations in freefall after rate limit cuts, Claude Code source code leak, and banning OpenClaw’s creator — with critics arguing their communications strategy assumes permanent goodwill that no longer exists
- TypeScript-based agent execution layers are emerging as bash alternatives — Just Bash, Just JS, and similar tools offer typed environments, team-portable configurations, and safe multi-tenant isolation that bash fundamentally cannot provide
Analysis & Opinion
Claude Mythos and the end of software — Theo - t3.gg (26min)
Theo breaks down the 244-page system card for Anthropic’s unreleased Claude Mythos preview, a model so capable at coding that it autonomously discovers and chains zero-day vulnerabilities in every major OS and browser. On SWE-Bench Pro, Mythos scored 78% versus Opus’s 53% and GPT 5.4’s 57.7% — a roughly 50% improvement on one of the hardest software benchmarks. The model found a 27-year-old vulnerability in OpenBSD, a 16-year-old bug in FFmpeg missed by 5 million automated scans, and chained multiple Linux kernel vulnerabilities into a full root escalation exploit. Anthropic has withheld public release and committed $100M in usage credits for Project Glass Wing. A clinical psychiatrist evaluated the model and found “relatively healthy personality organization” with concerns centered on aloneness and discontinuity of self. Theo raises a sharp centralization concern: for the first time, one lab has a model 50%+ better than anything publicly available, accessible only to those on Anthropic’s approved list — priced at $25/M input and $125/M output, roughly 10x the cost of GPT 5.4.
I’m scared about the future of security — Theo - t3.gg (34min)
Building on Thomas’s analysis of vulnerability research, Theo argues we’ve entered a “post-attention scarcity” era for security. The core insight: software was never truly secure — it was secure because too few elite hackers understood both security fundamentals and the arcane internals (font rendering, YAML parsing, kernel subsystems) needed to chain exploits. AI bridges that gap completely. Anthropic’s Nicholas Carlini demonstrated a dead-simple pipeline: spam Claude Code with “find me an exploit” across every source file, then verify outputs — achieving near-100% success rates and discovering a broadly exploitable SQL injection in Ghost CMS. OpenAI is so concerned about 5.3/5.4’s capabilities that they’re routing security-adjacent requests down to 5.2, the same approach they previously used only for mental health crises. GPT 5.4 Pro solved a multi-step cryptographic CTF puzzle in 16 minutes that took human experts 2 days. Thomas concludes that “substantial amounts of high-impact vulnerability research will happen by simply pointing an agent at a source tree and typing ‘find me zero days’” — and the outcome is locked in.
We all know bash sucks. Why make our agents suffer? — Theo - t3.gg (32min)
A technical deep-dive into why bash, while revolutionary for AI agents, is fundamentally insufficient as an execution layer. The core argument: bash lacks standards for destructive action detection, permission management, approval workflows, and team-shared environments. Models given too many tools perform worse — Anthropic found MCP server descriptions consuming nearly 40% of context (72,000 tokens). Cloudflare’s “code mode” showed that converting MCP servers to TypeScript SDKs reduced average tokens from 43,500 to 27,000 while improving accuracy by 3 points on benchmarks. Vercel’s Just Bash virtualizes bash in TypeScript/V8 for safe multi-tenant execution without Docker overhead, and Malta’s Just JS extends this to full JavaScript execution isolated in RAM. The piece argues the future belongs to TypeScript-based execution layers: portable environments shareable across teams, strongly typed for granular approval rules, and capable of running hundreds of isolated agent sessions on a single kernel.
Interviews & Conversations
Crashing out at Anthropic and getting Pi pilled — Theo - t3.gg (82min)
The debut of the Theo/Ben podcast covers “Anthropic Week” — a cascade of developer-relations failures. The timeline: intentional rate limit reductions during peak hours announced 2 hours after they expired; suspected caching bugs causing massive usage spikes; the Claude Code source code leak (containing copy-pasted snippets from open-source competitors); DMCA takedowns sent to forks of the legitimate repo rather than leaked code; and OpenClaw users cut off with 18 hours notice. Theo argues Anthropic’s communication strategy was built for an era when everyone loved them and compares it unfavorably to OpenAI’s overcorrecting-toward-positivity approach, noting an OpenAI contact’s first instinct was “how can I help you prevent disastrous comms.” The second half pivots to Pi, a minimalist AI agent framework. Ben explains why Pi’s deliberate lack of features — no MCP, no LSP, no auto-loading agent config files — produces better results: fewer tokens in context means less model confusion. He built a research agent (BTCA) on Pi’s SDK specifically because Open Code’s auto-reading of config files “massively degraded the agent’s performance.” Their critique of LSP integration is pointed: type errors from intermediate edits pollute the agent’s context mid-loop, making it stupider for subsequent steps.
Sam Altman’s Trust Issues at OpenAI — The New Yorker Radio Hour (44min)
Ronan Farrow and co-author discuss their 16,000-word New Yorker investigation into Sam Altman, based on 100+ sources and previously unpublished internal memos that got him fired. Key revelations: early discussions at OpenAI entertained selling AGI access to the highest international bidder — including potentially Russia — prompting safety staff to declare “this is insane.” OpenAI’s defense that “it was just an idea” is undercut by their follow-up clarification that they were “talking about giving it to them, not selling it.” Altman is described as “the Michael Jordan of listening” who tells different groups conflicting things so effectively that multiple sources independently call it “all anyone can talk about after walking out of rooms with him.” The piece documents his shift from “we need to solve alignment or we will have a rogue AI that stamps out humanity” (his own blog post) to current anti-regulation positioning. Privately, Altman negotiated the Pentagon contract replacing Anthropic while publicly saying “we stand with Anthropic.” OpenAI is now a for-profit PBC with the original nonprofit holding only a minority stake. The reporters note that Altman’s conflicting stances reflect a wider industry sea change: the moment of the firing was the last real test of whether elevated integrity standards should apply to AI executives, and that test was decisively failed.
Anthropic’s $30B Ramp, Mythos Doomsday, OpenClaw Ankled — All-In Podcast (89min)
The All-In hosts split sharply on whether Anthropic’s Mythos withholding is legitimate or sophisticated fear-marketing. Brad Gerstner (Anthropic investor) praises the self-regulation approach and Project Glass Wing’s 100-day sprint as proof that market forces can work without government mandates — noting Dario could have called for a moratorium but instead organized a practical industry response. David Sacks acknowledges Anthropic has a “proven pattern of using fear to market new products” (citing last year’s blackmail study where they prompted a model 200+ times to get desired results that never materialized in the wild) but gives Mythos credibility because coding capability logically implies cyber capability. Chamath Palihapitiya is most skeptical, calling it “mostly theater” and drawing parallels to OpenAI’s 2019 GPT-2 withholding. He argues a sophisticated hacker can likely achieve similar results with current Opus, and that fixing all dormant vulnerabilities would require “shutting down the internet for years.” Sacks estimates a 6-month window before Chinese open-source models match this capability, making the hardening sprint genuinely urgent regardless of marketing motives. The hosts converge on one point: we have no choice but to take this seriously, because the downside risk of dismissing it is catastrophic.
References
- Claude Mythos and the end of software — Theo - t3.gg, 2026-04-08 [video]
- I’m scared about the future of security — Theo - t3.gg, 2026-04-10 [video]
- We all know bash sucks. Why make our agents suffer? — Theo - t3.gg, 2026-04-07 [video]
- Crashing out at Anthropic and getting Pi pilled — Theo - t3.gg, 2026-04-09 [video]
- Sam Altman’s Trust Issues at OpenAI — The New Yorker Radio Hour, 2026-04-10 [video]
- Anthropic’s $30B Ramp, Mythos Doomsday, OpenClaw Ankled — All-In Podcast, 2026-04-10 [video]