Key Highlights

  • Microsoft and OpenAI’s exclusivity deal is effectively dead — the amended partnership announced last week strips Azure of its sole-cloud-provider status, lets OpenAI ship to AWS (which just signed a $100B/8-year compute extension on top of an existing $38B deal), and quietly loosens the AGI-trigger clause that would have ended the IP-sharing relationship. Theo’s read: Microsoft got 2032 IP rights as a consolation; OpenAI got everything else, especially the right to chase Anthropic in Bedrock-dominated enterprise accounts. The structural reason this matters: most enterprise AWS startup credits cannot be spent on Anthropic models — the rev-share Anthropic locked in with the hyperscalers is too expensive for AWS/Google to subsidize — which is the under-discussed reason Anthropic’s enterprise revenue is growing faster than OpenAI’s despite ostensibly weaker code models. Putting OpenAI on Bedrock attacks that moat directly.
  • Two studies converging on the same finding: LLMs collapse the variance in human writing. A new Lobsters-surfaced research project (“How LLMs Distort Our Written Language”) shows LLM-edited essays drift toward a common region of semantic space, become more neutral on argument stance, increase formality while reducing personal pronouns, and simultaneously amplify both emotional and analytical vocabulary — distortions human editors don’t make. Most striking: the same homogenization shows up in ICLR 2026 peer reviews, where AI-generated reviews systematically over-weight reproducibility/scalability and under-weight clarity/relevance. Users prefer the assisted output and still report “a statistically significant loss of voice and creativity” — the satisfaction signal is decoupled from the quality signal.
  • A 2024-era model is already outperforming attending physicians at ER triage. Harvard’s Science paper on o1-preview across 76 real ER cases: 67.1% diagnostic accuracy vs. 55.3% and 50.0% for two attendings, and in one case the model flagged a flesh-eating infection in a transplant patient 12–24 hours before the treating doctor caught it. Working from raw EHR text only, no extra context. Reviewers couldn’t distinguish AI from physician-generated assessments. The implication that hangs over the result: this is o1-preview, not a frontier model — the question isn’t whether AI will be deployed in formal patient care, it’s how long the regulatory lag holds.

Analysis & Opinion

Do AI summaries hurt critical thinking? — Blueprint for Disaster (via Lobsters)

The piece argues that AI summarization is a categorically different shortcut from older ones (CliffsNotes, Wikipedia, the Seinfeld-grade “just watch the movie”), because it removes the cognitive friction that those workarounds preserved. Reading a summary is still reading; having an LLM compress a piece you never opened is something else. The author’s frame — “machines generate content that other machines then condense” — lands harder when paired with today’s other research finding that LLM-edited prose collapses toward a homogenized semantic center. The implicit chain is uncomfortable: if generation flattens variance and summarization removes engagement, the medium-term failure mode isn’t bad reasoning, it’s no reasoning at all on the consumer end of the pipeline. The piece doesn’t offer a prescription, which is fair — the problem isn’t AI summaries, it’s that the demand for them is real and growing.

Research

How LLMs Distort Our Written Language — Lobsters

Multi-pronged study finding that LLM revisions move text toward a common semantic region in a way human editors don’t, increase formality while suppressing first-person voice, and inflate both emotional and analytical vocabulary simultaneously. Users self-report reduced voice and creativity even when satisfied with output. A separate analysis of ICLR 2026 peer reviews shows AI-generated reviews systematically weight reproducibility and scalability higher than human reviewers do, while underweighting clarity and relevance — meaning the homogenization effect leaks from drafting into evaluation, and could already be reshaping which papers get accepted at major ML venues.

AI shows its skills in the emergency room — The Rundown

Harvard study in Science: OpenAI’s o1-preview evaluated on 76 real ER cases reached 67.1% diagnostic accuracy vs. 55.3% and 50.0% for two attending physicians, working only from raw EHR text. In one case the model flagged a rare flesh-eating infection in a transplant patient 12–24 hours before the treating physician. Two physician reviewers couldn’t reliably tell AI from human assessments. The benchmark predates o1’s public release and uses none of the agentic scaffolding that has since become standard, which makes it a conservative floor rather than a ceiling on current capability.

Interviews & Conversations

Microsoft and OpenAI break up (Amazon is pumped) — Theo - t3.gg (37m)

Theo walks the full arc of the Microsoft-OpenAI relationship from the original 2019 $1B deal through the amended 2026 agreement, and the structural takeaway is that Sam Altman fleeced Microsoft. Microsoft’s only meaningful gain in the renegotiation was extending some IP rights from 2030 to 2032; it lost cloud exclusivity (OpenAI products now ship on any cloud), it lost rev-share from OpenAI (replaced with a profit-share, which Theo notes is effectively zero given OpenAI’s losses), and the AGI-trigger clause that would have terminated IP-sharing is now arbitrated by an “independent expert panel” — a face-saver after both sides realized the original AGI definition was unenforceable. The friction reportedly started in late 2024 when Sam yelled at Mira Murati on a video call for not handing over o1’s chain-of-thought research to Microsoft, and Nadella allegedly fumed at his 1,500-person research org for being lapped by OpenAI’s 250 people. The deeper story is the Anthropic angle: Theo argues Anthropic’s faster enterprise growth (~$30B run rate, vs. OpenAI’s larger but slower-growing book) is almost entirely a cloud distribution story — Anthropic is in Bedrock, AWS startup credits can’t be spent on Anthropic models because of the rev-share they negotiated, but enterprises already on AWS find it dramatically easier to deploy Claude through Bedrock than to broker a separate Microsoft-Azure relationship for OpenAI. Putting OpenAI on AWS (and likely GCP soon) directly attacks that moat. Side detail with broader implications: OpenAI is committing to 2 GW of Trainium 3/4 capacity, the first time it has trained or served on non-NVIDIA silicon — Theo flags this as a possible quality risk, citing his ongoing thesis that Anthropic’s model degradation correlates with their own Trainium migration. The video also includes a remarkable anecdote about Theo’s azure.t3.gg benchmark (o1-preview running 2.2× slower on average and 15× slower in worst case on Azure than on OpenAI’s own endpoints) that Microsoft fixed within a day after he threatened to repost it — illustrating the operational gap that made Azure exclusivity untenable to begin with.


References

  1. Blueprint for Disaster, “Do AI summaries hurt critical thinking?,” Medium / Lobsters, 2026-05-04 [blog]
  2. “How LLMs Distort Our Written Language,” Lobsters, 2026-05-04 [blog]
  3. The Rundown, “AI shows its skills in the emergency room,” 2026-05-04 [blog]
  4. Theo - t3.gg, “Microsoft and OpenAI break up (Amazon is pumped),” YouTube, 2026-05-04 [video]