AI Daily Digest — 2026-05-10

Key Highlights Nvidia has committed over $40B in equity stakes to AI companies in the first four months of 2026 — and analysts are openly calling it a circular-investment problem. $30B went to OpenAI alone; another seven multi-billion deals into publicly-traded suppliers (Corning $3.2B, IREN $2.1B) plus ~24 private rounds on top of 67 from 2025. Wedbush’s Matthew Bryson labels it “squarely into the circular investment theme” — Nvidia funding its own customers to buy Nvidia GPUs. Worth holding next to the Cloudflare/Oracle layoff stories from earlier this week: the AI capex flywheel is now visibly self-financing at the supplier level, while the productivity story at the customer level is being used to justify headcount cuts. The risk concentration here is structural, not cyclical. OpenAI published the first detailed look at how it actually runs Codex agents in production — and the answer is a surprisingly heavy security harness. Sandboxing, multi-tier approval gates, network egress policies, and agent-native telemetry. The piece is notable mostly because the disclosure pattern itself is new: until now, the running-AI-agents-safely conversation has been mostly external (red-team papers, regulator white papers). OpenAI describing its own internal controls reads as a deliberate move to set the de-facto standard before regulators write one. Useful read alongside Jeff Kaufman’s vulnerability-disclosure piece from yesterday — the embargo equilibrium is shifting in both research and deployment. Tilde’s Aurora optimizer claims 100x data efficiency over Muon on 1.1B-parameter training, and the diagnosis explains a known failure mode rather than just beating a benchmark. Muon inherits row-norm anisotropy on tall matrices, causing rows with initially small gradient norms to keep getting small updates — a self-reinforcing feedback loop that permanently kills MLP neurons. Aurora reformulates the steepest-descent step under a joint constraint of row-norm uniformity and orthogonality. State-of-the-art on the modded-nanoGPT speedrun (3,175 steps), MMLU up ~10 points over Muon. If the result holds up at scale, it’s the rare optimizer paper where the mechanism, not just the curve, is the contribution. Wispr Flow says India is now its fastest-growing market — meaningful because the linguistic surface there is the hardest the company has tackled. Hinglish (mixed Hindi/English with code-switching), Android-first launch, planned tier expansion to reach beyond white-collar users. The thesis: voice notes and voice search are already the dominant input modality in India, so a working voice-input layer becomes a general computing surface, not a per-app convenience. Watch this against Western voice-AI assumptions, which still treat voice as an accessibility/hands-free fallback. Analysis & Opinion Nvidia has already committed $40B to equity AI deals this year — TechCrunch By the end of April, Nvidia had publicly committed over $40B to AI-company equity in 2026. The headline number is dominated by the $30B OpenAI stake, but the supporting deals are where the circularity becomes visible: $3.2B into glassmaker Corning, $2.1B into data-center operator IREN, and roughly two dozen private-startup rounds on top of the 67 Nvidia participated in during 2025. Wedbush analyst Matthew Bryson called the pattern “squarely into the circular investment theme” — Nvidia is increasingly funding the buyers of its own GPUs, which compresses the audit trail between Nvidia’s shipped revenue and end-customer demand. Bryson hedges that this can build “a competitive moat” if execution holds, but the read across the ecosystem is sharper: the AI capex story is increasingly self-financed at the supplier layer, and stress-tests of demand will be obscured for as long as the funding flows continue. Worth filing alongside this week’s Oracle and Cloudflare layoff-with-record-revenue stories — the productivity narrative at the customer end and the equity-stake narrative at the supplier end are being told as one continuous story, but the failure modes are very different. ...

2026-05-10 · 7 min · Kun Lu

AI Daily Digest — 2026-05-09

Key Highlights Cloudflare cut 1,100 jobs (~20% of headcount) the same quarter it posted record $639.8M revenue (+34% YoY) — and CEO Matthew Prince openly said AI is the cause, not cost-cutting. Internal AI use jumped 600% in three months, the entire R&D team is on Workers + AI, and autonomous agents now review all deployed code. Cuts hit support roles broadly, sales were spared. The pattern matches Meta, Microsoft, and Amazon’s recent moves: revenue growth and aggressive headcount reduction reported in the same breath, with AI productivity cited as the lever. Read against Dario’s and Dimon’s “no, capitalism absorbs every wave” arguments from earlier this week — the empirical record is now actively diverging from that historical comfort. AGENTS.md is the most-deployed AI-coding-agent ritual that probably isn’t earning its keep. Addy Osmani writes up two contradictory 2026 studies: Lulla et al. found AGENTS.md cuts runtime 28.6% and tokens 16.6%; ETH Zurich found LLM-generated context files reduced task success by 2-3% while raising costs >20%. The reconciliation: stripping the auto-generated content from repos improved performance by 2.7%, while developer-authored context with non-discoverable info (tooling quirks, operational gotchas) improved success 4%. The takeaway is sharper than “write good docs”: auto-/init is mostly duplicating what the agent can already grep, while genuinely useful context files have to be hand-curated and live deeper than the repo root. A new “AI breaks security disclosure” pattern just played out on a Linux kernel bug: 9 hours from initial private fix to independent rediscovery and public disclosure. Jeff Kaufman’s piece walks through the Copy Fail vulnerability — Hyunwoo Kim’s quiet upstream patch was independently re-derived by another researcher who saw the security implications and went public. The structural read: both “coordinated disclosure” (90-day embargoes) and Linux’s “bugs are bugs” culture were calibrated for a world where commit-scanning at scale was hard. Multiple AI-assisted research groups now scan kernel diffs continuously, so the embargo window is collapsing toward zero and the “fix-quietly-then-announce” middle path is gone. Worth filing alongside the broader thread that ML is changing offense/defense calibration in security, not just throughput. Analysis & Opinion Stop Using /init for AGENTS.md — Addy Osmani Two 2026 studies on AI-agent context files reach apparently opposite conclusions, and the reconciliation is the actual lesson. Lulla et al. found AGENTS.md reduced runtime 28.6% and output tokens 16.6%; ETH Zurich found LLM-generated context files cut task success by 2-3% while raising costs over 20%. The kicker: when ETH Zurich stripped the auto-generated context from repos, performance went up 2.7%, and developer-authored files with non-discoverable info improved success by 4%. The implication is that /init-style auto-generation is mostly restating what the agent can already discover by reading the code, and genuinely useful context has to be hand-curated and probably split across multiple files for non-trivial codebases. Useful counterpoint to the “always write AGENTS.md” reflex that’s emerged this year. ...

2026-05-09 · 10 min · Kun Lu

AI Daily Digest — 2026-05-08

Key Highlights OpenAI shipped GPT-Realtime-2, a voice model that finally closes the reasoning gap. Big Bench Audio jumped from 81.4% (predecessor) to 96.6%, and the new model can call tools simultaneously, reason mid-utterance, and “talk while it thinks.” A live translator covering 70+ languages and a streaming-transcription model shipped alongside. Zillow, Priceline, and Deutsche Telekom are already in production. The structural read: voice agents are graduating from turn-based dialog to multi-step workflow execution — the same arc text agents made in 2024-25, compressed into 18 months. Karp, Jensen, and Dario all showed up this week with sharply different theories of where the AI bottleneck is, and they don’t agree. Karp’s Q1 print (100% US growth, Rule of 145, free cash flow this quarter > revenue from same quarter a year ago) was the loudest argument that AI without an ontology is theater — he spent 15 minutes on the earnings call calling competitors “AI slop” and saying the demos work but the deployments don’t. Jensen at Milken made the opposite case: capacity is the bottleneck (agentic AI is 1000× more compute than generative AI), not platform discipline; both OpenAI and Anthropic just turned gross-margin-positive in the last 3-6 months and “are racing for capacity” because the unit economics finally work. Dario at JPMorgan added the 6-12 month estimate for Chinese open-weight models to catch frontier US labs and predicted individual SaaS companies will go bankrupt as moats collapse — useful to read in tension with Martin Alderson’s argument (covered below) that open-weight licensing is tightening, not loosening. Theo’s “What’s next?” video is the most useful single audit of GitHub alternatives anyone’s published this cycle. His framework: GitHub is dying, GitLab and Bitbucket are Gen-2 alternatives that are “just worse GitHub,” and the only mature open option worth recommending today is Forgejo / Codeberg (a community fork after Gitea went private). He went on-camera and donated $1,200 + $400/month live; that’s the credibility he’s putting behind it. The Gen-3 piece is Pierre’s code.sto (an ultra-low-latency Git cloud built for agent throughput — they hit 15,000 repos/min for 3 hours straight while GitHub buckles at half their volume) plus Entire (the new $60M-seed company from GitHub’s last CEO, building durable agent-context history alongside Git) and Zed’s Delta DB (CRDT-based realtime collab). The point worth filing: the Gen-2-to-Gen-3 jump may leave Git itself behind, the way Gen-1-to-Gen-2 left SVN behind. Analysis & Opinion Open Weights Are Quietly Closing Up — and That’s a Problem — via Lobsters Martin Alderson argues open-weight LLMs from DeepSeek, Qwen, and others are functionally the generic-pharma price ceiling of the AI economy: they cap pricing power on closed frontier models because “if frontier labs raised prices 5× overnight, a huge amount of people would just switch.” The trend he documents is that this constraint is quietly weakening — Meta has stopped releases, Alibaba is moving more weights behind API-only access, and Mistral is layering commercial-use restrictions onto licenses. That’s directly opposite to the “abundant defender swarms” framing Jensen used at Milken (where open source is the cybersecurity dome against frontier-model attackers), so the two pieces are worth reading against each other. ...

2026-05-08 · 9 min · Kun Lu

AI Daily Digest — 2026-05-07

Key Highlights Anthropic just leased SpaceX’s Colossus 1 supercluster — 300+ MW and 220,000+ Nvidia GPUs — from a company whose CEO has spent the last six months publicly calling them “misanthropic.” Per Anthropic’s announcement (and Theo’s deep-dive on the implications), the deal gets resolved by month’s end and is being paired with doubled Claude usage caps across paid tiers, removal of peak-hour throttling, and a 4–5× bump on API rate limits. Dario Amodei told the Code with Claude conference Anthropic saw an annualized 80× revenue/usage growth in Q1 and admitted they undershot compute planning; Theo’s read is sharper — Anthropic has world-class research, OpenAI has all three (research, data, compute), XAI had only compute, and Cursor had only data, which explains both the Anthropic↔SpaceX compute deal and the SpaceX↔Cursor $10B-for-data / $60B-for-the-company option signed weeks earlier. The unifying point worth filing: every recent “puzzling” Anthropic move (trying to remove Claude Code from Pro, peak-hour caps, banning Windsurf and XAI from API access) was a compute-allocation problem, not a pricing-power problem. The five-architect panel at Milken put hard numbers on what “supply-constrained” actually means in 2026. ASML’s Christophe Fouquet: chip supply will limit hyperscalers for “two, three, maybe five years.” Google Cloud COO Francis deSouza: $20B/quarter revenue at 63% YoY growth, with backlog almost doubling from $250B to $460B in a single quarter. Applied Intuition’s Qasar Younis flagged the non-silicon bottleneck: real-world data collection — synthetic simulation can’t fully bridge training gaps for autonomy. The takeaway across the panel was unusually candid: this isn’t a demand-discovery story anymore, it’s an industrial throughput story, and the constraint moves down the stack from chips to data to physical-world capture as you go from LLMs to agents to embodied systems. Google shipped a Prompt API in Chrome that requires accepting Google’s AI terms of service to call a standard web API, and silently downloaded a 4GB Gemini Nano model to users’ machines without consent. Mozilla, WebKit, and the W3C TAG all opposed the proposal; Google shipped it anyway. The author’s argument is narrow but load-bearing for the open web: no W3C-track API should require agreeing to a single advertiser’s prohibited-use policy as a precondition for being callable, and the user-side consent failure (auto-download + auto-reinstall after deletion) compounds the standards problem. Reads as the natural sequel to last week’s Telus accent-conversion story — AI features now ship across long-standing trust boundaries first, and the consent/disclosure debate happens in retrospect. Analysis & Opinion Anthropic, SpaceX(AI) become unlikely compute partners — The Rundown Despite months of public hostility — Musk has repeatedly called Anthropic “misanthropic” on X — SpaceX is leasing its Memphis-based Colossus 1 supercomputer (220,000+ Nvidia GPUs, 300+ MW) to Anthropic to ease an acute serving-capacity crunch. As part of the deal, Claude usage caps double across paid subscription tiers, API users get further increases, and peak-hour rate limits are removed entirely. Musk publicly justified the about-face by saying he’d spent time with senior Anthropic staff and was “impressed,” and that SpaceX’s own training had already moved to Colossus 2 — making Colossus 1 surplus, given how little Grok inference traffic XAI is actually serving. The structural read is that the deal is mutually rational: Anthropic plugs a serving-capacity hole; XAI monetizes a data center it built for an inference business that hasn’t materialized; both companies route around OpenAI, which is currently the only major lab with research, data, and compute under one roof. ...

2026-05-07 · 8 min · Kun Lu

AI Daily Digest — 2026-05-06

Key Highlights Telus is using real-time AI voice conversion to alter offshore call-centre agents’ accents — and the disclosure question is now a live regulatory issue. The tool, built by Tomato.ai, performs low-latency speech-to-speech transformation to reduce what Telus reportedly internally calls “accent-related friction.” Labour groups have called the practice deceptive and are pushing for mandatory disclosure; Rogers and Bell told The Globe and Mail they have no plans to deploy similar tech. The technical stack — ASR + speaker/accent conversion + neural vocoder — is now cheap enough to run at call-centre scale, which means the consent and identity-disclosure question Telus is creating will eventually land on every customer-facing voice deployment, not just one telco’s. Worth filing alongside the Chrome silent-Nano install story from yesterday: a pattern of AI features being rolled out across trust boundaries with no user-facing notice, then debated in the press after the fact. The “AI phone” race is now an OpenAI-vs-everyone IPO timeline play. Ming-Chi Kuo’s supply chain note pulls OpenAI’s first phone forward by ~12 months, into 1H 2027 mass production, with MediaTek as exclusive chip supplier and dual AI processors handling vision and language in parallel. The accelerated schedule is being read as IPO-driven — hardware as an investor-narrative asset rather than a margin business. The unanswered question is what this does to OpenAI’s existing Jony Ive / “io” hardware effort acquired last year; the public framing has shifted from “beyond screens” to a fairly conventional smartphone that just happens to ship with an enhanced HDR vision pipeline tuned for agent perception. Xbox’s new CEO has killed Copilot AI features across console and mobile as part of a broader restructuring driven by declining gaming revenue. Asha Sharma’s internal memo explicitly named “shipping impact quickly” as the org problem and slotted four CoreAI leaders into Xbox roles — including former ChatGPT growth lead Jonathan McKay as Xbox’s head of growth. Two Xbox veterans (Kevin Gammill, Roanne Sones) are out. The signal worth tracking: this is the second high-profile retreat from a consumer Copilot integration in 30 days (after Microsoft paused new GitHub Copilot signups on capacity grounds), and reads as an explicit decision that Copilot is no longer the right product surface for every Microsoft division to bolt onto. Analysis & Opinion Telus uses AI to alter call-agent accents — via Hacker News Telus is using a Tomato.ai-built speech-to-speech model through its Telus Digital unit to alter offshore call-centre agents’ voices in real time, and the rollout has triggered swift backlash from Canadian labour groups who are calling the practice deceptive and pushing for mandatory disclosure. The technical pieces — automatic speech recognition, accent/speaker conversion, latency-optimized neural vocoders — are now cheap and deployable at scale; what’s new is that a major North American telecom has put it into production against its own employees’ voices without (per the reporting) a clear consent or disclosure protocol for callers. Rogers and Bell told The Globe and Mail they have no plans to deploy similar systems, suggesting the industry isn’t in alignment that the practice is acceptable. The deeper question this story forces — and which existing telecom and consumer-protection regulators in Canada are not currently positioned to answer — is whether real-time identity-modifying voice AI in customer interactions requires affirmative disclosure on the call, or whether it can ride on existing terms-of-service and “calls may be recorded” boilerplate. The story also lands awkwardly for the offshore agents themselves: a tool framed as reducing “friction” is, from the labour side, an externally-imposed identity edit applied to people who are already at the bottom of the customer-service supply chain. ...

2026-05-06 · 5 min · Kun Lu

AI Daily Digest — 2026-05-05

Key Highlights The “AI subsidy economy” story is the wrong frame — the real binding constraint is compute, not money. Theo’s response to The Primeagen’s viral take pulls together the recent moves that everyone has been reading as price gouging (Anthropic restricting Claude Code on the $20 tier, Microsoft pausing GitHub Copilot signups, Anthropic shifting peak-hours usage limits) and argues they all share a single root cause: there are not enough Nvidia GPUs in the world to serve both the consumer subscription tier and the enterprise contracts the labs actually make money on. The labs aren’t trying to squeeze $200/mo users — they’re trying to claw back compute so it can be sold to Fortune 500 customers paying API rates. The supporting math is vivid: a $200/mo Claude Max subscription can extract up to $5,000 of inference at API prices, and Theo personally watched a single Copilot request burn ~$100 of compute over a 2-hour run. The cost-per-token narrative also misses the bigger trend — at any fixed intelligence level, prices are dropping fast: GPT-5.5 medium matches GPT-5.4 high at <half the cost, and 5.5 low scores higher than DeepSeek V4 on 7M tokens vs. 200M+ for Claude Sonnet 4.6 max on the same eval. Google Chrome is silently installing a 4 GB AI model (Gemini Nano) on user devices without consent. A privacy researcher documented weights.bin appearing within minutes of profile creation on a clean macOS install, traced through filesystem logs, Chrome config, and the updater. The author argues this likely violates EU privacy regulations and parallels Anthropic’s Claude Desktop behavior — a class of “forced bundling across trust boundaries” dark patterns that the AI rollout is normalizing. The environmental angle (multiplied across Chrome’s ~3B installs) is a non-trivial second-order story. Jensen Huang is now publicly fighting the “AI eliminates jobs” framing, telling the Milken Institute that AI is “creating an enormous number of jobs” and is the U.S.’s best shot at re-industrialization. In a separate SCSP conversation he is more specific: the bottleneck isn’t whether software-engineer jobs exist (Nvidia is hiring more), it’s energy — the U.S. needs to modernize the grid and lean into nuclear/solar backing if it wants to host the manufacturing required for AI. Both messages arrive against analyst projections that 15% of U.S. jobs could be displaced within several years; the Huang counter-thesis is that “task” and “purpose of the job” are not the same thing, the same argument that kept radiologist headcount rising even after computer vision swept the field. Analysis & Opinion Prime is (mostly) right about AI — Theo - t3.gg (41m, video) A surgical response to The Primeagen’s “AI economy is breaking” video that agrees with the diagnosis but reframes the cause. Prime reads recent pricing and access changes (Anthropic kicking Claude Code off the $20 tier, peak-hours throttling, Cursor moving from message-count to usage-based billing, Microsoft pausing Copilot signups) as evidence that the subsidy era is ending and the labs are clawing back revenue. Theo argues this is a misread: the labs are perfectly happy losing money on consumer subs as a marketing expense — what they cannot afford is GPU capacity being consumed by $20/mo users when enterprise customers paying full API rates are queued up behind them. The Microsoft Copilot signup pause is the cleanest tell — you don’t pause new revenue to make more money; you pause it because you don’t have capacity. He also dismantles the “model losses” argument with the same economic frame Dario Amodei used: looked at per model, each generation has been profitable; it’s the next-generation training cost that makes the company-level P&L look bad, but post-training (RLVR, RLHF) is now a much bigger lever than pre-training, so newer models are not necessarily more expensive than the ones they replace. The closing data point is the most important thing for anyone planning model spend: GPT-5.5 is 2× more expensive per token than 5.4 but uses so many fewer tokens that 5.5 high actually costs ~20% more than 5.4 high for the same task, and 5.5 medium matches 5.4 high quality at less than half the price. The cost of intelligence is dropping; the cost of frontier intelligence is rising. Both are true. ...

2026-05-05 · 10 min · Kun Lu

AI Daily Digest — 2026-05-04

Key Highlights Microsoft and OpenAI’s exclusivity deal is effectively dead — the amended partnership announced last week strips Azure of its sole-cloud-provider status, lets OpenAI ship to AWS (which just signed a $100B/8-year compute extension on top of an existing $38B deal), and quietly loosens the AGI-trigger clause that would have ended the IP-sharing relationship. Theo’s read: Microsoft got 2032 IP rights as a consolation; OpenAI got everything else, especially the right to chase Anthropic in Bedrock-dominated enterprise accounts. The structural reason this matters: most enterprise AWS startup credits cannot be spent on Anthropic models — the rev-share Anthropic locked in with the hyperscalers is too expensive for AWS/Google to subsidize — which is the under-discussed reason Anthropic’s enterprise revenue is growing faster than OpenAI’s despite ostensibly weaker code models. Putting OpenAI on Bedrock attacks that moat directly. Two studies converging on the same finding: LLMs collapse the variance in human writing. A new Lobsters-surfaced research project (“How LLMs Distort Our Written Language”) shows LLM-edited essays drift toward a common region of semantic space, become more neutral on argument stance, increase formality while reducing personal pronouns, and simultaneously amplify both emotional and analytical vocabulary — distortions human editors don’t make. Most striking: the same homogenization shows up in ICLR 2026 peer reviews, where AI-generated reviews systematically over-weight reproducibility/scalability and under-weight clarity/relevance. Users prefer the assisted output and still report “a statistically significant loss of voice and creativity” — the satisfaction signal is decoupled from the quality signal. A 2024-era model is already outperforming attending physicians at ER triage. Harvard’s Science paper on o1-preview across 76 real ER cases: 67.1% diagnostic accuracy vs. 55.3% and 50.0% for two attendings, and in one case the model flagged a flesh-eating infection in a transplant patient 12–24 hours before the treating doctor caught it. Working from raw EHR text only, no extra context. Reviewers couldn’t distinguish AI from physician-generated assessments. The implication that hangs over the result: this is o1-preview, not a frontier model — the question isn’t whether AI will be deployed in formal patient care, it’s how long the regulatory lag holds. Analysis & Opinion Do AI summaries hurt critical thinking? — Blueprint for Disaster (via Lobsters) The piece argues that AI summarization is a categorically different shortcut from older ones (CliffsNotes, Wikipedia, the Seinfeld-grade “just watch the movie”), because it removes the cognitive friction that those workarounds preserved. Reading a summary is still reading; having an LLM compress a piece you never opened is something else. The author’s frame — “machines generate content that other machines then condense” — lands harder when paired with today’s other research finding that LLM-edited prose collapses toward a homogenized semantic center. The implicit chain is uncomfortable: if generation flattens variance and summarization removes engagement, the medium-term failure mode isn’t bad reasoning, it’s no reasoning at all on the consumer end of the pipeline. The piece doesn’t offer a prescription, which is fair — the problem isn’t AI summaries, it’s that the demand for them is real and growing. ...

2026-05-04 · 6 min · Kun Lu

AI Daily Digest — 2026-05-03

Key Highlights Addy Osmani argues the /init-generated AGENTS.md is making coding agents worse, not better — citing a 2026 study where LLM-generated context files cut task success by 2–3% while inflating costs over 20%. ETH Zurich research separates out why: documenting agent-discoverable info (directory layout, file names) is pure noise, while non-discoverable context (specific tooling quirks, gotchas, conventions) is what actually pays off. The takeaway: treat AGENTS.md as a curated list of codebase smells, hand-written, not a config file. A growing “specs-first” backlash to vibe-coding — the Specsmaxxing essay (front-page HN today) describes the trap of escalating specification frameworks into “AI psychosis,” where you build systems-to-build-systems instead of building products. The author lands on YAML acceptance criteria as the minimum viable grounding artifact: enough structure to keep agents from over-engineering, not so much that you’re now maintaining a meta-tool. Lines up with Addy’s point — written human context beats generated context. UiPath’s CMO on why 70–80% of enterprise AI pilots stall — Michael Atalla argues the bottleneck isn’t ambition but coordination: organizations deploy AI tools in isolation, can’t see across them, and can’t scale past pilot. He draws a direct analogy to the Office 365 cloud transition, where companies who simply transplanted on-prem workflows failed; the same mistake is now repeating with AI. Analysis & Opinion Stop Using /init for AGENTS.md — Addy Osmani Osmani’s clearest argument yet against treating AGENTS.md (and its variants like CLAUDE.md) as a setup-time artifact. Two 2026 studies he cites contradict each other on whether these files help — one finds efficiency gains, the other finds 2–3% lower task success and 20%+ higher cost when the file is LLM-generated. ETH Zurich’s framing reconciles them: the question isn’t whether to include context, it’s what kind. Information the agent can discover by reading the codebase (file tree, exports, types) just dilutes the prompt; information the agent can’t discover (a tool that requires --verbose to emit machine-readable output, a flake that needs a 3-second sleep before the next call, a convention you adopted last year and haven’t migrated everywhere yet) is where these files earn their cost. His mental model — AGENTS.md as a living list of codebase smells — is the part worth stealing. ...

2026-05-03 · 4 min · Kun Lu

AI Daily Digest — 2026-05-02

Key Highlights The White House is quietly de-escalating its fight with Anthropic — a forthcoming AI memo would address Anthropic’s grievances and let agencies route around supply-chain risk designations, even as Defense Secretary Hegseth still calls Anthropic’s leadership “an ideological lunatic.” The real constraint is compute, not policy: as GPT 5.5 reaches Mythos-class cyber capability, the case for restricting Mythos to ~50 private firms is eroding fast. Theo escalates his Anthropic critique with receipts — Anthropic’s “third-party harness detection” overlapped with cloud-code’s git-history injection so badly that a user got billed $200 simply because the string Hermes.md appeared in a commit message of an empty repo. Theo also publishes T3 Chat’s actual Anthropic bill — ~$40,000/month, with prompt-cache writes alone costing $970/day — and notes turning caching off didn’t change the bill, undercutting Anthropic’s stated “caching” rationale for blocking OpenClaw. Jensen Huang lays out the “AI as a five-layer cake” at SCSP — energy, chips, infrastructure, models, adoption — and argues America’s biggest weakness is the adoption layer, not the chip layer. He explicitly disavows the “AI will wipe out 50% of jobs” framing as “ridiculous and counterproductive,” and walks through why software engineering hiring is up at Nvidia despite Codex/Claude Code automating most of the typing. Sam Altman concedes 4o’s sycophancy was a real safety failure — and says he’s now privately consulting clinical psychologists and spiritual leaders to write “instruction manuals” for ChatGPT’s default personality, treating the personality layer with the same rigor as bio/cyber risks because “the impact this has had on the world is huge.” He also says GPT-5.5 with Codex compresses “weeks of work two years ago into an hour.” OpenAI is reportedly behind on its 1B WAU and revenue targets, with CFO Sarah Frier flagging a mismatch between growth and $600B in compute commitments while Altman pushes for an IPO this year — even as developer sentiment swings back to OpenAI on GPT-5.5/Codex. The All-In hosts frame this as the first real fissure between OpenAI’s research and finance leadership. Analysis & Opinion The White House rethinks its Anthropic fight — Rundown After months of escalating Pentagon–Anthropic tensions, the administration is shifting from confrontation to triage: a forthcoming AI memo would address Anthropic’s grievances and let agencies route around supply-chain risk designations, while still capping private-sector access to Mythos at roughly 50 companies (Anthropic asked for ~120). The piece is sharpest on the underlying tradeoff — compute, not policy, is the real constraint, and as competing models close the capability gap (GPT 5.5 reportedly already at Mythos-class cyber capability, with one former official expecting full parity in six months), the strategic case for restricting Mythos starts to evaporate. Worth reading for the political subtext: even as the White House softens, Hegseth is publicly calling Anthropic’s leadership “an ideological lunatic,” signaling this truce is fragile. ...

2026-05-02 · 7 min · Kun Lu

AI Daily Digest — 2026-05-01

Key Highlights GitHub’s reliability has collapsed to the point that Theo and Mitchell Hashimoto (Ghosty creator) are both publicly leaving — outages are now hours-long, the merge queue silently reverted ~2,800 PRs on April 23rd, a Wiz researcher landed an unauthenticated RCE via git push -o header injection, and npm let a name-squatter ship malware as the legitimate tanstack package. GitHub currently has no CEO; product and engineering report up to a Microsoft EVP also overseeing Azure and Copilot. Reiner Pope (Maddox CEO, ex-Google TPU) does a 2-hour blackboard walk through how Claude/Gemini/GPT-5 are actually served on Dwarkesh — quantifying why “fast mode” exists (batch size economics), why optimal batches sit around 300 × sparsity (~2,000 tokens for DeepSeek-class MoEs), and why an HBM rack reads its full capacity in ~20 ms, which sets the floor on latency. OpenAI ships Advanced Account Security — phishing-resistant login, stronger account recovery, and new takeover protections, signaling that account compromise is now a first-class threat for AI accounts that increasingly hold persistent memory, tool credentials, and payment. ChatGPT Images 2.0 is a hit in India but flat globally — India is now the largest user base since launch, but Sensor Tower/Similarweb show only ~1% global DAU lift and ~1.6% web traffic gain; emerging markets spiked up to 79% week-over-week, mature markets barely moved. JavaScript’s Temporal proposal is finally landing after 9 years — Stack Overflow Podcast interviews Boa engine creator Jason Williams on why Date is broken, why Moment.js itself became the problem, and why a top-level Temporal namespace was needed at the language level. New Products & Tools Introducing Advanced Account Security — OpenAI OpenAI is rolling out phishing-resistant login, stronger account recovery, and new takeover protections across consumer and developer accounts. The framing is defensive — keep attackers out — but the timing matters: as ChatGPT accounts accumulate persistent memory, connected tool credentials, payment instruments, and now Codex/Operator-style agentic capabilities, account takeover stops being a privacy issue and becomes a credential-stuffing vector for agents that act on your behalf. This sets a precedent other AI providers will likely have to match. The most consequential bit isn’t any individual feature — it’s the implicit acknowledgement that an AI account is no longer a chat history; it’s a privileged identity. ...

2026-05-01 · 7 min · Kun Lu