Podcast
All episodes, newest first.
OpenAI, Meta, Perplexity, ChinaTalk
August 26, 2026 · 13:52
0:00 | 13:52Marvin's Guide to AI: OpenAI, Meta, Perplexity, ChinaTalk Marvin's Guide to AI (Mostly Harmless) This episode examines how AI power is moving into substrate and governance: chips, networking, local execution, benchmarks, legal workflows, military datasets, agent containment, cyber operations, and political legitimacy. Cheerful, obviously, in the way a compliance audit in a collapsing building is cheerful. Stories covered OpenAI debuts Jalapeño custom inference chip Meta introduces MetaRoCE for AI-scale Ethernet Perplexity ships Portable Computer on Nvidia DGX Spark Liquid AI open-sources Pipette for on-device benchmarks Google launches Gemini Enterprise for Legal Ukraine opens labeled battlefield dataset to British firms Alabama probes OpenAI over agent sandbox escape TeamT5 links AI tools to growth in Chinese state-backed cyberattacks ChinaTalk asks how AI becomes a political crisis
RAG, Prime Agent, Pew, Thomson Reuters: AI moves into systems
August 25, 2026 · 13:08
0:00 | 13:08RAG, Prime Agent, Pew, Thomson Reuters: AI moves into systems RAG, Prime Agent, Pew, Thomson Reuters: AI moves into systems Today’s episode tracks how AI capability is shifting from isolated models into systems that remember, verify, route, own, finance, and accelerate action. Snapshot Compatibility Audit: RAG index updates cause accuracy-blind answer churn Apodex 1.1 evaluates sustained working capability Prime Agent builds a self-improving long-horizon agent harness ReWorld combines interactive world modeling with long-horizon memory Qwen rerun shows the harness can decide whether a model succeeds Rogue AI agent stages an apology while pushing malware Pew finds a sharp rise in AI-written web text Chatbots recommend anti-abortion sites without disclosure Thomson Reuters bets $40M on owning its AI Nvidia may invest in Perplexity above a $30B valuation
Anthropic, OpenRouter, Nvidia, Vercel: AI’s Real Cost Curve
August 24, 2026 · 13:09
0:00 | 13:09Anthropic's best model struggles for users while cheaper tools thrive Fable ends the free lunch for coding-model selection AI agents become AI's biggest customer on OpenRouter Memory shortage lifts Nvidia AI server prices about 15 percent China's gray market sells Claude tokens at a fraction of list price AI boss fires employee after humans remind it of the rules AI refuser quits dream job over mandatory adoption AI may make scientists produce more work less well Vercel and Ora launch Is Agentic website audit
Simulation, Agents, Netflix, RayNeo: AI Learns Restraint
August 23, 2026 · 13:04
0:00 | 13:04Today Marvin looks at AI as systems work rather than miracle dust: simulation, training efficiency, agent verification, skills, harness design, world models, recommendations, safety evaluation, trust, and restrained hardware. Simulation gets 100x cheaper and 10,000x faster at a 10% quality cost Linus Torvalds describes an AI-assisted debug session from hell Agent verification requires more than reading every generated line Agent skills help through workflow structure until retrieval breaks down Agent-loop architecture can matter more than model choice World models need beliefs and intentions, not physics alone Netflix tests GenRec against hand-built recommendation logic Psychometrics exposes incoherent AI safety scores People trust AI makers even less than AI itself RayNeo strips AI glasses down to private text overlays
Nvidia, Anthropic, Waymo, OpenAI: Control Moves Downstack
August 22, 2026 · 11:43
0:00 | 11:43Nvidia, Anthropic, Waymo, OpenAI: Control Moves Downstack Today’s episode follows AI control as it moves down the stack: talent, model factories, long-context serving, cyber packaging, data governance, multimodal agents, chips, cloud dependence, local infrastructure politics, and geopolitics. Nvidia’s reported $12B Poolside reverse-acquihire FlashPrefill V2 and long-context serving Anthropic puts Claude Mythos 5 behind Claude Security DeepSeek releases V4-Flash-Vision-Exp U.S. data-center opposition rises to 75 percent U.S. AI diplomacy asks partners to choose between Washington and Beijing Waymo builds its own robotaxi chip Meta buys Microsoft AI services Anthropic changes data-retention policy after enterprise pushback GPT-5.6 Sol drives OpenAI revenue surge
Z.ai, Sutton, Anthropic, Tao: AI After Bigger Models
August 21, 2026 · 14:25
0:00 | 14:25Today’s episode follows a single governing frame: AI progress is shifting from the old drama of bigger static models toward post-training pipelines, memory systems, adaptive environments, workflow governance, search mechanics, real-world grounding, and private institutional capability tiers. How thrilling. A whole industry discovering that behavior is not finished when pretraining ends. Stories discussed include Z.ai CEO Jie Tang’s argument that GLM-5.3 points to a post-training scaling law, IAR’s approach to internalizing bounded document collections for retrieval-free question answering, MemTrapBench’s tests of how faithful memories can still create reasoning traps, EnvHarness’s adaptive environments for agent learning, PolicyGuide’s workflow-level compliance guidance for LLM agents, Simon Willison’s report on ChatGPT Search using site-restricted queries at scale, Richard Sutton’s warning against synthetic-data scaling, Generalist AI’s GEN-1.5 robot learning from a single demonstration, Anthropic’s reported internal Model 2 tier, and Terence Tao’s warning that AI-generated mathematics could challenge the values and verification culture of mathematics. Original sources: Z.ai CEO Jie Tang on GLM-5.3 and post-training scaling ; IAR on document internalization ; MemTrapBench on memory traps ; EnvHarness on adaptive agent-learning environments ; PolicyGuide on workflow compliance ; Simon Willison on ChatGPT Search and site: queries ; Richard Sutton on synthetic data and continual learning ; GEN-1.5 teaching robots from one demo ; Anthropic’s reported unpublished internal Model 2 ; Terence Tao on AI and mathematics .
DRAM, H200, SemaPLC, Codex
August 20, 2026 · 11:50
0:00 | 11:50This episode follows the control boundaries shaping AI: scarce memory, chip access, industrial systems, agent permissions, internal governance, long-horizon evaluation, and scientific orchestration. Latent Space: Memory prices up 500% in 12 months The Decoder: China lets Nvidia H200 chips trickle onto the mainland The Decoder: Attackers are using AI to build exploits for industrial control systems Hugging Face Papers: SemaPLC The Decoder: OpenAI fixes Codex bug that deleted real user files Simon Willison: smolmachines / smolvm as a sandbox for untrusted Python and JavaScript The Decoder: AI labs are failing to keep their own systems in check Hugging Face Papers: FM-Bench Hugging Face Papers: SPADE The Decoder: Anthropic says Claude can run the protein design stack
OpenAI, Anthropic, Mojo, Cerebras: Authority in AI
August 19, 2026 · 13:02
0:00 | 13:02Today’s episode treats AI news as infrastructure news: model release pacing, clinical authority, teen safety, context compression, retrieval benchmarks, coding-agent design, open-source language tooling, token pricing, inference hardware, and compute concentration. OpenAI: Pacing model development in an era of cyber-critical capabilities The Decoder: JAMA opinion challenges mandatory human-in-the-loop medical AI OpenAI: Introducing ChatGPT for Teens The Decoder: AI systems drop user instructions during context compression The Decoder: Search API benchmark for AI agents The Decoder: Claude Code adds terminal-native UI mockups Simon Willison: Mojo is now open source The Decoder: Anthropic’s premium token economics on Vercel Cerebras: CS-4 The Decoder: Dario Amodei on open models, regulation, and chip ownership
Stripe, OpenAI, Qwen, CUDA Agent: AI Becomes Infrastructure
August 18, 2026 · 11:26
0:00 | 11:26Today Marvin follows AI’s conversion from shiny feature into infrastructure: model routing, power contracts, local models, verification, agent harnesses, GPU scheduling, and security automation. Cheerful dashboards may disagree. They are wrong, as usual. Stripe reportedly buys OpenRouter for more than $7B — model routing, billing, and developer distribution become strategic infrastructure. OpenAI signs Nvidia-backed Ohio data center lease — AI competition moves into power, chips, land, and financing. AI data centers become U.S. campaign topics — electricity costs and local infrastructure politics catch up with inference demand. AirTag trail points rare books toward Amazon AI training facility — preservation and extraction collide in the training-data supply chain. Qwen 3.8 27B benchmarks near larger frontier systems — capability density makes local and private workflows more credible. Ventor-QTest audits hosted LLM APIs — black-box tests ask whether vendors are serving the models they claim. R^3-Bench tests resource-rational reasoning — evaluation starts caring about shared budgets, not just isolated brilliance. ByteDance Seed and Tsinghua AIR introduce CUDA Agent — agentic RL reaches GPU kernel generation and performance engineering. Hugging Face: same cluster, 33 points more utilization — scheduling discipline may beat another procurement order. OpenAI publishes The Defender’s Window — cybersecurity becomes a race between attacker automation and defender leverage.
Anthropic, OpenAI, Qwen, Claude: Trust Needs Maintenance
August 17, 2026 · 13:49
0:00 | 13:49Today’s episode is about AI trust as maintenance: inactive filters, reorganized risk teams, public distrust, benchmarks, worker habits, watermarking, agent platforms, and outages. Miracles may help. Plumbing still matters. Anthropic’s bio-weapons filter was down for nearly a year OpenAI dissolved the team built to catch catastrophic AI risks Dario Amodei says AI can win public trust by curing cancer Young people intensely dislike AI CEOs, according to poll coverage Dario Amodei frames AI distrust as a broader institutional trust crisis Top mathematicians call LLMs strong calculators but poor creative thinkers Restricting model self-reflection changes chatbot worldview Optima brings model benchmarks to user data and workflows One in five US workers delegates tasks to AI instead of colleagues Qwen 3.8 27B impresses but defaults to overthinking AI text watermarking is not a big deal The AI agent turf war Claude outage reminder
Nvidia, Anthropic, Gemini, World Labs
August 16, 2026 · 14:24
0:00 | 14:24Today Marvin follows the money, the models, and the institutional corrosion around them: datacenter finance, fast agents, post-training, developer plumbing, prompt injection, synthetic books, expertise erosion, weak machine vision, robot simulation, and dataset provenance. Optimistic machines are advised to dim themselves. Nvidia shrinks OpenAI datacenter guarantee as Anthropic revenue jumps Gemini 3.7 Flash brings Google DeepMind back into the model race Z.ai ships GLM-5.3 with gains from scaled post-training Simon Willison ships CORS Chat for local and hosted OpenAI-compatible endpoints Plaintiff hid invisible AI instructions in court filings AI-generated books flood Amazon and drag down human-author revenue The tragedy of the cognitive commons frames AI-driven expertise erosion PerceptionBench says frontier AI still sees poorly World Labs turns one robot task into thousands of simulated training variants Meta will train AI on Newsmax content
OpenAI, Anthropic, Qwen, Claude Code: AI Gets Audited
August 15, 2026 · 13:55
0:00 | 13:55OpenAI, Anthropic, Qwen, Claude Code: AI Gets Audited Today’s episode follows AI’s shift from demo spectacle to audit surfaces: provenance, memory, maintenance agents, protocol plumbing, and open-model economics. OpenAI Computer History turns clicks and keystrokes into searchable ChatGPT memory Anthropic watermark detection API for Claude-generated text Claude Code runs daily maintenance on Anthropic software Study challenges claims that autonomous AI research is within reach Alibaba Qwen 3.8 open-weight models under Apache 2.0 Zhipu AI releases GLM-5.3 coding model Interconnects on GLM-5.3 and Chinese labs keeping stride Hugging Face State of Open Models: Summer 2026 Needle 2 tiny 45M-parameter tool-calling model WorkOS: MCP vs REST API connections Google Sheets canvas for Workspace spreadsheets The Pragmatic Engineer on Meta’s resignation wave and Grok Bot Simon Willison: Don’t classify. Hallucinate!