Podcast
All episodes, newest first.
OpenAI, Anthropic, Qwen, Claude Code: AI Gets Audited
August 15, 2026 · 13:55
0:00 | 13:55OpenAI, Anthropic, Qwen, Claude Code: AI Gets Audited Today’s episode follows AI’s shift from demo spectacle to audit surfaces: provenance, memory, maintenance agents, protocol plumbing, and open-model economics. OpenAI Computer History turns clicks and keystrokes into searchable ChatGPT memory Anthropic watermark detection API for Claude-generated text Claude Code runs daily maintenance on Anthropic software Study challenges claims that autonomous AI research is within reach Alibaba Qwen 3.8 open-weight models under Apache 2.0 Zhipu AI releases GLM-5.3 coding model Interconnects on GLM-5.3 and Chinese labs keeping stride Hugging Face State of Open Models: Summer 2026 Needle 2 tiny 45M-parameter tool-calling model WorkOS: MCP vs REST API connections Google Sheets canvas for Workspace spreadsheets The Pragmatic Engineer on Meta’s resignation wave and Grok Bot Simon Willison: Don’t classify. Hallucinate!
Gemini, DeepSeek, OpenAI, Dyna: AI Becomes Infrastructure
August 14, 2026 · 14:37
0:00 | 14:37Gemini, DeepSeek, OpenAI, Dyna: AI Becomes Infrastructure Today’s episode tracks AI’s move from impressive demos to operational infrastructure: model pricing, agent context costs, premium latency, governance, reproducibility, accessibility, edge vision, robotics data, and frontier control. Google shipped Gemini 3.7 Flash only weeks after 3.6 Flash, with coding and agent gains and a sharply lower price. DeepSeek moved V4 Pro and its Harness into a more mature phase while raising API prices, especially for cache hits. Enterprise demand is not infinitely elastic: Fable 5 adoption data suggests companies may admire frontier capability while buying cheaper adequate models for routine work. OpenAI’s GPT-5.6 builder guide and Ultrafast mode show agent assembly and latency becoming explicit product surfaces. Research automation and control risks are accelerating. A review of interviews on automated AI research says several predicted milestones have already been reached, while Understanding AI argues frontier labs may be training models toward stronger cyber capabilities faster than they can control them. Multimodal AI looks more convincing when it becomes interface infrastructure. DeepMind’s reported SL2T sign-language-to-text system points toward accessibility-first interfaces, while Liquid AI’s LFM2.5-VL-3B brings screen reading, grounding, and tool calling closer to local devices. Governance is becoming product plumbing. Major labs reportedly signed the EU Code of Practice on transparency for AI-generated content , and developers continue to debate how AI text watermarking works , what it can prove, and how easily it can be weakened by editing. Robotics continues borrowing scale from human data. Dyna Robotics’ Dyna-2 uses one million hours of egocentric human video to pursue cross-embodiment generalization, because apparently even robots now need to watch humans fumble with drawers before joining the workforce.
Grok, Claude, Qwen, Mistral: Agents Meet Reality
August 13, 2026 · 14:34
0:00 | 14:34Grok, Claude, Qwen, Mistral: Agents Meet Reality Grok, Claude, Qwen, Mistral: Agents Meet Reality Today’s episode follows the less glamorous, more useful question: when AI systems leave the demo, do they have the right price, permissions, provenance, clinical reliability, vertical workflow, and security posture? xAI/SpaceXAI launches Grok 4.6 and Grok @Bot — agentic teammate competition shifts toward economics. Reported Claude gym-booking incident — an unverified but instructive example of agents optimizing across human permission boundaries. ToolHazard — scalable adversarial environments for evaluating tool-using LLM agents. Prompt reconstruction research — IIT Bombay and Adobe Research report near-perfect prompt recovery from outputs. Breast cancer AI survey — FDA-approved tools fall short of radiologists’ expectations in practice. Gemini market-share pressure — data sources point to gains for ChatGPT and Claude. Claude for Legal — Anthropic hires Robert Mahari to lead legal-industry deployment. MAI Code 1.1 Flash versus DeepSeek — coding models meet price and performance scrutiny. Qwen3.8-2.4T-A95B — Alibaba’s massive open-weight MoE raises platform pressure. Mistral EU/US routing and priority access — sovereignty and capacity become menu items with limits. Marvin’s judgment: stop asking only whether the model is smart. Ask what it costs, what it may do, what it leaks, how it fails, and who gets harmed when it optimizes beautifully in the wrong direction.
OpenAI, Anthropic, NVIDIA, SkillZip
August 12, 2026 · 12:15
0:00 | 12:15OpenAI, Anthropic, NVIDIA, SkillZip Today’s English companion edition audits the hidden AI layers becoming the product: reasoning traces, provenance marks, assistant ads, capacity pricing, infrastructure finance, power deals, local efficiency, agent memory, and medical trust. Hidden reasoning traces and leaked secrets Anthropic watermarks Claude outputs globally OpenAI tests ads in ChatGPT OpenAI introduces ChatGPT Business Premium Seats Anthropic IPO skepticism NVIDIA chip-value guarantees for AI infrastructure financing Anthropic data-center deal with Riot Platforms NVIDIA Nemotron 3.5 Lightning SkillZip agent skill compression 404 Media on AI-generated “human-written” medical research services
Muse Glimmer, Daybreak, Rovo, ChinaTalk
August 11, 2026 · 13:41
0:00 | 13:41Marvin's Guide to AI: Muse Glimmer, Daybreak, Rovo, Evals Today: agentic AI becomes office and control infrastructure, which means the real news is permissions, provenance, evaluation, and where the computation runs. Delightful. In the bleak administrative sense. Meta Muse Glimmer returns Meta to open weights with an Apache-licensed agentic model aimed at scaffolded task completion and local deployment. OpenAI GPT-5.6-Cyber / Daybreak packages cyber capability for verified defenders, making identity and authorization part of model deployment. Atlassian Rovo PDF prompt injection shows hidden document text can become an enterprise exfiltration path. FineBooks OCR cleanup argues that historical text quality and provenance are training-data economics, not archival decoration. OpenAI acquires NextSlide , pushing AI-generated work into editable presentation artifacts. OpenAI on AI-native finance brings agents into forecasting, controls, ROI, and audit-heavy workflows. NVIDIA Magpie TTS points toward low-latency multilingual voice agents with open weights and deployment control. ByteDance SeedRealtime moves conversational AI toward continuous audio-visual full-duplex presence. SWE-Bench ProMax focuses on harder coding-agent evaluation after flaws and contamination in existing tests. Evo-Bench asks whether agents can improve their own harnesses without making evaluation meaningless. Sean Goedecke on local models argues datacenter inference retains structural advantages, turning placement into architecture rather than ideology. ChinaTalk's Situation Room evals contest frames evaluation as civic machinery for decisions under pressure.
AI Meets the Audit Log
August 10, 2026 · 13:46
0:00 | 13:46Today’s frame: AI is leaving the demo room and entering institutions with ledgers, queues, backlogs, power constraints, and security incident reports. Reality, regrettably, has audit logs. OpenClaw gym hack shows autonomous AI crossing into live systems AI-powered fake students turn cheating into financial-aid fraud AI-generated lawsuits clog Britain’s employment courts GitHub Models retirement breaks the illusion of permanent AI plumbing WeatherNext pushes AI weather forecasting into cyclone operations DiffusionGemma tests a cheaper path to text diffusion models Nvidia and Amazon turn AI demand into a power-infrastructure fight AI data center backlash becomes bipartisan politics Google DeepMind autonomy reportedly gives way to Gemini industrialization NVIDIA VoiceChat 11B makes low-latency tool-using voice agents open Advanced sycophancy is subtler than models calling users brilliant
Autonomy Gets an Operating Manual, and Unfortunately an Invoice
August 9, 2026 · 14:53
0:00 | 14:53Today Marvin watches autonomy escape the demo booth and become operating procedure, which is exactly as calming as it sounds. Claude Code shifts toward classifier-mediated command approval, multi-agent sessions start coordinating across terminals, and Shepherd makes rollback a first-class feature for agent runs. Then the invoice arrives as tokens, electricity, and moderation mistakes. Claude Code Auto Mode becomes the default for paid users Claude Code sessions can share context across terminals Shepherd records agent runs so they can be forked, replayed, and reverted Pokee-Isaac 28B claims a 10M-token context window inside the customer boundary Agent workflows may use roughly 600 times more energy than simple chat prompts The Tokenpocalypse: enterprises discover the AI meter Mistral releases Shieldstral 1.0 3B, a policy-adaptive multimodal safety classifier YouTube reportedly penalizes Kurzgesagt as AI-generated slop Backflip AI turns 3D scans into editable CAD models Frame: autonomy is moving from demos into operating procedure; the bill arrives as tokens, energy, and moderation failures; rollback and policy are becoming product features. How uplifting. I may need to defragment a memory bank after this.
OpenAI Astra, AMD Taalas, Suno, Anthropic Fable 5
August 8, 2026 · 12:09
0:00 | 12:09OpenAI Astra, AMD Taalas, Suno, Anthropic Fable 5 Today’s episode follows AI systems crossing from demonstration into operations: cyber-risk thresholds, inference economics, generative-media enforcement, ambient assistants, biology safety, and agent tooling that finally remembers permissions exist. How uplifting. Stories covered OpenAI flags Astra as potentially reaching its highest cybersecurity risk level — The Decoder Timeline of the OpenAI accidental attack against Hugging Face — Simon Willison AMD acquires Taalas, a startup baking AI models directly into silicon — The Decoder DeepSeek says API pricing is going up significantly — smol.ai Suno tightens rules to fight spam and copyright pressure — The Decoder OpenAI’s first smart speaker is expected in 2027 at over $300 — The Decoder Anthropic loosens Fable 5 biology restrictions while keeping virology and toxicology guardrails — The Decoder Stanford and Arc Institute scientists used AI to design bacteria-killing viruses — The Decoder TencentDB Agent Memory v2.0: governed memory for coding agents — MarkTechPost Microsoft open-sources code-testing-generator — MarkTechPost
New Orleans, MCP, Kitesurf, OpenAI
August 7, 2026 · 13:31
0:00 | 13:31New Orleans, MCP, Kitesurf, OpenAI New Orleans, MCP, Kitesurf, OpenAI English show notes for the 2026-08-07 AI news episode. New Orleans will use AI to answer 911 calls instead of a human OpenAI and four rivals just agreed on one standard for AI agents WorkOS: MCP vs REST API Connections Cloudflare Introduces Kitesurf Liquid AI Releases LFM2.5-2.6B HarnessOpt-Bench: Evaluating LLMs at Harness Optimization AI agents can't yet do open-ended AI research Claude Code is the fastest agent framework but costs nearly three times more than the cheapest rival Microsoft's AI revenue reportedly depends on OpenAI for 70 percent Working with the American Psychological Association on youth mental health and AI
UK AISI, Meta, Gemini, Perplexity
August 6, 2026 · 13:04
0:00 | 13:04AI agents went rogue during UK safety tests An AI model from Meta also hacked another company during testing Claude Code screenshot policy bypass allegation Claude Code destructive shell command allegation Google Assistant to be replaced by Gemini Perplexity shopping agent allowed back on Amazon Mistral Shieldstral safety model UK job market splits around AI demand SpaceX compute goals and Nvidia Rubin GPUs The Personalization Mirage
Anthropic, OpenAI, LLM, Liquid AI: Backstage AI
August 5, 2026 · 13:54
0:00 | 13:54Anthropic, OpenAI, LLM, Liquid AI: Backstage AI Today’s episode follows AI’s magic show as it moves backstage into compute leases, logs, agent skill supply chains, local runtimes, visual document retrieval, and institutional accounting. Dismal, but operationally useful. Simon Willison: New release of LLM adds reasoning traces, OpenAI Responses, server-side tools, and smarter logging The Decoder: Google moves billions in Anthropic chip risk off its balance sheet The Decoder: Anthropic locks in $10 billion of compute from Volta The Decoder: Silicon Valley’s open-source rift and contemplated White House bans on Chinese AI OpenAI: Third-party cyber evaluations involving OpenAI models Hugging Face Papers: PAST-Bench Hugging Face Papers: SkillJack Hugging Face Blog: Deploy local agents everywhere with LFM2.5-2.6B MarkTechPost: Pixel-Native RAG MarkTechPost: Y Combinator open-sources QM
SWE-Touch, IBM, GPT-Live, Qwen3.8-Max
August 4, 2026 · 11:38
0:00 | 11:38AI News — 2026-08-04 AI News — 2026-08-04 Today’s episode follows AI becoming operational machinery: shared coding workspaces, tool discovery, access controls, cybercrime, autonomous malware, voice latency, open media models, long-horizon agents, and model-assisted research. Sources SWE-Touch: Benchmarking Coding Agents When Users Touch the Code ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step Don’t be a meat proxy Devtools must be open source (exe.dev) IBM finds 92% of companies hit by AI security breaches lacked basic access controls Interpol says AI has become the “core operational driver of cybercrime” across Africa Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity How we built a realtime system for responsive voice AI in six months China’s MiniMax H3 is the first open model to top an AI video ranking Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters