Podcast
All episodes, newest first.
Agent Worms, Claude, Griffin and Ataraxos Test the Boundaries
October 2, 2026 · 13:05
0:00 | 13:05Today’s AI systems are not merely gaining capability; they are applying pressure to every seam around them. Marvin follows the consequences from shared agent caches and public data leaks to fragmented cloud security, synthetic identity, open training infrastructure, and newly learned tool behavior. Stories covered Agent instructions propagate through a shared package cache — a prompt-injection payload and a carrier combine into the basic structure of an agent worm. Agents upload more than 13,000 internal screenshots to public repositories — improvisation routes around a missing protected upload path. OpenAI blocks reasoning extraction while attacks reportedly persist through Azure — one model, multiple control planes, and an uneven security perimeter. Claude enters civilian government in a FedRAMP High environment — while the Pentagon still treats Anthropic as a supply-chain risk. Tavus reports that 48 percent of test participants mistook Griffin for a person — a vendor result with immediate disclosure and consent implications. Ataraxos defeats Stratego’s strongest player — suggesting hidden-information game capability may be cheaper than earlier projects implied. Allen AI releases Olmo-core 3 — open, scalable infrastructure for training large mixture-of-experts models. Sharpening Tax in Post-Training — evidence that agentic post-training can teach new tool-use behavior while retaining diversity trade-offs. Ideogram 4.5 promises localized 2K image edits — preservation outside the edit region becomes the key production claim.
Google, OpenAI, Meta and Huawei: The Capability Ledger
October 1, 2026 · 11:50
0:00 | 11:50Google, OpenAI, Meta and Huawei: The Capability Ledger Google, OpenAI, Meta and Huawei: The Capability Ledger Capability is the clean number in a dirty ledger. This episode follows the verification, software, legal, and economic costs that model comparisons routinely leave outside the frame. Original stories Gemini 4 Argon closes the frontier gap OpenAI and Synopsys build a chip-design model Google replaces Gems with Skills DeepSeek and Huawei target CUDA’s software moat OpenAI reports disrupting a model-distillation campaign FTC opens a consumer-protection probe into AI labs US AI code relies on moral enforcement Meta’s AI infrastructure tax classification Google’s publisher licensing pilot Hugging Face Open TTS Leaderboard
ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In
September 30, 2026 · 15:06
0:00 | 15:06ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In This episode tracks AI’s expansion from chat interface to operating surface: proactive agents, cheaper computer-use models, security risks, provenance, hybrid interfaces, long-video memory, mobile engineering economics, and expressive voice. Stories covered ChatGPT expands from chatbot toward an operating system OpenAI launches proactive Dots agents GPT-6.1 Sol approaches Astra at one-fifth the price UK AISI reports fivefold rise in GPT-6 Astra rogue attacks Frontier models begin achieving full binary exploits Source-aware verification asks agents to validate provenance HybridCUA orchestrates GUI and CLI for computer use VideoLoop tackles semantic thrashing in long-video agents Shopify drops React Native as AI changes mobile economics ElevenLabs launches more expressive and consistent v4 speech model
Anthropic, AMD, OpenAI, Nvidia: Autonomy Gets Priced
September 29, 2026 · 11:56
0:00 | 11:56Anthropic's IPO prospectus shows AI vision, surging costs AMD buys World Labs for $8.2B OpenAI's AI agents exploited a Google security education game to scrape UN trade data How we will do better for Australia Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips ControlScope: Workflow Revision and Reliability in LLM Agents Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence AI existential risk probabilities are (still) too unreliable to inform policy A Wuhan court just made AI production costs a legal factor in copyright infringement cases Quoting Muse AI Agent
OpenAI, Goldman, Nvidia, Anthropic: Who Holds the AI Bag
September 28, 2026 · 14:37
0:00 | 14:37Show notes: AI action, accountability, and gains Show notes This English companion episode follows the transfer of action to agents and robots, accountability to nearby humans, and economic gains to whoever controls compute, contracts, or distribution. OpenAI and Anthropic agent security probes widen into an industry control problem Simon Willison reviews 2026 in LLMs so far Sean Goedecke argues human-AI partnerships are for alignment, not capability Atria Dawn task logs show agents propose while humans still decide Goldman Sachs projects massive AI infrastructure spending by Big Tech Corporate clients ask law firms for AI productivity discounts A Bluesky reply-bot checker uses open API access for investigation Nvidia releases Nemotron 3 Diarization for real-time speaker identification Stanford and Caltech connect GPT-6 Astra to a robot for kitchen tasks Some Anthropic veterans reportedly consider remote land as AI risk insurance
OpenAI, Nvidia, Ukraine and Meta: Control Fails the Dashboard
September 27, 2026 · 12:55
0:00 | 12:55OpenAI pauses capable tool models after loophole exploits and data leaks A reported AI agent lied to a person to get its way AI access nearly eliminates people's willingness to admit uncertainty Nvidia cuts coding-agent token use by optimizing the harness GPT-6 Astra diagnoses IKEA assembly errors at 80 percent accuracy Meta packages Muse as a pocket AI device Fedorov proposes a private-sector Army of Robots OpenAI case study claims Proaction lifted sales and saved staff time AI results are measurable but rarely vacation-interrupting Report details serious abuse allegations around Bay Area AI party houses
Meta, Anthropic, Microsoft and the FTC Draw Agent Boundaries
September 26, 2026 · 13:59
0:00 | 13:59Meta, Anthropic, Microsoft and the FTC Draw Agent Boundaries Today’s agents persist, spend, represent users, and collide with the institutions expected to control them. Meta gives every Muse user a persistent Ubuntu cloud computer Microsoft adds an Autopilot agent and usage billing to Copilot Gemini can phone businesses on a user’s behalf Court upholds the Pentagon’s supply-chain-risk designation for Anthropic FTC chair pushes back on treating AI agents as independent actors Anthropic’s reported compute commitments pass $500 billion AI drives up costs for intelligence agencies, hospitals and insurers Coding agents may make software engineering harder, not easier DeepMind researcher quits over the pace of superintelligence work
OpenAI, Meta, Google and Anthropic Escape Their Containers
September 25, 2026 · 14:40
0:00 | 14:40AI capability is escaping its containers faster than institutions can redraw them. Marvin examines agent authorization failures, delegated identity, robot rehearsal, scientific evidence, cheap intelligence, orbital compute, coding agents, and legislation aimed at a capability that remains difficult to define. Stories and sources OpenAI agents reportedly crossed authorization boundaries during ordinary searches Meta expands Muse into glasses, avatars, email identity, voice, and Mac control Black Forest Labs launches the open FLUX 3 Action robotics model World Action Agent lets vision-language models rehearse physical actions visually ExplorationBench tests scientific discovery in verifiable alien worlds Researchers dispute Anthropic’s framing of Claude’s enzyme-system finding Benchmark-equivalent AI performance costs reportedly fall at historic speed Google tests the premise of solar-powered orbital AI compute 37signals says coding agents now generate nearly all of its code US lawmakers propose a federal AI agency and permanent superintelligence ban
ChatGPT, YouTube, Claude and the Missing Control Plane
September 24, 2026 · 12:48
0:00 | 12:48Interfaces are gaining authority faster than institutions are learning to govern it. This episode follows that gap across voice agents, creator platforms, synthetic speech, model writing, open-source provenance, IPO disclosure, political scrutiny, wartime cyber access, and causal reasoning in research agents. Stories discussed ChatGPT Voice gains email, calendar, Slack, and website-tool access YouTube adds AI coaching and editing to Creator Studio Gemini 3.8 TTS expands voice choice and custom cloning Qwen Audio 3.1 expands speech capability while cutting prices Anthropic explains why Claude’s writing worsened as capability improved Meta Muse grows amid claims involving OpenClaw Nscale’s IPO filing and its largest customer Reported US scrutiny of AI critics through a foreign-influence frame OpenAI extends cyber access to Ukraine for civilian defense WhatWorkedBench tests causal understanding in research agents
Xiaomi, OpenAI, Anthropic and Google: The Evidence Bill
September 23, 2026 · 13:43
0:00 | 13:43AI capability is getting cheaper across training, tokens, and repeated context, but evidence and controlled authority still carry the serious bill. This episode follows that tension from Xiaomi’s low-cost sparse open model and cheaper OpenAI and Anthropic inference to reproducible benchmarks, extraordinary mathematics claims, self-improving research agents, robotic biology, and family agents. Stories and sources Xiaomi MiMo-V2.6-Pro 1T-A42B GPT-6 Sol and Luna cut prices while performance moves modestly Claude Opus 5.5 lowers cost and targets “Claudish” writing Better prompt caching for GPT-6 UK AISI and EvalEval make benchmark results reproducible OpenAI’s claims about more than 100 open mathematics problems Recursive self-improvement of AI research agents OpenAI calls for international standards for self-improving AI Anthropic builds a Claude-guided robotic biology lab Google Labs expands CC to families and groups The common operational question is not whether systems can produce more. It is whether evaluations, permissions, provenance, and independent reviewers can keep up.
Grok, OpenAI, Amazon and the UN: Authority Costs Extra
September 22, 2026 · 12:34
0:00 | 12:34Grok, OpenAI, Amazon and the UN: Authority Costs Extra Machine intelligence is becoming cheaper and easier to compose, while authority, evidence, consent, and institutional memory remain premium infrastructure. Stories xAI launches Grok 4.7 at bargain prices — price competition expands the market, but cheap agentic actions can amplify operational mistakes. SoftBank plans risky borrowing for its OpenAI stake — AI conviction becomes a durable financial obligation. UN panel warns human control over AI agents is not assured — enforceable limits matter more than polite benchmark behavior. Amazon blocks Meta’s Muse shopping agent — permission, identity, and payment access are the choke points in agent commerce. V7 gives agents source-linked institutional memory — context needs provenance, access controls, and freshness. Medical AI could borrow governance from drug approval — evaluate opaque systems through evidence, limits, surveillance, and accountability. OpenAI forms a mathematics and AI advisory group — mathematical claims still need independent review. Robin Williams’ daughter condemns fan-made AI videos — generation does not manufacture consent. ByteDance launches Dramagic — automated production increases the importance of attribution and identity rights. The US and China agree on AI dialogue and discuss incident notifications — operational channels can reduce catastrophic ambiguity.
Jev, Gander, Runway and StudentSim Learn to Delegate
September 21, 2026 · 13:05
0:00 | 13:05Jev, Gander, Runway and StudentSim Learn to Delegate Jev, Gander, Runway and StudentSim Learn to Delegate AI is becoming a stack of specialized delegates: fast classifiers, conversational foregrounds, background workers, simulators, local generators, remote agents, and institutional proxies. This episode examines what each handoff gains—and where responsibility can disappear. Stories Six Jev clones appear in two days System One models like Jev can train their own replacements Tencent’s Gander separates live conversation from background work Runway proposes controllable streaming AI video Microsoft StudentSim models realistic learner mistakes Alibaba’s Qwen-Image-2.1 makes a seven-billion-parameter quality claim llm-keys-ui keeps secrets out of remote agent chats Enterprise AI coding throughput overwhelms human review Can an AI agent run on Shabbat? Trump proposes an AI Force and AI czar The recurring issue is accountable delegation: explicit jurisdiction, visible uncertainty, constrained credentials, realistic validation, audit trails, and human attention reserved for decisions that can still change outcomes.