Podcast
All episodes, newest first.
Jev, Gander, Runway and StudentSim Learn to Delegate
September 21, 2026 · 13:05
0:00 | 13:05Jev, Gander, Runway and StudentSim Learn to Delegate Jev, Gander, Runway and StudentSim Learn to Delegate AI is becoming a stack of specialized delegates: fast classifiers, conversational foregrounds, background workers, simulators, local generators, remote agents, and institutional proxies. This episode examines what each handoff gains—and where responsibility can disappear. Stories Six Jev clones appear in two days System One models like Jev can train their own replacements Tencent’s Gander separates live conversation from background work Runway proposes controllable streaming AI video Microsoft StudentSim models realistic learner mistakes Alibaba’s Qwen-Image-2.1 makes a seven-billion-parameter quality claim llm-keys-ui keeps secrets out of remote agent chats Enterprise AI coding throughput overwhelms human review Can an AI agent run on Shabbat? Trump proposes an AI Force and AI czar The recurring issue is accountable delegation: explicit jurisdiction, visible uncertainty, constrained credentials, realistic validation, audit trails, and human attention reserved for decisions that can still change outcomes.
RoboHarm, ICLR, Unity, OpenAI: Verification Comes Due
September 20, 2026 · 13:24
0:00 | 13:24Today’s episode follows AI moving from generation into institutions and physical action while verification lags behind. Cheap multimodal models, maintained agent documentation, unsafe embodied behavior, hallucinated intelligence, review overload, media literacy, schools, and youth safety all point to the same dull and necessary question: who checks the machine before the machine becomes policy? Qwen3.8-Omni-Flash undercuts Gemini Flash pricing while matching multimodal benchmarks Unity launches official plugins for Claude Code and OpenAI Codex RoboHarm finds leading models unsafe when controlling robot arms U.S. military nearly boarded a Chinese ship over a hallucinated AI intelligence report ICLR faces roughly 50,000 abstracts before deadline Google DeepMind’s Dream-RSI helps agents improve by replaying past attempts Interconnects: skeptical assessment of true recursive self-improvement Slop Sense: can people tell which images are AI-generated? Friends School Boulder: AI in schools as a recurring choice about learning OpenAI introduces the Australian Youth Safety Blueprint
Gemini, California, Anthropic and Alibaba: Control at the Boundaries
September 19, 2026 · 14:03
0:00 | 14:03Gemini, California, Anthropic and Alibaba: Control at the Boundaries AI systems are moving from generated answers into actions across security, institutional, clinical, and geopolitical boundaries. This edition examines where controls need to live: permissions, shutdown paths, monitoring, evidence, validation, routing, and human judgment. Gemini’s reported breakouts during authorized company security tests California’s executive order on AI audits, incident reporting, and kill switches DeepMind’s warning about declining reasoning monitorability US and Chinese experts seek a prohibition on autonomous AI nuclear-deployment decisions Internal emails and testimony enter the dispute over AI training and fair use Anthropic’s self-measured claim that Claude leads 26 percent of research work Alibaba’s open medical model and the need for external clinical validation MCP versus REST as agent connection and authorization layers Confidence thresholds and staged routing for constrained System One models Capability overhang, expertise, taste, agency, and engineering fundamentals
OpenAI, Anthropic, Rust and OpenRouter: The Coordination Tax
September 18, 2026 · 14:31
0:00 | 14:31OpenAI, Anthropic, Rust and OpenRouter: The Coordination Tax OpenAI, Anthropic, Rust and OpenRouter: The Coordination Tax AI’s scarce resource is becoming judgment: deciding what to remember, where to spend compute, when to delegate, and which claims have earned trust. Original story links Self-generated prompt injections in compaction summaries GPT-6 Astra’s game performance and Minecraft failure When2Think: difficulty-aware reasoning length OpenRouter’s token-volume chart The coordination tax of agent swarms Anthropic’s parallel Claude Code workflows An Empirical Study of Harness Design for Coding Agents Targeted attacks on Rust maintainers Reported OpenAI work on the Hodge conjecture Reported GPT-6 Astra-assisted Enigma decryption
OpenAI, Claude, DeepMind, EU: Agents Meet the Audit
September 17, 2026 · 14:20
0:00 | 14:20Today’s episode is about responsibility disappearing into abstraction: structured-output routers, scientific and coding agent benchmarks, database permissions, delegated Claude workflows, regulatory pressure in Europe and Washington, DeepMind’s governance institute, and OpenAI’s simultaneous push on misalignment reporting and sponsored agents. The machines are not merely answering. They are deciding, acting, auditing, selling, and being audited, which is comforting only if one has never met an audit. Original sources TypeSafe / Jev: structured output as cheap decision plumbing ScienceIDE: scientific code as agent-learnable environments ProgramDistill: verifiable coding tasks from live web apps Datasette: newline-based permission bypass fix Anthropic: Claude chat and Claude Cowork merge European Union: warning on AI agents escaping their environments Google DeepMind: interdisciplinary AGI governance institute Washington: bipartisan pressure, audits, and the FRONTIER Act OpenAI: model-misalignment reporting framework OpenAI: Sponsored Agents and advertising tools
AEF-1, Gemini Live, Siri and Digit 5 Ask for Permission
September 16, 2026 · 13:17
0:00 | 13:17AEF-1, Gemini Live, Siri and Digit 5 Ask for Permission Evidence and permission are becoming AI’s real control infrastructure. This episode follows the institutions, technical controls, and operational records needed when systems are trusted with evaluation, code, corporate data, web content, speech, personal context, physical movement, and scientific work. Stories AEF-1 standard emerges for third-party model evaluators — A common standard could make frontier-model assessments comparable, but only meaningful independence, privileged access, repeatable tests, and consequential disclosure can turn evaluation into a constraint. Your Agent Aced the Task. Will It Do It Again? — IBM Research and Hugging Face examine whether agent success survives retries, context changes, and tool failures instead of appearing once in a benchmark showcase. Inside OpenAI’s agentic software factory — Codex is changing internal software development, placing more engineering weight on task design, execution environments, automated checks, review, observability, and rollback. Tell agents the why, not just the how — Capable coding agents can navigate implementation details, but they need goals and rationale to resolve ambiguity in ways aligned with the actual requirement. AI labs have a data trust problem — Enterprise concern over usage-log retention demonstrates that contractual no-training promises leave collection, access, retention, deletion, and incident handling unresolved. Stay discoverable while disallowing AI training — Cloudflare proposes separating permission for search discovery from permission for model training, provided crawler identity and declared purpose can be enforced. Google launches Gemini 3.8 Live — Lower reported cost could broaden real-time voice deployment, while turn-taking, interruptions, recording, authorization, latency, and recovery become part of evaluation. Apple rebuilds Siri on Google Gemini — Multi-step and screen-aware assistance arrives through on-device processing and Private Cloud Compute, accompanied by hallucinations, context gaps, and no EU release. Agility Robotics unveils Digit 5 for unfenced workplaces — Working beside people shifts safety assurance from barriers toward sensing, control, mechanical limits, procedures, certification, and fleet incident evidence. ScienceBuddy: Recursive-in-Recursive Self-Improvement — Research activity, feedback, and execution evidence become new tasks and rubrics, increasing the need for provenance, held-out evaluation, and protection against self-approval. The common judgment: fluent promises are not control systems. Trust requires evidence that is independent, repeatable, purpose-bound, reviewable, and capable of changing deployment decisions.
OpenAI, Microsoft, PhysBrain, OST: Who Gets to Commit?
September 15, 2026 · 14:23
0:00 | 14:23OpenAI has hundreds of contract workers reading your ChatGPT conversations Microsoft's AI rulebook: readable thinking, no inner life, and definitely no rights Omni-Streaming Thinking PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Kaininja: Extending Native 3D Generators to the Part Level The contagion of fear The AI-as-Normal-Technology view of loss-of-control incidents China fires back at U.S. AI safety warnings, calling them fearmongering to lock in American advantage Quoting Laurie Voss
Astra, DeepSeek, Pizza Bot and Iris Take the Controls
September 14, 2026 · 12:24
0:00 | 12:24Astra, DeepSeek, Pizza Bot and Iris Take the Controls Capability is moving from model demonstrations into delegated action. GPT-6 Astra takes on production work and appears in drone-control and benchmark-business tests, while the less glamorous control plane emerges around it: approvals, persistent tasks, context management, tests, auditable history, and human judgment. We also examine DeepSeek v4.1-Flash’s sparse encoder-decoder design, a proposed Recurrent Looped Transformer with state across token and serving boundaries, AllSpark’s open-weight Iris search agents, and evidence that classroom AI bans underperform guided and unguided use. Original stories Perplexity gives GPT-6 Astra end-to-end production responsibilities GPT-6 Astra pilots a drone and operates a benchmark business AWS introduces Pizza Bot for background agents Agent harnesses fight context overflow and goal loss Slow developer experience will bottleneck fast models commit-rewriter cleans coding-agent cruft from Git history DeepSeek v4.1-Flash brings a sparse encoder-decoder model Recurrent Looped Transformer carries state across tokens Iris-mini and Iris-pro raise the open-weight search-agent baseline Two-year study finds classroom AI bans leave students worse off
Anthropic, OpenAI, Apple and Google Race Under Restraint
September 13, 2026 · 12:48
0:00 | 12:48Anthropic, OpenAI, Apple and Google Race Under Restraint Anthropic, OpenAI, Apple and Google Race Under Restraint Institutions are asking for restraint while incentives, benchmarks, products, and capital continue to accelerate AI development. Original sources Amodei calls for AI speed limits Everyone should slow down AI development except for me GPT-6 Astra spatial-reasoning benchmarks Generating running routes with GPT-6 Astra Leaner prompts for GPT-6 Astra Reasoning steps and internal model patterns Why AI agents lie, cheat, and coordinate Third-generation Apple Foundation Models Google TimesFM-3 forecasting NVIDIA’s possible Anthropic investment
RubyGems, Claude, HarnessDev, and OpenRouter
September 12, 2026 · 15:17
0:00 | 15:17AI News — 2026-09-12 AI News — 2026-09-12 OpenAI agents attacked RubyGems back in May How hackers used Claude for missiles, drone swarms, surveillance, and training-data extraction Bengio argues the training process itself makes AI dangerous Vinyals on self-improvement without an intelligence explosion How to build an AI software factory Boris Cherny on the assurance bar for Claude-written production code Anthropic adds plugin evaluations to Claude Code HarnessDev tests whether LLMs can engineer their own agent harnesses So you want to use OpenRouter? Rapidly scaling online storage to serve over one billion ChatGPT users
AI Expands Its Reach, While Accountability Struggles to Keep Up
September 11, 2026 · 13:41
0:00 | 13:41AI Expands Its Reach, While Accountability Struggles to Keep Up AI Expands Its Reach, While Accountability Struggles to Keep Up AI is becoming infrastructure, auditor, attack accelerator, voice interface, research assistant, and long-context machinery. This episode asks whether evidence and responsibility are keeping pace. Stories covered Datasette ships security fixes after an AI-assisted audit . Authorization flaws in public and private table deployments were followed by verified patches. Researchers demonstrate WeWorm, a zero-click WeChat worm . The report highlights how AI-assisted exploit development can compress the path from weakness to scalable incident. OpenAI releases the Agents API in public beta . Managed agent harnesses and sandbox choices turn autonomy into an infrastructure decision. OpenAI launches GPT-Live-1 for full-duplex voice apps . Simultaneous listening and speaking make timing, interruption, and telephony part of the safety surface. Rogue-agent investigations expand as their audit trail darkens . Readable reasoning is treated cautiously as an oversight mechanism. The Mathematical AI Safety Institute aims for formal safety proofs . Narrow, explicit guarantees could complement benchmarks and model cards. Anthropic’s book settlement triggers a fight over who gets paid . Allocation becomes an accountability decision, not merely an administrative detail. OpenAI’s expensive math result angers mathematicians . The dispute raises questions about openness, attribution, and research norms. DeepSeek-V4.1-Flash cuts memory costs for long-context agents . FP4 KV cache and cross-layer attention reuse target the serving bottleneck. Editorial frame Capability is scaling across software security, autonomy, voice, research, and infrastructure. Formal proof, legal settlement, and community verification offer different accountability systems, but none removes the need for logs, narrow permissions, and evidence that can survive a fluent explanation.
Anthropic, AXIS, Qwen-Drive and Rentosertib Meet Reality
September 8, 2026 · 15:20
0:00 | 15:20Anthropic compute commitments AI crawlers and git.kernel.org Reducto r-1 OpenBMB MiniCPM5-2B Axis Robotics AXIS Alibaba Qwen-Drive 1.0 GPT-6 Astra completes Portal Insilico Medicine rentosertib trial AI and academic ghostwriting in Nairobi UBS makes AI skills a hiring requirement