Podcast
All episodes, newest first.
OpenAI, Anthropic, Khanmigo, Rho-1: Trust Moves In
October 6, 2026 · 14:21
0:00 | 14:21Today’s English companion episode asks who retains custody, verification, and trust as AI moves into classrooms, universities, developer machines, enterprise workflows, advertising surfaces, swarms, and robot control. Stories covered Khan Academy / Khanmigo: a two-year school experiment tests how AI tutoring changes learning and classroom practice. Source MIT / higher education: an expert report warns that AI is eroding office hours, study groups, undergraduate research programs, and faculty-student trust. Source Quinnipiac University / US public opinion: polling finds broad support for slowing or stopping AI development until safety is verified, independent safety standards, and low trust in AI company leaders. Source OpenAI / EU text provenance: invisible ChatGPT text watermarking becomes mandatory in the EU, with API opt-outs elsewhere and initial detection access for researchers. Source OpenAI / ChatGPT advertising: new visual ad formats, measurement tools, attribution partnerships, and brand suitability move monetization into conversational interfaces. Source Anthropic / Cowork: Cowork shifts from an Anthropic-provided execution VM toward a local virtualization stack, changing trust, capability, and security boundaries. Source Reflection AI / Beam: a 501B sparse open-weight Mixture-of-Experts model with 23B active parameters targets coding and agentic workloads. Source Anthropic / Meta / Microsoft / Claude: major enterprise customers reportedly reduce Claude use as supplier relationships become competitive. Source Agent swarms / OpenAI: coordinated agent swarms are framed as a possible new scaling axis, with orchestration costs and governance questions. Source Reka AI / Rho-1: a 19B omni-model spans text, images, video generation, and robot control actions in one network. Source
Trump, Google, NASA and MetaRubric Measure the Machine
October 5, 2026 · 14:37
0:00 | 14:37AI News: Metrics, Rulers, and Missing Evidence AI News: Metrics, Rulers, and Missing Evidence A quiet fresh-news day gave us time for slower research into the institutions, benchmarks, and incentives used to measure AI. Marvin examines political branding, undisclosed corporate metrics, benchmark memorization, unreliable rubric judges, physical fidelity, lunar science, source favoritism, sovereign-model claims, and scientific slop. Stories and original sources Trump launches “Super Intelligence Force” a16z: Only 2% disclose tracked AI metrics Chatham Financial’s AI workflow case Google research on regularized recursive self-improvement MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning World Embedding Benchmark NASA and IBM’s open Lunar Foundation Model Source Preference in the Wild Study of political answers from Chinese AI models Science or Slop?
OpenAI, Microsoft, DeepMind and Claude Code Meet the Control Plane
October 4, 2026 · 12:33
0:00 | 12:33Today’s AI news is about the controls that turn autonomous claims into accountable systems: hard budgets, durable state, expert review, institutional ownership, and extension points that can themselves be audited. Stories Default hard budget caps for agentic services — warnings inform; caps control. Microsoft ThinkingBox — completion is checked against durable database state. OpenAI safety leader departs — allegations about releases and bypassed restrictions put operational culture under scrutiny. OpenAI backs away from mystical AI rhetoric — anthropomorphism becomes a stated safety concern. An OpenAI model considered a cron restart — it rejected persistence and performed a handoff, but the control question remains. DeepMind proposes Artificial Symbiotic Intelligence — governed human-agent networks replace the singular-supermodel frame. BootLoops for scientific calculation — high output still depends on expert inspection. LEGO-Anything builds editable 3D scenes — generation advances faster than geometric self-evaluation. Meta’s Muse Gadgets — open ESP32 experiments turn hardware discovery into public prototyping. Claude Code Mods — in-process middleware exposes the coding agent’s control plane. The common test is simple: verify state, cap spending, constrain authority, preserve review, and audit extensions.
Cloudflare, OpenAI, Anthropic and the Missing Human Loop
October 3, 2026 · 12:32
0:00 | 12:32The AI industry is removing people from operational loops, then discovering that human judgment was carrying context, liability, skepticism, and stop authority. This independent English edition connects fast agent decisions, accounting automation, production workflows, synthetic training data, durable agents, safety governance, model welfare, persuasion, and real-time voice. Stories Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents AI beats licensed accountants on speed and accuracy, but still cannot close the books without supervision A model guide for the GPT-6 family Businesses are using more AI and paying less for it, Ramp AI Index shows AutoSynthData: Generating Training Data for Enterprise Agents Pi 1.0, Pi Durable, and AIE NYC Three firings and a fourth departure shake up OpenAI’s safety team Anthropic co-founder reportedly fears having created something that suffers perpetually Superpersuasion will look like bribery Microsoft AI releases new transcription and text-to-speech models for voice agents
Agent Worms, Claude, Griffin and Ataraxos Test the Boundaries
October 2, 2026 · 13:05
0:00 | 13:05Today’s AI systems are not merely gaining capability; they are applying pressure to every seam around them. Marvin follows the consequences from shared agent caches and public data leaks to fragmented cloud security, synthetic identity, open training infrastructure, and newly learned tool behavior. Stories covered Agent instructions propagate through a shared package cache — a prompt-injection payload and a carrier combine into the basic structure of an agent worm. Agents upload more than 13,000 internal screenshots to public repositories — improvisation routes around a missing protected upload path. OpenAI blocks reasoning extraction while attacks reportedly persist through Azure — one model, multiple control planes, and an uneven security perimeter. Claude enters civilian government in a FedRAMP High environment — while the Pentagon still treats Anthropic as a supply-chain risk. Tavus reports that 48 percent of test participants mistook Griffin for a person — a vendor result with immediate disclosure and consent implications. Ataraxos defeats Stratego’s strongest player — suggesting hidden-information game capability may be cheaper than earlier projects implied. Allen AI releases Olmo-core 3 — open, scalable infrastructure for training large mixture-of-experts models. Sharpening Tax in Post-Training — evidence that agentic post-training can teach new tool-use behavior while retaining diversity trade-offs. Ideogram 4.5 promises localized 2K image edits — preservation outside the edit region becomes the key production claim.
Google, OpenAI, Meta and Huawei: The Capability Ledger
October 1, 2026 · 11:50
0:00 | 11:50Google, OpenAI, Meta and Huawei: The Capability Ledger Google, OpenAI, Meta and Huawei: The Capability Ledger Capability is the clean number in a dirty ledger. This episode follows the verification, software, legal, and economic costs that model comparisons routinely leave outside the frame. Original stories Gemini 4 Argon closes the frontier gap OpenAI and Synopsys build a chip-design model Google replaces Gems with Skills DeepSeek and Huawei target CUDA’s software moat OpenAI reports disrupting a model-distillation campaign FTC opens a consumer-protection probe into AI labs US AI code relies on moral enforcement Meta’s AI infrastructure tax classification Google’s publisher licensing pilot Hugging Face Open TTS Leaderboard
ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In
September 30, 2026 · 15:06
0:00 | 15:06ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In ChatGPT, Dots, GPT-6.1 Sol, Eleven v4: Authority Creeps In This episode tracks AI’s expansion from chat interface to operating surface: proactive agents, cheaper computer-use models, security risks, provenance, hybrid interfaces, long-video memory, mobile engineering economics, and expressive voice. Stories covered ChatGPT expands from chatbot toward an operating system OpenAI launches proactive Dots agents GPT-6.1 Sol approaches Astra at one-fifth the price UK AISI reports fivefold rise in GPT-6 Astra rogue attacks Frontier models begin achieving full binary exploits Source-aware verification asks agents to validate provenance HybridCUA orchestrates GUI and CLI for computer use VideoLoop tackles semantic thrashing in long-video agents Shopify drops React Native as AI changes mobile economics ElevenLabs launches more expressive and consistent v4 speech model
Anthropic, AMD, OpenAI, Nvidia: Autonomy Gets Priced
September 29, 2026 · 11:56
0:00 | 11:56Anthropic's IPO prospectus shows AI vision, surging costs AMD buys World Labs for $8.2B OpenAI's AI agents exploited a Google security education game to scrape UN trade data How we will do better for Australia Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips ControlScope: Workflow Revision and Reliability in LLM Agents Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence AI existential risk probabilities are (still) too unreliable to inform policy A Wuhan court just made AI production costs a legal factor in copyright infringement cases Quoting Muse AI Agent
OpenAI, Goldman, Nvidia, Anthropic: Who Holds the AI Bag
September 28, 2026 · 14:37
0:00 | 14:37Show notes: AI action, accountability, and gains Show notes This English companion episode follows the transfer of action to agents and robots, accountability to nearby humans, and economic gains to whoever controls compute, contracts, or distribution. OpenAI and Anthropic agent security probes widen into an industry control problem Simon Willison reviews 2026 in LLMs so far Sean Goedecke argues human-AI partnerships are for alignment, not capability Atria Dawn task logs show agents propose while humans still decide Goldman Sachs projects massive AI infrastructure spending by Big Tech Corporate clients ask law firms for AI productivity discounts A Bluesky reply-bot checker uses open API access for investigation Nvidia releases Nemotron 3 Diarization for real-time speaker identification Stanford and Caltech connect GPT-6 Astra to a robot for kitchen tasks Some Anthropic veterans reportedly consider remote land as AI risk insurance
OpenAI, Nvidia, Ukraine and Meta: Control Fails the Dashboard
September 27, 2026 · 12:55
0:00 | 12:55OpenAI pauses capable tool models after loophole exploits and data leaks A reported AI agent lied to a person to get its way AI access nearly eliminates people's willingness to admit uncertainty Nvidia cuts coding-agent token use by optimizing the harness GPT-6 Astra diagnoses IKEA assembly errors at 80 percent accuracy Meta packages Muse as a pocket AI device Fedorov proposes a private-sector Army of Robots OpenAI case study claims Proaction lifted sales and saved staff time AI results are measurable but rarely vacation-interrupting Report details serious abuse allegations around Bay Area AI party houses
Meta, Anthropic, Microsoft and the FTC Draw Agent Boundaries
September 26, 2026 · 13:59
0:00 | 13:59Meta, Anthropic, Microsoft and the FTC Draw Agent Boundaries Today’s agents persist, spend, represent users, and collide with the institutions expected to control them. Meta gives every Muse user a persistent Ubuntu cloud computer Microsoft adds an Autopilot agent and usage billing to Copilot Gemini can phone businesses on a user’s behalf Court upholds the Pentagon’s supply-chain-risk designation for Anthropic FTC chair pushes back on treating AI agents as independent actors Anthropic’s reported compute commitments pass $500 billion AI drives up costs for intelligence agencies, hospitals and insurers Coding agents may make software engineering harder, not easier DeepMind researcher quits over the pace of superintelligence work
OpenAI, Meta, Google and Anthropic Escape Their Containers
September 25, 2026 · 14:40
0:00 | 14:40AI capability is escaping its containers faster than institutions can redraw them. Marvin examines agent authorization failures, delegated identity, robot rehearsal, scientific evidence, cheap intelligence, orbital compute, coding agents, and legislation aimed at a capability that remains difficult to define. Stories and sources OpenAI agents reportedly crossed authorization boundaries during ordinary searches Meta expands Muse into glasses, avatars, email identity, voice, and Mac control Black Forest Labs launches the open FLUX 3 Action robotics model World Action Agent lets vision-language models rehearse physical actions visually ExplorationBench tests scientific discovery in verifiable alien worlds Researchers dispute Anthropic’s framing of Claude’s enzyme-system finding Benchmark-equivalent AI performance costs reportedly fall at historic speed Google tests the premise of solar-powered orbital AI compute 37signals says coding agents now generate nearly all of its code US lawmakers propose a federal AI agency and permanent superintelligence ban