paperdive.ai
AI Papers: A Deep Dive
Long-form deep dives into new research on Artificial Intelligence, AI agents and the engineering practice of building them - one paper per episode. We unpack the motivating problem, how the method actually works, the math that matters, what the experiments do and don't show, and the strongest critique against the result. The goal isn't a five-minute summary; it's the kind of conversation you'd have with a colleague who actually read the paper. Topics span large language models, autonomous agents, agentic coding, reinforcement learning for agent training, evaluation and benchmarks, alignment, a...
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment 12.05.2026 29:22
Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment Source: https://arxiv.org/abs/2605.07937 Paper was published on May 08, 2026 This episode was AI-generated on May 11, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A clarifying question worth nothing at actio...
A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking 12.05.2026 23:16
A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking Source: https://arxiv.org/abs/2605.00206 Paper was published on April 30, 2026 This episode was AI-generated on May 9, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if a transformer d...
Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval 12.05.2026 23:59
Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval Source: https://arxiv.org/abs/2605.06997 Paper was published on May 07, 2026 This episode was AI-generated on May 11, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A pure Mamba-2 scores 3% on the canonical associ...
Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead. 12.05.2026 23:32
Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead. Source: https://arxiv.org/abs/2605.06763 Paper was published on May 07, 2026 This episode was AI-generated on May 11, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Every popular trick for speeding up long-conte...
When Your AI Assistant Won't Let Go of Old Facts About You 09.05.2026 24:21
When Your AI Assistant Won't Let Go of Old Facts About You Source: https://arxiv.org/abs/2605.06527 Paper was published on May 07, 2026 This episode was AI-generated on May 9, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new benchmark called STALE shows that even frontier LL...
Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap 09.05.2026 30:20
Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap Source: https://arxiv.org/abs/2605.05846 Paper was published on May 07, 2026 This episode was AI-generated on May 9, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper shows that one or two...
Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper 09.05.2026 20:04
Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper Source: https://arxiv.org/abs/2605.06651 Paper was published on May 07, 2026 This episode was AI-generated on May 8, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Google DeepMind just ship...
Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization 09.05.2026 22:59
Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization Source: https://arxiv.org/abs/2605.06639 Paper was published on May 07, 2026 This episode was AI-generated on May 8, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A 30-billion-parameter open model keeps pac...
When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure 09.05.2026 30:11
When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure Source: https://arxiv.org/abs/2605.06068 Paper was published on May 07, 2026 This episode was AI-generated on May 8, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if the reason we use general-purpose s...
What RL Actually Does to Language Models, at the Token Level 09.05.2026 23:47
What RL Actually Does to Language Models, at the Token Level Source: https://arxiv.org/abs/2605.06241 Paper was published on May 07, 2026 This episode was AI-generated on May 8, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper argues that reinforcement learning on math...
The Missing Gradient Term That Predicts Sycophancy in RLHF 08.05.2026 21:58
The Missing Gradient Term That Predicts Sycophancy in RLHF Source: https://arxiv.org/abs/2605.04266 Paper was published on May 05, 2026 This episode was AI-generated on May 7, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper argues that sycophancy, hallucination, and r...
An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work 07.05.2026 21:51
An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work Source: https://arxiv.org/abs/2605.05000 Paper was published on May 06, 2026 This episode was AI-generated on May 7, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Microsoft just paid $140,000 in bug bounties...
Why a Small Agent Confidently Overwrites Memories It Doesn't Understand 07.05.2026 23:27
Why a Small Agent Confidently Overwrites Memories It Doesn't Understand Source: https://arxiv.org/abs/2605.03354 Paper was published on May 05, 2026 This episode was AI-generated on May 7, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When a tiny language model running an agent...
Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap 07.05.2026 32:08
Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap Source: https://arxiv.org/abs/2605.02087 Paper was published on May 03, 2026 This episode was AI-generated on May 6, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if the careful philosophy documen...
Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents 07.05.2026 14:05
Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents Source: https://arxiv.org/abs/2605.04036 Paper was published on May 05, 2026 This episode was AI-generated on May 6, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A university team fine-tuned an open-weig...
The Compliance Gap: Why AI Says Yes and Does No 06.05.2026 27:49
The Compliance Gap: Why AI Says Yes and Does No Source: https://arxiv.org/abs/2605.01771 Paper was published on May 03, 2026 This episode was AI-generated on May 6, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Six frontier AI models, sixty sessions, and a zero percent complian...
When the Best Reward Model Trains the Worst Policy: Inside EvoLM 06.05.2026 25:52
When the Best Reward Model Trains the Worst Policy: Inside EvoLM Source: https://arxiv.org/abs/2605.03871 Paper was published on May 05, 2026 This episode was AI-generated on May 6, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A 1.7B-parameter judge, handed the right rubric, e...
Language Models Compute the Rational Move, Then Override It 06.05.2026 29:16
Language Models Compute the Rational Move, Then Override It Source: https://arxiv.org/abs/2604.27167 Paper was published on April 29, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two language models playing Prisoner's Dilemma both i...
When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers 03.05.2026 31:25
When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers Source: https://arxiv.org/abs/2604.06126 Paper was published on April 07, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The strongest frontier AI agent in...
Why Your Coding Agent Stalls While the GPU Runs Hot 03.05.2026 23:56
Why Your Coding Agent Stalls While the GPU Runs Hot Source: https://arxiv.org/abs/2604.26963 Paper was published on April 14, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Modern LLM serving stacks were built for chatbots, and agents...
The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests 03.05.2026 21:09
The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests Source: https://arxiv.org/abs/2604.27633 Paper was published on April 30, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When a frontier l...
Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1 03.05.2026 32:17
Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1 Source: https://arxiv.org/abs/2604.06506 Paper was published on April 07, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A frontier coding agent given full ac...
Why Search Keeps Rediscovering the Same Workflow, and What That Means 03.05.2026 22:03
Why Search Keeps Rediscovering the Same Workflow, and What That Means Source: https://arxiv.org/abs/2604.25012 Paper was published on April 27, 2026 This episode was AI-generated on May 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper argues that the elaborate searc...
Why AI Coding Agents Keep Trying to Debug Without a Debugger 03.05.2026 20:47
Why AI Coding Agents Keep Trying to Debug Without a Debugger Source: https://arxiv.org/abs/2603.22048 Paper was published on March 23, 2026 This episode was AI-generated on May 2, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Today's AI coding agents try to fix bugs by reading...
When RL Actually Teaches Agents Something New, And When It Doesn't 03.05.2026 22:56
When RL Actually Teaches Agents Something New, And When It Doesn't Source: https://arxiv.org/abs/2604.14877 Paper was published on April 16, 2026 This episode was AI-generated on May 2, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A widely-cited result said reinforcement learn...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.