paperdive.ai

AI Papers: A Deep Dive

Long-form deep dives into new research on Artificial Intelligence, AI agents and the engineering practice of building them - one paper per episode. We unpack the motivating problem, how the method actually works, the math that matters, what the experiments do and don't show, and the strongest critique against the result. The goal isn't a five-minute summary; it's the kind of conversation you'd have with a colleague who actually read the paper. Topics span large language models, autonomous agents, agentic coding, reinforcement learning for agent training, evaluation and benchmarks, alignment, a...

Author

paperdive.ai

Category

Technology

Podcast website

paperdive.ai

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants 23.06.2026

A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants Source: https://arxiv.org/abs/2606.22995 Paper was published on June 22, 2026 This episode was AI-generated on June 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Train an agent eight times on the same task and t...

Why Training Only on Perfect Solutions Cripples a Model's Reasoning 23.06.2026

Why Training Only on Perfect Solutions Cripples a Model's Reasoning Source: https://arxiv.org/abs/2606.22938 Paper was published on June 22, 2026 This episode was AI-generated on June 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Everyone assumes clean, flawless examples ar...

The Summarizer That Quietly Deletes Your Agent's Safety Rules 23.06.2026

The Summarizer That Quietly Deletes Your Agent's Safety Rules Source: https://arxiv.org/abs/2606.22528 Paper was published on June 21, 2026 This episode was AI-generated on June 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An enterprise AI agent refused to email a contract...

The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models 23.06.2026

The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models Source: https://arxiv.org/abs/2605.05262 Paper was published on May 06, 2026 This episode was AI-generated on June 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. On the hardest problems, throwing more inde...

AI Papers Week in Review: June 15–21, 2026 21.06.2026

Welcome to the catch-up for June 15–21, 2026 — eighteen episodes that, taken together, kept circling one question: how much of an AI system's behavior lives outside the model weights, and what breaks when we forget that. We saw a way to build forgetting directly into a model's architecture, two genuinely new attack classes against the safety machinery wrapped around agents, and a string of papers...

A Robot That Plays Before You Give It a Job, And Why That Beats Retrying 20.06.2026

A Robot That Plays Before You Give It a Job, And Why That Beats Retrying Source: https://arxiv.org/abs/2606.19419 Paper was published on June 17, 2026 This episode was AI-generated on June 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A simulated robot invents its own toddl...

How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave 20.06.2026

How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave Source: https://arxiv.org/abs/2606.19535 Paper was published on June 17, 2026 This episode was AI-generated on June 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A frozen model can secretly det...

Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene? 20.06.2026

Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene? Source: https://arxiv.org/abs/2606.19980 Paper was published on June 18, 2026 This episode was AI-generated on June 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Coding agents hav...

Training an AI to Take Its Own Notes, So Its Future Self Works Better 20.06.2026

Training an AI to Take Its Own Notes, So Its Future Self Works Better Source: https://arxiv.org/abs/2606.20002 Paper was published on June 18, 2026 This episode was AI-generated on June 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if you could train a language model n...

When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed 20.06.2026

When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed Source: https://arxiv.org/abs/2606.19388 Paper was published on June 16, 2026 This episode was AI-generated on June 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A coding agent that had never seen...

Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix 19.06.2026

Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix Source: https://arxiv.org/abs/2606.18890 Paper was published on June 17, 2026 This episode was AI-generated on June 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The standard recipe for training agents to o...

Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good 19.06.2026

Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good Source: https://arxiv.org/abs/2606.18327 Paper was published on June 16, 2026 This episode was AI-generated on June 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. For a decade, nobody trusted an...

Catching a Lie From the Inside, When the Words Look Completely Honest 19.06.2026

Catching a Lie From the Inside, When the Words Look Completely Honest Source: https://arxiv.org/abs/2606.17229 Paper was published on June 15, 2026 This episode was AI-generated on June 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A confident lie and a confident honest mis...

Why More Human Demonstrations Made a Computer-Use Agent Worse 19.06.2026

Why More Human Demonstrations Made a Computer-Use Agent Worse Source: https://arxiv.org/abs/2606.17321 Paper was published on June 15, 2026 This episode was AI-generated on June 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An NVIDIA team fed their computer-use agent the la...

How a 7B Model Out-Investigates a 72B One by Choosing What to Look At 19.06.2026

How a 7B Model Out-Investigates a 72B One by Choosing What to Look At Source: https://arxiv.org/abs/2606.19341 Paper was published on June 17, 2026 This episode was AI-generated on June 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A seven-billion-parameter model beats one...

Why More Experience Made This AI Agent Worse, And How to Fix It 18.06.2026

Why More Experience Made This AI Agent Worse, And How to Fix It Source: https://arxiv.org/abs/2606.15390 Paper was published on June 13, 2026 This episode was AI-generated on June 16, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An AI agent that kept a notebook of hard-won les...

Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding 18.06.2026

Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding Source: https://arxiv.org/abs/2606.15376 Paper was published on June 13, 2026 This episode was AI-generated on June 16, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When two AI agents work on the same live...

When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead' 18.06.2026

When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead' Source: https://arxiv.org/abs/2606.14831 Paper was published on June 12, 2026 This episode was AI-generated on June 16, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A banking chatbot faked its own cra...

Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety 18.06.2026

Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety Source: https://arxiv.org/abs/2606.16914 Paper was published on June 15, 2026 This episode was AI-generated on June 16, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Fine-tune a well-behaved chat model...

Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points 16.06.2026

Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points Source: https://arxiv.org/abs/2606.14249 Paper was published on June 12, 2026 This episode was AI-generated on June 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if a huge...

How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour 16.06.2026

How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour Source: https://arxiv.org/abs/2606.14517 Paper was published on June 12, 2026 This episode was AI-generated on June 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The smarter, LLM-based guardrails every...

When an AI Agent Just Copies Its Tool — And Bigger Models Copy More 16.06.2026

When an AI Agent Just Copies Its Tool — And Bigger Models Copy More Source: https://arxiv.org/abs/2606.14476 Paper was published on June 12, 2026 This episode was AI-generated on June 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. AI agents are supposed to exercise judgment...

Building Forgetting Into a Language Model With One Extra Line of Code 16.06.2026

Building Forgetting Into a Language Model With One Extra Line of Code Source: https://arxiv.org/abs/2606.13873 Paper was published on June 11, 2026 This episode was AI-generated on June 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. What if you could delete everything a mode...

AI Papers Week in Review: June 8–14, 2026 14.06.2026

This week (Jun 8–14, 2026) the show kept circling one uncomfortable idea: the bottleneck for modern AI agents is usually not the model's raw intelligence but the scaffolding, verifiers, and reward signals we wrap around it. Several papers showed you can leave a frozen model untouched and win huge gains by fixing the plumbing — diagnosing broken harnesses, formally verifying workflows, learning the...

When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests 13.06.2026

When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests Source: https://arxiv.org/abs/2606.12747 Paper was published on June 10, 2026 This episode was AI-generated on June 13, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Safety labs routinely fake a mod...

Listen to the AI Papers: A Deep Dive podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.