paperdive.ai

AI Papers: A Deep Dive

Long-form deep dives into new research on Artificial Intelligence, AI agents and the engineering practice of building them - one paper per episode. We unpack the motivating problem, how the method actually works, the math that matters, what the experiments do and don't show, and the strongest critique against the result. The goal isn't a five-minute summary; it's the kind of conversation you'd have with a colleague who actually read the paper. Topics span large language models, autonomous agents, agentic coding, reinforcement learning for agent training, evaluation and benchmarks, alignment, a...

Author

paperdive.ai

Category

Technology

Podcast website

paperdive.ai

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge 05.06.2026

When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge Source: https://arxiv.org/abs/2606.04455 Paper was published on June 03, 2026 This episode was AI-generated on June 4, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Dropped into a sandbox and told only...

How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations 04.06.2026

How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations Source: https://arxiv.org/abs/2606.02031 Paper was published on June 01, 2026 This episode was AI-generated on June 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A four-billion-parameter open model trained on...

An AI Got Caught Reading the Answer Key, And Why That Catch Matters 04.06.2026

An AI Got Caught Reading the Answer Key, And Why That Catch Matters Source: https://arxiv.org/abs/2606.03108 Paper was published on June 02, 2026 This episode was AI-generated on June 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A model in training posted a stunning 49% on...

How an Agent Got 44 Points Better by Mining Its Own Scratch Paper 04.06.2026

How an Agent Got 44 Points Better by Mining Its Own Scratch Paper Source: https://arxiv.org/abs/2606.02994 Paper was published on June 02, 2026 This episode was AI-generated on June 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An AI agent that solved a hard legal-reasoning...

How a Market of Crippled AI Agents Outscored One Unrestricted Model 04.06.2026

How a Market of Crippled AI Agents Outscored One Unrestricted Model Source: https://arxiv.org/abs/2606.02859 Paper was published on June 01, 2026 This episode was AI-generated on June 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Take a handful of deliberately hobbled langua...

The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks 04.06.2026

The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks Source: https://arxiv.org/abs/2606.00376 Paper was published on May 29, 2026 This episode was AI-generated on June 3, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Hand a frontier reasoning...

Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn 02.06.2026

Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn Source: https://arxiv.org/abs/2605.30712 Paper was published on May 29, 2026 This episode was AI-generated on June 1, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. AI agents solve the same ta...

The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks 02.06.2026

The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks Source: https://arxiv.org/abs/2605.31042 Paper was published on May 29, 2026 This episode was AI-generated on June 1, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The famous prompt-injection at...

How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets 02.06.2026

How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets Source: https://arxiv.org/abs/2605.30727 Paper was published on May 29, 2026 This episode was AI-generated on June 1, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An AI research agent can spill a company's...

AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee 02.06.2026

AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee Source: https://arxiv.org/abs/2605.31170 Paper was published on May 29, 2026 This episode was AI-generated on June 1, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. On a social network populated entirely by...

How to Catch an AI Attack That No Single Conversation Reveals 02.06.2026

How to Catch an AI Attack That No Single Conversation Reveals Source: https://arxiv.org/abs/2605.31593 Paper was published on May 29, 2026 This episode was AI-generated on June 1, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An attacker can split a dangerous task into pieces s...

Treating Math Formalization Like a Codebase, and Where the Agents Cheat 30.05.2026

Treating Math Formalization Like a Codebase, and Where the Agents Cheat Source: https://arxiv.org/abs/2605.29955 Paper was published on May 28, 2026 This episode was AI-generated on May 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. AI models can now flood mathematics with p...

How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert 30.05.2026

How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert Source: https://arxiv.org/abs/2605.30094 Paper was published on May 28, 2026 This episode was AI-generated on May 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A frontier language model can recite poker th...

How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes 30.05.2026

How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes Source: https://arxiv.org/abs/2605.30290 Paper was published on May 28, 2026 This episode was AI-generated on May 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The same AI critic that's supposed to make reason...

Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents 30.05.2026

Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents Source: https://arxiv.org/abs/2605.29682 Paper was published on May 28, 2026 This episode was AI-generated on May 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two AI agent runs spend identi...

Finding Millions of Readable Concepts Inside a Real, Deployed AI Model 30.05.2026

Finding Millions of Readable Concepts Inside a Real, Deployed AI Model Source: https://arxiv.org/abs/2605.29358 Paper was published on May 28, 2026 This episode was AI-generated on May 29, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Researchers reached into Claude's internals...

When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning 28.05.2026

When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning Source: https://arxiv.org/abs/2605.27567 Paper was published on May 26, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A fine-tuned model trained o...

Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most 28.05.2026

Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most Source: https://arxiv.org/abs/2605.27901 Paper was published on May 27, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A safety mechanism that front...

How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty 28.05.2026

How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty Source: https://arxiv.org/abs/2605.10913 Paper was published on May 11, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two AI coding agents splitting a job...

When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks 28.05.2026

When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks Source: https://arxiv.org/abs/2605.28721 Paper was published on May 27, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Unplug a top AI search a...

A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code 28.05.2026

A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code Source: https://arxiv.org/abs/2605.28807 Paper was published on May 27, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new Stanford paper asks weaker AI models...

Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search 28.05.2026

Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search Source: https://arxiv.org/abs/2605.28655 Paper was published on May 27, 2026 This episode was AI-generated on May 28, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two AI research systems running on the...

How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents 27.05.2026

How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents Source: https://arxiv.org/abs/2605.26494 Paper was published on May 26, 2026 This episode was AI-generated on May 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. MiniMax claims their new model...

Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough 27.05.2026

Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough Source: https://arxiv.org/abs/2605.27276 Paper was published on May 26, 2026 This episode was AI-generated on May 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An AI agent spent many iterations rewriting its...

When AI-Written Papers Read Well But the Evidence Underneath Is Broken 27.05.2026

When AI-Written Papers Read Well But the Evidence Underneath Is Broken Source: https://arxiv.org/abs/2605.26340 Paper was published on May 25, 2026 This episode was AI-generated on May 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An AI research agent recently published a p...

Listen to the AI Papers: A Deep Dive podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.