paperdive.ai
AI Papers: A Deep Dive
Long-form deep dives into new research on Artificial Intelligence, AI agents and the engineering practice of building them - one paper per episode. We unpack the motivating problem, how the method actually works, the math that matters, what the experiments do and don't show, and the strongest critique against the result. The goal isn't a five-minute summary; it's the kind of conversation you'd have with a colleague who actually read the paper. Topics span large language models, autonomous agents, agentic coding, reinforcement learning for agent training, evaluation and benchmarks, alignment, a...
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review 27.05.2026 25:32
When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review Source: https://arxiv.org/abs/2605.26174 Paper was published on May 25, 2026 This episode was AI-generated on May 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When long documents get partitione...
Why Frozen-Weight Agents Still Get Worse Over Time 27.05.2026 22:56
Why Frozen-Weight Agents Still Get Worse Over Time Source: https://arxiv.org/abs/2605.26302 Paper was published on May 25, 2026 This episode was AI-generated on May 27, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A deployed AI agent's model weights never change — but the agen...
When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence 26.05.2026 24:35
When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence Source: https://arxiv.org/abs/2605.24396 Paper was published on May 23, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper argues that...
Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick 26.05.2026 31:08
Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick Source: https://arxiv.org/abs/2605.24218 Paper was published on May 22, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. An academic lab just matched OpenAI...
Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction 26.05.2026 24:13
Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction Source: https://arxiv.org/abs/2605.26099 Paper was published on May 25, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. For two years the long-context modeling...
Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away 26.05.2026 25:50
Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away Source: https://arxiv.org/abs/2605.24517 Paper was published on May 23, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Standard agent RL throws away 85% of...
How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents 26.05.2026 31:43
How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents Source: https://arxiv.org/abs/2605.25624 Paper was published on May 25, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Computer-use agents have been stuck wh...
Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves 26.05.2026 23:31
Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves Source: https://arxiv.org/abs/2605.24486 Paper was published on May 23, 2026 This episode was AI-generated on May 26, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Most multi-agent s...
An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models 25.05.2026 28:42
An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models Source: https://arxiv.org/abs/2605.23384 Paper was published on May 22, 2026 This episode was AI-generated on May 25, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper takes John Flavell's 197...
Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training 25.05.2026 27:54
Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training Source: https://arxiv.org/abs/2605.23904 Paper was published on May 22, 2026 This episode was AI-generated on May 25, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A team at Micros...
Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math 25.05.2026 22:23
Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math Source: https://arxiv.org/abs/2605.22875 Paper was published on May 20, 2026 This episode was AI-generated on May 25, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A university r...
Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It 25.05.2026 22:27
Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It Source: https://arxiv.org/abs/2605.22873 Paper was published on May 20, 2026 This episode was AI-generated on May 25, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Telling a language model to 'think...
Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year 25.05.2026 28:29
Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year Source: https://arxiv.org/abs/2605.23109 Paper was published on May 22, 2026 This episode was AI-generated on May 25, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper claims to compress ni...
How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning 24.05.2026 21:07
How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning Source: https://arxiv.org/abs/2605.20613 Paper was published on May 20, 2026 This episode was AI-generated on May 24, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A team at Sapient Intelligence and...
A Robot Made Graphene Without Help, And Caught Itself Hallucinating 24.05.2026 28:45
A Robot Made Graphene Without Help, And Caught Itself Hallucinating Source: https://arxiv.org/abs/2605.18407 Paper was published on May 18, 2026 This episode was AI-generated on May 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. For twenty years, every graphene flake in ever...
When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving 24.05.2026 28:03
When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving Source: https://arxiv.org/abs/2605.17193 Paper was published on May 16, 2026 This episode was AI-generated on May 23, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Put three large language models in a room with...
When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions 23.05.2026 30:23
When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions Source: https://arxiv.org/abs/2605.22672 Paper was published on May 21, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Claude Opus 4.6 looked at Brazil's...
When Models Know the Answer But Say the Wrong Thing Anyway 23.05.2026 21:33
When Models Know the Answer But Say the Wrong Thing Anyway Source: https://arxiv.org/abs/2605.22007 Paper was published on May 21, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper shows that up to 47% of hallucinations from...
The OS Trick That Makes Tree Search Practical for Coding Agents 23.05.2026 26:59
The OS Trick That Makes Tree Search Practical for Coding Agents Source: https://arxiv.org/abs/2605.22781 Paper was published on May 21, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Almost nobody runs Monte Carlo tree search on real...
An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won 23.05.2026 31:01
An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won Source: https://arxiv.org/abs/2605.22763 Paper was published on May 21, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A Google DeepMind system autonomously cracked ni...
When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface 23.05.2026 23:15
When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface Source: https://arxiv.org/abs/2605.22166 Paper was published on May 21, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A four-billion-parameter model c...
Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer 22.05.2026 26:30
Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer Source: https://arxiv.org/abs/2605.12481 Paper was published on May 12, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Hand Claude 4.5 Sonnet a more powerful act...
One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery 22.05.2026 26:34
One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery Source: https://arxiv.org/abs/2605.19633 Paper was published on May 19, 2026 This episode was AI-generated on May 22, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Five separate LLM optimization frameworks...
When Agent Memory Stops Being a Database and Starts Being a Skill 22.05.2026 29:54
When Agent Memory Stops Being a Database and Starts Being a Skill Source: https://arxiv.org/abs/2605.20616 Paper was published on May 20, 2026 This episode was AI-generated on May 21, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two language agents solve the same science tasks...
Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency 22.05.2026 26:16
Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency Source: https://arxiv.org/abs/2605.21470 Paper was published on May 20, 2026 This episode was AI-generated on May 21, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Most of the time a web agent spends on your...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.