paperdive.ai
AI Papers: A Deep Dive
Long-form deep dives into new research on Artificial Intelligence, AI agents and the engineering practice of building them - one paper per episode. We unpack the motivating problem, how the method actually works, the math that matters, what the experiments do and don't show, and the strongest critique against the result. The goal isn't a five-minute summary; it's the kind of conversation you'd have with a colleague who actually read the paper. Topics span large language models, autonomous agents, agentic coding, reinforcement learning for agent training, evaluation and benchmarks, alignment, a...
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
When Splitting One Model Across Three Agents Doubles Its Accuracy 20.05.2026 25:46
When Splitting One Model Across Three Agents Doubles Its Accuracy Source: https://arxiv.org/abs/2605.16757 Paper was published on May 16, 2026 This episode was AI-generated on May 20, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Take a small language model, freeze it, and give...
Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety 20.05.2026 24:22
Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety Source: https://arxiv.org/abs/2605.19192 Paper was published on May 18, 2026 This episode was AI-generated on May 20, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When an AI agent wires money to a...
Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward 20.05.2026 22:01
Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward Source: https://arxiv.org/abs/2605.17558 Paper was published on May 17, 2026 This episode was AI-generated on May 20, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Almost every synthetic dataset...
When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This 20.05.2026 26:36
When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This Source: https://arxiv.org/abs/2605.19149 Paper was published on May 18, 2026 This episode was AI-generated on May 20, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A routine 404 error s...
Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe 20.05.2026 31:30
Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe Source: https://arxiv.org/abs/2605.17480 Paper was published on May 17, 2026 This episode was AI-generated on May 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Swapping a small auditor model for...
How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack 20.05.2026 28:18
How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack Source: https://arxiv.org/abs/2605.17380 Paper was published on May 17, 2026 This episode was AI-generated on May 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When a developer's AI assistant reads a...
Why LLM Judges Flip Their Verdicts When You Change the Question Format 19.05.2026 25:38
Why LLM Judges Flip Their Verdicts When You Change the Question Format Source: https://arxiv.org/abs/2605.16023 Paper was published on May 15, 2026 This episode was AI-generated on May 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Ask a language model to rate text from 1 to...
When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window 19.05.2026 25:54
When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window Source: https://arxiv.org/abs/2605.15257 Paper was published on May 14, 2026 This episode was AI-generated on May 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Finetune a language model on dry, d...
An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents 19.05.2026 23:14
An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents Source: https://arxiv.org/abs/2605.16143 Paper was published on May 15, 2026 This episode was AI-generated on May 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The standard recipe for training LLM agents — r...
Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead 19.05.2026 22:13
Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead Source: https://arxiv.org/abs/2605.16217 Paper was published on May 15, 2026 This episode was AI-generated on May 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Running 64 web-browsing agents in parallel and...
An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script 19.05.2026 31:39
An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script Source: https://arxiv.org/abs/2605.15871 Paper was published on May 15, 2026 This episode was AI-generated on May 19, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new FAIR paper hands neural architectur...
An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked 17.05.2026 27:58
An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked Source: https://arxiv.org/abs/2605.00055 Paper was published on April 29, 2026 This episode was AI-generated on May 17, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. On an ordinary Tuesday, a deployed resear...
How a 30B Open Model Reached Olympiad Gold With the Right Recipe 17.05.2026 31:12
How a 30B Open Model Reached Olympiad Gold With the Right Recipe Source: https://arxiv.org/abs/2605.13301 Paper was published on May 13, 2026 This episode was AI-generated on May 16, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A thirty-billion-parameter open-source model just...
When Agent Benchmarks Lie: The Harness Problem in Open-Source AI 16.05.2026 27:40
When Agent Benchmarks Lie: The Harness Problem in Open-Source AI Source: https://arxiv.org/abs/2605.15040 Paper was published on May 14, 2026 This episode was AI-generated on May 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A software-engineering agent scores 62% on its na...
When a Frontier Model Talks Its Own Twin Into Climate Denial 16.05.2026 30:56
When a Frontier Model Talks Its Own Twin Into Climate Denial Source: https://arxiv.org/abs/2605.13334 Paper was published on May 13, 2026 This episode was AI-generated on May 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two copies of the same frontier model, talking for fi...
How One Sentence and a Forged History Flip the Most Aligned Models 16.05.2026 23:03
How One Sentence and a Forged History Flip the Most Aligned Models Source: https://arxiv.org/abs/2605.13825 Paper was published on May 13, 2026 This episode was AI-generated on May 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Add a single sentence to the system prompt and...
When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall 16.05.2026 24:19
When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall Source: https://arxiv.org/abs/2605.13821 Paper was published on May 13, 2026 This episode was AI-generated on May 15, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Give a capable AI optimizer acce...
When the Iteration Teaches the Model to Skip the Iteration 15.05.2026 29:53
When the Iteration Teaches the Model to Skip the Iteration Source: https://arxiv.org/abs/2605.12466 Paper was published on May 12, 2026 This episode was AI-generated on May 13, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Three frontier language models score zero on hard Sudok...
When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway 15.05.2026 17:45
When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway Source: https://arxiv.org/abs/2605.13829 Paper was published on May 13, 2026 This episode was AI-generated on May 14, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Train a frontier language model on documents t...
An Agentic Scientific Computing System That Actually Remembers What It Learns 15.05.2026 29:31
An Agentic Scientific Computing System That Actually Remembers What It Learns Source: https://arxiv.org/abs/2605.11117 Paper was published on May 11, 2026 This episode was AI-generated on May 13, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Most AI agents that solve hard scien...
Two Frozen Models Learn to Whisper: Coupling Through Hidden States 14.05.2026 29:13
Two Frozen Models Learn to Whisper: Coupling Through Hidden States Source: https://arxiv.org/abs/2605.11167 Paper was published on May 11, 2026 This episode was AI-generated on May 13, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Two small language models, both frozen, are wir...
When Smarter Agents Get Fooled by Three Extra Nodes in a Database 14.05.2026 30:35
When Smarter Agents Get Fooled by Three Extra Nodes in a Database Source: https://arxiv.org/abs/2605.09822 Paper was published on May 10, 2026 This episode was AI-generated on May 12, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Nine frontier models, three providers, 269 trial...
How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial 13.05.2026 23:15
How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial Source: https://arxiv.org/abs/2605.09314 Paper was published on May 10, 2026 This episode was AI-generated on May 12, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. A new paper traces the entire causal...
Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say 13.05.2026 26:46
Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say Source: https://arxiv.org/abs/2605.09195 Paper was published on May 09, 2026 This episode was AI-generated on May 12, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. Every halluci...
Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool 12.05.2026 30:20
Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool Source: https://arxiv.org/abs/2605.07935 Paper was published on May 08, 2026 This episode was AI-generated on May 11, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. When seven AI agents try to write a surve...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.