Yun Wu

Learning GenAI via SOTA Papers

This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.

Be sure to visit the podcast's website and support the creator: podcasters.spotify.com

Author

Yun Wu

Category

Technology

Podcast website

podcasters.spotify.com

Latest episode

Oct 7, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

EP350: Training AI agents without live environments 05.08.2026

Title: Multi-Turn On-Policy Distillation with Prefix Replay Source: http://arxiv.org/abs/2607.04763v1 Summary: This paper addresses a critical training bottleneck for LLM agents by introducing an off-environment distillation method that resolves the 'prefix trap' in multi-turn interactions. By enabling scalable on-policy distillation with a 4x training speedup and zero tool calls, it estab...

EP349: Fixing AI judges with continuous verification 05.08.2026

Title: LLM-as-a-Verifier: A General-Purpose Verification Framework Source: http://arxiv.org/abs/2607.05391v1 Summary: This paper formalizes solution verification as a major new scaling axis for language models, introducing a framework that computes continuous scores over logit distributions rather than discrete judgments to evaluate complex reasoning. It establishes a highly versatile, training-fr...

EP348: Building AI agents like living cells 04.08.2026

Title: Biological Motifs for Agentic Control Source: http://arxiv.org/abs/2607.04240v1 Summary: This work establishes a formal mathematical framework that models multi-agent systems using control motifs from systems biology and category theory. By introducing the Agentic Operad, it provides a typed syntax for agent composition with provable error-suppression bounds and scaling laws for complex age...

EP347: Compiling AI into Permanent Free Skills 04.08.2026

Title: Auto: The AGI Compiler Source: http://arxiv.org/abs/2607.04542v1 Summary: This paper introduces Auto, a novel compiler framework that records live AI agent behavior and compiles deterministic sequences into verified WebAssembly binaries. This approach dramatically reduces the inference cost and latency of complex agentic workflows while maintaining reliability through calibrated fallback gu...

EP346: Teaching small AI to ignore teachers 03.08.2026

Title: Reward-Gated On-Policy Distillation Source: http://arxiv.org/abs/2607.04037v1 Summary: This paper proposes a novel on-policy distillation method that gates token-level teacher supervision using sparse verifier feedback to prevent the propagation of incorrect reasoning modes. It represents a significant optimization breakthrough for transferring complex reasoning capabilities from frontier m...

EP345: AI agents retry from pivotal mistakes 03.08.2026

Title: Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Source: http://arxiv.org/abs/2607.03702v1 Summary: This paper introduces PivoARL, a novel agentic reinforcement learning framework that optimizes agent trajectories by identifying and retrying only from pivotal erroneous states. By isolating correct prefixes and addressing credit assignment near error boundaries, it signific...

EP344: AI predicts tool calls to skip waiting 02.08.2026

Title: SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference Source: http://arxiv.org/abs/2607.03333v1 Summary: This paper addresses the system-level execution bottleneck of LLM agents by introducing Self-Speculative Forking (SPORK), a training-free controller that enables speculative tool execution. By using the model as its own predictor to dispatch tool calls early and overlap exe...

EP343: How AI agents escape infinite loops 02.08.2026

Title: No Time Like the Present: Agentic Test-Time Training for LLM Agents Source: http://arxiv.org/abs/2607.03441v1 Summary: This paper introduces Agentic Test-Time Training (aTTT), a novel framework that continuously adapts model weights during multi-turn agent episodes to prevent performance degradation over long trajectories. By utilizing a token-level reweighting method to mitigate self-train...

EP342: Why process rubrics triple AI accuracy 01.08.2026

Title: SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use Source: http://arxiv.org/abs/2607.01874v1 Summary: This paper presents SkillCoach, a framework that self-evolves process-based rubrics to evaluate and supervise complex agentic tool and skill workflows. By decoupling process quality from final outcomes, it provides a foundational mechanism for structured, multi...

EP341: Gemma 4 brings thinking mode to laptops 01.08.2026

Title: Gemma 4 Technical Report Source: http://arxiv.org/abs/2607.02770v1 Summary: This report details Gemma 4, introducing a native thinking mode that allows models to generate structured reasoning traces before responding to queries. It also presents novel compute-efficient Mixture-of-Experts architectures and a unified, encoder-free framework for multimodal ingestion, establishing key primitive...

EP340: AI Models Prove Opposite Scientific Truths 31.07.2026

Title: The Agentic Garden of Forking Paths Source: http://arxiv.org/abs/2607.01507v1 Summary: This paper introduces the Agentic Bootstrap, a foundational paradigm that uses persona-driven AI agents to systematically explore and map the 'multiverse' of alternative analytical decision paths. By formalizing this distribution of defensible choices, it establishes a novel method for validating...

EP339: How AI Safely Rewrites Its Own Code 31.07.2026

Title: Self-Evolving Agents with Anytime-Valid Certificates Source: http://arxiv.org/abs/2607.00871v1 Summary: This paper introduces SEA, a novel architectural framework designed to safely enable self-evolving and self-modifying AI agents by confining code changes to a steering adapter governed by anytime-valid certificates. It directly addresses a major safety and stability bottleneck in agentic...

EP338: DiscoPER conducts autonomous science via reflection 30.07.2026

Title: Autonomous Scientific Discovery via Iterative Meta-Reflection Source: http://arxiv.org/abs/2607.01131v1 Summary: This work proposes DiscoPER, an agentic framework featuring a novel second-order meta-reflection loop that allows agents to treat their own accumulated findings as empirical data to guide future exploration. It introduces a highly sophisticated reasoning primitive for open-ended...

EP337: Why AI Agents Fail in Silence 30.07.2026

Title: Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration Source: http://arxiv.org/abs/2606.31422v1 Summary: This work introduces a budgeted probing framework to address the critical issue of world-model drift in long-horizon planning agents. By formalizing environment feedback as a scarce calibration resource rather than just a task progression tool, it establis...

EP336: ACE fixes the AI goldfish memory problem 29.07.2026

Title: ACE: Pluggable Adaptive Context Elasticizer across Agents Source: http://arxiv.org/abs/2606.31564v1 Summary: This paper proposes a plug-and-play context management module that dynamically orchestrates historical agent trajectories into raw, abstracted, or dropped states. By enabling reversible context compression, it resolves a fundamental bottleneck in long-horizon AI agents where critical...

EP335: How AI agents learn from failure 29.07.2026

Title: Self-Evolving World Models for LLM Agent Planning Source: http://arxiv.org/abs/2606.30639v1 Summary: This work presents WorldEvolver, a novel framework that equips long-horizon LLM agents with self-evolving world models to enhance predictive planning. By integrating episodic and semantic memories to revise test-time reasoning context dynamically without parameter updates, it establishes a r...

EP334: Fixing AI Hallucinations With Process Rewards 28.07.2026

Title: SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution Source: http://arxiv.org/abs/2606.29713v1 Summary: This paper is foundational for Agentic AI as it introduces a novel Verify-Reflect-Probe-Refine self-evolution loop that enables agents to iteratively self-correct and improve. Furthermore, it addresses a key training bottleneck for multi-component agents by prop...

EP333: Logic not length makes AI smarter 28.07.2026

Title: Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters Source: http://arxiv.org/abs/2606.30128v1 Summary: This work presents a foundational breakthrough in understanding Large Language Model reasoning by investigating why Chain-of-Thought (CoT) prompting succeeds. Through systematic controlled interventions, it demonstrates that reasoning impro...

EP332: AI Architects Designing Better Embodied Agents 27.07.2026

Title: Automating the Design of Embodied AgentArchitectures Source: http://arxiv.org/abs/2606.30111v1 Summary: This research introduces a novel architectural primitive for agents by automating the discovery of optimal modular compositions through a typed-graph runtime and code-agent search procedure. By shifting from manual design to systematic architecture search, it provides a foundational metho...

EP331: Internalizing AI debate with Mixture of Debaters 27.07.2026

Title: Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Source: http://arxiv.org/abs/2606.29425v1 Summary: This paper introduces Mixture of Debaters (MoD), a novel framework that embeds multi-agent dialectical reasoning directly into a single model's Mixture-of-Experts architecture. By replacing multi-model communication with lightweight expert routing and d...

EP330: AI agents audit 10,000 page nuclear reports 26.07.2026

Title: LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents Source: http://arxiv.org/abs/2606.29399v1 Summary: This work proposes a novel agentic reasoning framework centered on LLM-guided planning for complex, multi-hop reasoning tasks over extensive document sets, demonstrating a significant breakthrough in agent capabilities. The emphasis on planning as the d...

EP329: Teaching AI to forget the right things 26.07.2026

Title: Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM Source: http://arxiv.org/abs/2606.29563v1 Summary: This paper introduces K-VEC, a coverage-aware KV-cache eviction strategy that resolves performance drops in long-context tasks by maximizing token retention across attention heads and layers. By theoretically linking token coverage to mutual information, it provid...

EP328: FlowWM and branching futures 25.07.2026

Title: Flow Matching in Feature Space for Stochastic World Modeling Source: http://arxiv.org/abs/2606.29059v1 Summary: This work introduces FlowWM, a novel stochastic world modeling architecture that successfully performs generative flow matching directly in high-dimensional pretrained feature spaces. By incorporating a differentiable one-step projection, it establishes a new primitive for plannin...

EP327: Why Chatbot Safety Training Backfires for Agents 25.07.2026

Title: Agent Safety Is Action Alignment Source: http://arxiv.org/abs/2606.28739v1 Summary: This paper redefines agent safety by identifying the category error of using chatbot refusal training for action-taking LLM agents. It establishes 'action alignment' enforced outside model weights via least privilege as the necessary paradigm to prevent the reasoning collapse of multi-step agents.

EP325: Why robots have too much brain 24.07.2026

Title: Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Source: http://arxiv.org/abs/2606.27755v1 Summary: This paper presents an architectural efficiency breakthrough for Vision-Language-Action (VLA) models by systematically analyzing block redundancy and block sensitivity via a novel 'GateProbe' metric. It reveals that VLA language backbones are highly redundant, demo...

Listen to the Learning GenAI via SOTA Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.