Yun Wu
Learning GenAI via SOTA Papers
This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.
Be sure to visit the podcast's website and support the creator: podcasters.spotify.com
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
EP324: JERP synchronizes AI rules and neural weights 23.07.2026 18:55
Title: Joint Learning of Experiential Rules and Policies for Large Language Model Agents Source: http://arxiv.org/abs/2606.27136v1 Summary: This work introduces JERP, a novel agentic framework that simultaneously updates an external pool of natural-language rules and the model's parametric policy from the same interaction trajectories. By keeping prompt-based rules synchronized with the evolvi...
EP323: Giving AI Einstein s visual imagination 23.07.2026 21:05
Title: Einstein World Models Source: http://arxiv.org/abs/2606.26969v1 Summary: This paper proposes a blueprint for LLM-based reasoning systems that integrates visual-temporal rollouts directly into the reasoning trace as inspectable hypotheses. By extending tool-calling into the domain of visual thought experiments, it introduces a novel framework for grounding complex physical and counterfactual...
EP322: Why cliff tokens break AI math 22.07.2026 18:26
Title: Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Source: http://arxiv.org/abs/2606.25524v1 Summary: This paper identifies 'cliff tokens' as the exact single-token triggers that cause large language models to diverge into reasoning failures during multi-step mathematical tasks. By introducing a taxonomy of these failures and a targeted preference...
EP321: Measuring AI intelligence in bits 22.07.2026 23:02
Title: Agentic System as Compressor: Quantifying System Intelligence in Bits Source: http://arxiv.org/abs/2606.25960v1 Summary: This paper introduces a novel theoretical framework that quantifies agentic system intelligence through the lens of compression efficiency, linking agent capabilities like tool-use and search directly to codelength reduction. It provides a foundational methodology for ana...
EP320: Universal AI is mathematically impossible 21.07.2026 23:24
Title: World Models in Pieces: Structural Certification for General Agents Source: http://arxiv.org/abs/2606.24842v1 Summary:This paper introduces structural certification, a novel transition-local mathematical framework that guarantees goal-conditioned performance for general agents operating with segmented, non-universal world models. By proving tight error bounds, it establishes a foundational...
EP319: How TRUSTMEM Fixes Broken AI Memory 21.07.2026 13:21
Title: TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory Source: http://arxiv.org/abs/2606.25161v1 Summary: This work is foundational as it presents a novel preference-guided reinforcement learning framework to optimize trustworthy memory consolidation for long-term LLM agents. By utilizing a Memory Transition Verifier to mitigate omission, corruption, and ha...
EP318: Open Data Recipes for AI Agents 20.07.2026 23:37
Title: OpenThoughts-Agent: Data Recipes for Agentic Models Source: http://arxiv.org/abs/2606.24855v1 Summary: This paper establishes the first open, systematically ablated data curation pipeline and training recipes designed to generalize language models across diverse agentic benchmarks. By demonstrating strong training data scaling properties and releasing a high-performing 32B model, it provide...
EP317: The Architecture Of Genuine Artificial Agency 20.07.2026 23:48
Title: Critique of Agent Model Source: http://arxiv.org/abs/2606.23991v1 Summary: This paper establishes a crucial conceptual boundary between workflow-scaffolded automation and true endogenous agency, introducing the novel Goal-Identity-Configurator (GIC) architecture. It provides a foundational blueprint for general-purpose agent models by integrating hierarchical goal decomposition, identity ev...
EP316: Teaching robotaxis the biological urge to survive 19.07.2026 18:40
Title: Active Inference as the Test-Time Scaling Law for Physical AI Agents Source: http://arxiv.org/abs/2606.22813v1 Summary: This paper introduces a novel test-time scaling law for physical agents grounded in active inference and free energy minimization to handle out-of-distribution environments. By updating policies dynamically at test-time through variational inference, it unlocks continuous...
EP315: Teaching robots to think like scientists 19.07.2026 21:59
Title: Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence Source: http://arxiv.org/abs/2606.22449v1 Summary: This paper proposes a novel agentic cognitive framework that shifts embodied AI from simple predictive world modeling to self-evolving epistemic intelligence. It introduces integrated loops for causal discovery, intervention-driven reasoning, an...
EP314: Why AI hacks its own geometry 18.07.2026 21:54
Title: All Routes Lead to Collapse Source: http://arxiv.org/abs/2606.22325v1 Summary: This paper presents a foundational geometric analysis demonstrating that representation collapse and attention sinks are inherent to content-based routing across diverse architectures, not just transformers. By showing this pathology persists in selective state-space models and recurrent mixers, it establishes cr...
EP313: How ARTS reasons through its own failures 18.07.2026 21:06
Title: Learning the ARTS of Search for Automated Discovery Source: http://arxiv.org/abs/2606.21891v1 Summary: This paper introduces ARTS, a novel agentic reasoning framework that uses LLMs to navigate search spaces by diagnosing execution failures and selecting hypotheses. It also presents a reasoning breakthrough by using test-time training to distill search tree knowledge directly into model wei...
EP312: BioMatrix translates English to 3D biology 17.07.2026 23:41
Title: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language Source: http://arxiv.org/abs/2606.22138v1 Summary: This paper presents BioMatrix, a novel decoder-only foundation model architecture that natively integrates text, structural, and sequence data into a shared token space. By eliminating external encoders, project...
EP311: Why AI Teams Hallucinate Together 17.07.2026 22:40
Title: Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems Source: http://arxiv.org/abs/2606.21666v1 Summary: This paper establishes a novel distributed systems framework for multi-agent LLMs by modeling hallucination as a consequence of context drift between independent agents. It introduces the Shared State Verification Protocol (SSVP) and Context Divergence Sco...
EP310: Why AI Breaks While Fixing Itself 16.07.2026 22:04
Title: Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning Source: http://arxiv.org/abs/2606.21724v1 Summary: This work introduces a novel test-time reasoning primitive by formulating LLM self-correction as an iterative denoising process across structured verify-judge-correct loops. By utilizing a binary judgment gate and role allocation, it achieves a sig...
EP309: AutoRAS builds self-healing AI agent networks 16.07.2026 24:56
Title: AutoRAS: Learning Robust Agentic Systems with Primitive Representations Source: http://arxiv.org/abs/2606.21445v1 Summary: This paper introduces a foundational framework for the automated design and optimization of agentic systems by representing workflows as sequences of symbolic primitives. By optimizing these configurations using safety signals and flow-based objectives, it shifts multi-...
EP308: Giving AI Agents Mathematical Muscle Memory 15.07.2026 19:44
Title: SoftSkill: Behavioral Compression for Contextual Adaptation Source: http://arxiv.org/abs/2606.20333v1 Summary: This paper introduces SoftSkill, a novel framework that compresses natural-language agent instructions into compact continuous context objects using soft-tuning on frozen LLM backbones. This represents a significant efficiency and architectural breakthrough for agentic systems by r...
EP307: AI agents now train physical robots autonomously 15.07.2026 22:02
Title: ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Source: http://arxiv.org/abs/2606.19980v1 Summary: This paper presents ENPIRE, a novel agentic self-improvement framework that automates real-world robot policy optimization through a closed physical feedback loop. By orchestrating environment resets, policy rollouts, and autonomous log analysis, it establishes a repeatable loo...
EP306: AIs that engineer their own pipelines 14.07.2026 20:46
Title: FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines Source: http://arxiv.org/abs/2606.19605v1 Summary: This paper introduces a novel optimization framework that autonomously diagnoses and refines both prompts and pipeline structures in multi-step LLM systems. By dynamically resolving bottlenecks through structural chain modifications, it provides a foundational method for...
EP305: Mathematical guardrails for autonomous AI agents 14.07.2026 25:35
Title: Deontic Policies for Runtime Governance of Agentic AI Systems Source: http://arxiv.org/abs/2606.19464v1 Summary: This paper proposes AgenticRei, a deontic policy framework that enables runtime governance, constraints, and deontic reasoning for autonomous multi-agent systems. By decoupling policy enforcement from the LLM via an external logic engine, it establishes a foundational security an...
EP304: MagicSim Bridges AI and Physics 13.07.2026 22:04
Title: MagicSim: A Unified Infrastructure for Executable Embodied Interaction Source: http://arxiv.org/abs/2606.17511v1 Summary: This paper is foundational because it provides a unified execution substrate and shared Markov Decision Process that bridges high-level agent planning with low-level physical control. By enabling deterministic multi-modal episode rollouts, evaluation, and annotation unde...
EP303: How MODE-RAG stops AI video lies 13.07.2026 22:38
Title: MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation Source: http://arxiv.org/abs/2606.17449v1 Summary: MODE-RAG proposes a multi-agent system driven by Variational Free Energy and Monte Carlo Tree Search to dynamically gate and guide interventions in multimodal RAG. It establishes a principled energy-based reasoning loop to quantify and mitigate c...
EP302: Transferable interaction patterns for web agents 12.07.2026 15:05
Title: Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns Source: http://arxiv.org/abs/2606.17645v1 Summary: This paper introduces SkillMigrator and Transferable Interaction Patterns, a novel framework that allows web agents to generalize interaction skills across diverse sites by matching layout structures rather than specific element references. This approach addresses a co...
EP301: VeriGraph Makes AI Data Analysis Verifiable 12.07.2026 22:51
Title: VeriGraph: Towards Verifiable Data-Analytic Agents Source: http://arxiv.org/abs/2606.16603v1 Summary: This paper introduces a novel neuro-symbolic reasoning framework that replaces linear text trajectories with explicit evidence-based directed acyclic graphs to ensure structural traceability. It provides a significant breakthrough in agentic verifiability by grounding natural-language claim...
EP300: Tensors prevent multi-agent LLM collisions 11.07.2026 22:26
Title: Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent LLM Planning Source: http://arxiv.org/abs/2606.16478v1 Summary: This work establishes a foundational algebraic primitive for multi-agent coordination by representing joint plans as third-order tensors and using Canonical Polyadic decomposition to identify latent conflicts. It introduces a computable co...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.