Yun Wu
Learning GenAI via SOTA Papers
This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.
Be sure to visit the podcast's website and support the creator: podcasters.spotify.com
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
EP375: M2GDT solves multimodal knowledge graph completion 18.08.2026 23:59
Title: MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion Source: http://arxiv.org/abs/2607.15592v1 Summary: This paper proposes a new architectural primitive for multimodal generative AI by combining MLLMs, Diffusion Transformers, and a Relation-Adaptive Mixture-of-Experts. This innovative architecture offers a foundational...
EP374: TopoAgent Outperforms GPT-5 in Science 17.08.2026 20:51
Title: TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning Source: http://arxiv.org/abs/2607.14658v1 Summary: TopoAgent proposes a self-evolving topological agent framework, introducing novel agentic reasoning loops that leverage topological structures for multimodal scientific reasoning. This represents a foundational advancement in agent design, enabling more adaptiv...
EP373: Middle Layer Recurrence Fixes AI Amnesia 17.08.2026 15:03
Title: T^2MLR: Transformer with Temporal Middle-Layer Recurrence Source: http://arxiv.org/abs/2607.15178v1 Summary: This paper introduces T^2MLR, a novel Transformer architecture incorporating temporal recurrence directly within its middle layers. This design represents a new architectural primitive, enhancing Transformers' ability to process sequential information and potentially leading to s...
EP372: How UrbanAgent profiles unseen cities 16.08.2026 20:21
Title: Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling Source: http://arxiv.org/abs/2607.13558v1 Summary: This paper introduces a novel framework for multi-agent collaborative reasoning, significantly advancing Agentic AI by enabling agents to work together and leverage external tools. This approach is foundational for developing sophisticated AI agents...
EP371: Groc-PO Stops Multimodal AI Hallucinations 16.08.2026 23:28
Title: Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs Source: http://arxiv.org/abs/2607.13712v1 Summary: This work proposes a new preference optimization method, Groc-PO, specifically designed to enhance the truthfulness and grounding of Multimodal Large Language Models. This represents a significant breakthrough in improving the reliability and reducing hallucinati...
EP370: SLEUTH fixes AI multi-hop reasoning failures 15.08.2026 23:14
Title: Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents Source: http://arxiv.org/abs/2607.12267v1 Summary: This work introduces a novel agentic reasoning framework centered on 'Epistemic Working Memory' and a 'Track, Rank, Crack' mechanism. It provides a significant breakthrough in scaling multi-hop reasoning capabilities, which is crucial...
EP369: How Atomic Units Scale Intelligence 15.08.2026 23:24
Title: Atomic Units of X: The Compression Layer of Intelligence Source: http://arxiv.org/abs/2607.12634v1 Summary: This paper proposes a foundational concept of 'Atomic Units of X' as a compression layer, which could introduce new architectural primitives for intelligence itself. Such a fundamental theoretical framework has the potential to underpin the design of both future Generative AI...
EP368: Samba Framework for Audio-Visual Navigation 14.08.2026 20:35
Title: A Hybrid Mamba for Audio-Visual Navigation Source: http://arxiv.org/abs/2607.13110v1 Summary: This paper proposes a 'Hybrid Mamba' architecture, establishing a new architectural primitive or a significant evolution of the Mamba architecture for GenAI. This innovation promises substantial efficiency and capability gains across various domains, exemplified by its application to comple...
EP367: Why AI sounds so painfully corporate 14.08.2026 20:24
Title: Optimization Is Not All You Need Source: http://arxiv.org/abs/2607.11977v2 Summary: This work proposes a paradigm shift beyond traditional optimization-centric approaches in AI, suggesting new fundamental principles for learning and intelligence. By challenging a core tenet of modern AI, it could unlock significant reasoning and efficiency breakthroughs, impacting the very foundation of how...
EP366: Autonomous AI Writes Its Own Hacking Tools 13.08.2026 23:59
Title: Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation Source: http://arxiv.org/abs/2607.11288v1 Summary: This work proposes a groundbreaking self-evolving agentic operating system (SE-AOS), presenting a new architectural primitive for agentic AI. This paradigm allows agents to autonomously manage and improve their foundational operational environment, signi...
EP365: Smarter managers beat bigger AI brains 13.08.2026 22:28
Title: Agentic Routing: The Harness-Native Data Flywheel Source: http://arxiv.org/abs/2607.11399v1 Summary: This paper introduces a foundational step-level routing paradigm that treats multi-model coordination within agent execution harnesses as a dynamic, state-conditioned systems problem. It establishes a harness-native data flywheel where execution traces and environmental feedback automaticall...
EP364: Capability Trees for Scalable AI Agents 12.08.2026 22:35
Title: A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery Source: http://arxiv.org/abs/2607.11138v1 Summary: This paper presents a novel hierarchical, stack-based execution architecture for agentic orchestration that models complex, nested agent states as Pushdown Automata. By substituting flat tool registries with localized stack frames and...
EP363: How Logos Architecture Stops AI Misevolution 12.08.2026 21:44
Title: LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans Source: http://arxiv.org/abs/2607.10878v1 Summary: LOGOS introduces a pluggable layer for self-evolution and governance within multi-agent frameworks, enabling verifiable human-agent loop engineering. This framework is foundational for controlling, auditing, and safely evolving AI agent teams at scale, addressing critical ques...
EP362: How Agentic-DPO fixes brittle AI agents 11.08.2026 25:53
Title: Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Source: http://arxiv.org/abs/2607.10601v1 Summary: Agentic-DPO proposes a lightweight offline policy optimization method that transforms expert trajectories into state-conditioned preference supervision for DPO-style training. This approach represents a significant breakthrough in efficiently training robust L...
EP361: How Riemannian geometry fixes AI reasoning 11.08.2026 21:51
Title: Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Source: http://arxiv.org/abs/2607.10169v1 Summary: This paper introduces Riemannian Isometric Policy Optimization (RIPO), a foundational mathematical advancement that corrects the geometric mismatch in standard PPO-Clip's Euclidean-based policy distance measurements. By ensu...
EP360: How ARMOR stops AI reasoning collapse 10.08.2026 21:39
Title: ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Source: http://arxiv.org/abs/2607.10481v1 Summary: This paper presents ARMOR, a novel post-training reinforcement learning framework that addresses the persistent issue of training instability and over-optimization in reasoning LLMs. By introducing active anchor rollouts from reference policies paired with a mixed optimizati...
EP359: Why your AI should forget 10.08.2026 21:47
Title: Shared Selective Persistent Memory for Agentic LLM Systems Source: http://arxiv.org/abs/2607.09493v1 Summary: This work introduces a foundational shared selective persistent memory architecture that solves the critical context-bloat and token-inefficiency bottlenecks in multi-session LLM agents. By isolating reusable context parameters and decoupling runtime data, the framework achieves dra...
EP358: Europe s Transparent Soofi S AI Blueprint 09.08.2026 20:01
Title: A Sovereign, Open-Source Foundation Model for German and English Source: http://arxiv.org/abs/2607.09424v1 Summary: This paper presents a novel hybrid MoE Mamba-Transformer architecture at scale, representing a significant architectural milestone for foundation model design. The hybrid model maintains a near-constant inference cache during context growth, demonstrating a major throughput an...
EP357: Copying Smart Experts Makes AI Worse 09.08.2026 19:31
Title: Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation Source: http://arxiv.org/abs/2607.08255v1 Summary: This work delivers a foundational breakthrough in LLM post-training by demonstrating that traditional supervised fine-tuning on teacher-generated data can degrade capable student models, whereas reinforcement learning wi...
EP356: CMA solves the visual token explosion 08.08.2026 19:15
Title: Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Source: http://arxiv.org/abs/2607.08497v1 Summary: This paper introduces a novel architectural framework for multimodal agents that replaces monolithic context-window scaling with an Episodic Visual Memory and selective reactivation system to prevent visual token explosion. By decoupling perception,...
EP355: RL builds compositional reasoning strategies 08.08.2026 18:33
Title: RL Post-Training Builds Compositional Reasoning Strategies Source: http://arxiv.org/abs/2607.07646v1 Summary: This paper provides foundational insights into how reinforcement learning post-training enables models to transition from simple skills to complex, multi-step compositional reasoning strategies. By demonstrating how RL systematically organizes and compresses primitive actions into s...
EP354: How AI Agents Code Their Own Habits 07.08.2026 23:52
Title: From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents Source: http://arxiv.org/abs/2607.07321v1 Summary: This work introduces a novel framework for self-evolving agents by enabling them to autonomously synthesize low-level, atomic actions into reusable, higher-order Standard Operating Procedures (SOPs). This dynamic tool-optimization...
EP353: How IGRPO stops AI search distractions 07.08.2026 20:45
Title: Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Source: http://arxiv.org/abs/2607.06223v1 Summary: This paper introduces a novel policy optimization framework (IGRPO) that dynamically allocates rollout budgets based on node-level information gain during tree-structured exploration. By unifying adaptive search-tree ex...
EP352: Hidden states predict AI agent failure 06.08.2026 12:21
Title: Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade Source: http://arxiv.org/abs/2607.06503v1 Summary: This paper introduces a novel framework that uses internal activation probes to detect and early-abort doomed agent trajectories, saving up to 47% of inference compute. This represents a significant efficiency breakthrough for agentic loops, addre...
EP351: Direct-OPD slashes AI reasoning compute costs 06.08.2026 19:59
Title: Weak-to-Strong Generalization via Direct On-Policy Distillation Source: http://arxiv.org/abs/2607.05394v1 Summary: This work introduces a novel post-training paradigm that transfers the reinforcement learning policy shift of a smaller, cheaper weak model as an implicit reward signal to a stronger target model. By bypassing the need for explicit reward modeling or expensive on-policy rollout...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.