Yun Wu
Learning GenAI via SOTA Papers
This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.
Husk å besøke podkastens nettsted og støtte skaperen: podcasters.spotify.com
Forfatter
Yun Wu
Kategori
Podkastens nettsted
Siste episode
6. okt 2026
Hvor kan du lytte?
Podkaster i appen Replaio Radio Kommer snartPodkaster kommer snart til appen. Installer nå, og bli den første som ser en helt ny tilnærming til podkaster
Episoder
EP474: How SkillGLoW Fixes AI Memory Hoarding 06.10.2026 23:52
Title: SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams Source: http://arxiv.org/abs/2609.02217v1 Summary:This research presents 'procedural-family skill consolidation' as a novel framework for self-improving agents. It addresses fundamental challenges in agent autonomy and generalization by enabling agents to learn, adapt, and consoli...
EP473: How six templates make AI 120x faster 06.10.2026 24:30
Title: Codebook Agent: Amortized Topology Design for LLM Multi-Agent SystemsSource: http://arxiv.org/abs/2609.02264v1 Summary: This paper introduces 'amortized topology design' as a novel architectural primitive for structuring interactions within LLM multi-agent systems. This approach promises significant efficiency and reasoning breakthroughs by optimizing how agents communicate and coor...
EP472: Why AI Reasons Better in Silence 05.10.2026 22:22
Title: Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs Source: http://arxiv.org/abs/2609.01117v1Summary: This work proposes a novel reasoning loop that employs recurrent refinement of latent representations to enhance reasoning capabilities, even when utilizing frozen large language models. This represents a significant breakthrough in improving t...
EP471: Self evolving autonomous agents with ARISE-RL 05.10.2026 22:49
Title: ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement LearningSource: http://arxiv.org/abs/2609.01058v1 Summary: This paper introduces a novel agentic reasoning framework that enables iterative self-evolution in AI agents using reinforcement learning and rubric-grounded evaluation. This provides a foundational mechanism for agents to autonomously learn, improve, and...
EP470: 1933 math fixes AI context limits 04.10.2026 27:11
Title: Higher-Dimensional Rotary Position Embedding Source: http://arxiv.org/abs/2608.29715v1 Summary: This paper introduces an extension to Rotary Position Embedding (RoPE), a critical architectural primitive in Transformer models. This advancement can significantly enhance GenAI by improving how models process and understand complex, multi-dimensional data, leading to more capable and efficient...
EP469: How AutoCRAT stops AI from overthinking 04.10.2026 23:15
Title: AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM ReasoningSource: http://arxiv.org/abs/2608.29988v1 Summary: This research proposes a novel framework for adaptively controlling stochasticity and computational resources during LLM reasoning. This represents a significant breakthrough in agentic reasoning, enabling more efficient, targeted, and adaptable decision...
EP468: Steering AI logic with latent vectors 03.10.2026 23:09
Title: Toward Latent Language Model Skills Steering and Optimization: An Empirical Study Source: http://arxiv.org/abs/2608.29459v1Summary: This research explores methods for steering and optimizing latent language model skills, offering a deeper level of control over LLM capabilities. This represents a significant reasoning breakthrough for GenAI by enabling more efficient and targeted utilization...
EP467: LiteSearch-VL gives small AI search powers 03.10.2026 23:21
Title: LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO Source: http://arxiv.org/abs/2608.29357v1 Summary: This paper introduces novel techniques like trajectory distillation and synthetic Step-DPO to develop small, efficient multimodal search agents. Such advancements offer a significant efficiency and scaling breakthrough for Agentic AI, making sop...
EP466: Sony OmniUE replaces typing with physical interaction 02.10.2026 24:27
Title: Omni-Interactive Universal EmbedderSource: http://arxiv.org/abs/2608.27044v1 Summary: This paper proposes a "universal embedder" capable of processing and unifying diverse, interactive data modalities into a singular representation. Such a primitive could fundamentally change how GenAI and Agentic AI systems perceive and represent the complex, multimodal world, enabling more versa...
EP465: AI agents finally ditch computer screens 02.10.2026 21:54
Title: ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions Source: http://arxiv.org/abs/2608.26991v1Summary: This work introduces a novel framework for agent-environment interaction, moving beyond brittle pixel-based methods to structured state and semantic actions. This paradigm shift is foundational for creating robust and intelligent AI agents capable of reliably per...
EP464: PROGROUTER manages the cost of AI agents 01.10.2026 22:08
Title: ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs Source: http://arxiv.org/abs/2608.25992v1 Summary: This work proposes a novel framework for orchestrating complex multi-agent LLM workflows. By intelligently guiding agent interactions based on progress and optimizing for quality-cost tradeoffs, it provides a foundational mechanism fo...
EP463: How PolyMemDB Solves AI Memory Hallucinations 01.10.2026 20:52
Title: PolyMemDB: A Polyglot Database System for AI Memory Management Source: http://arxiv.org/abs/2608.25577v1 Summary: This paper introduces a polyglot database system specifically designed for comprehensive AI memory management, addressing a critical and pervasive bottleneck for scaling Generative AI and Agentic AI systems. It provides foundational infrastructure that enables agents to handle v...
EP462: OpsHarness turns general AI into SRE experts 30.09.2026 20:15
Title: From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis Source: http://arxiv.org/abs/2608.25661v1 Summary:This work proposes a "self-evolving harness" framework that enables general agents to specialize and become experts in specific domains, exemplified by Root Cause Analysis. This represents a novel agentic reasoning loop for continuous self-improveme...
EP461: AsymSpec Slashes AI Agent Compute Costs 30.09.2026 14:47
Title: AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMsSource: http://arxiv.org/abs/2608.26004v1 Summary: This paper introduces a novel speculative decoding technique optimized for agentic LLMs. It represents a significant efficiency breakthrough by intelligently managing context to speed up LLM inference, which is crucial for the iterative reasoning processes of AI agents.
EP460: Teaching AI to Ignore Its Teacher 29.09.2026 21:55
Title: On-policy Distillation with Verifiable Reward Source: http://arxiv.org/abs/2608.24696v1Summary: This work presents a crucial advance in AI training by combining on-policy distillation with verifiable reward mechanisms. It offers a significant efficiency breakthrough for training generative models and agents while also addressing the foundational challenge of ensuring reliable and aligned AI...
EP459: Small models outsmart giants by coding 29.09.2026 21:11
Title: Joint Optimization of Tool Creation and Use for Large Language Model Agents Source: http://arxiv.org/abs/2608.24571v1 Summary:This paper introduces a foundational framework for LLM agents to not only utilize existing tools but also autonomously create new ones, then optimize both their creation and use. This represents a significant breakthrough in agentic reasoning, enabling more autonomou...
EP458: Building AI agents that know their limits 28.09.2026 22:24
Title: TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents Source: http://arxiv.org/abs/2608.22793v1Summary: This research introduces a novel agentic reasoning framework centered on a 'Self-Evolving Skill Bank,' which allows LLM agents to autonomously acquire, refine, and manage their capabilities. This directly addresses foundational challenges in Agentic AI by enabli...
EP457: How AI finally masters 3D space 28.09.2026 25:20
Title: Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation Source: http://arxiv.org/abs/2608.22757v1 Summary: This paper proposes a unified, object-centric model for spatial understanding and controllable generation, representing a new architectural primitive for GenAI. By enabling models to intrinsically reason about and manipulate individual objects,...
EP456: Why AI robot crews need a bouncer 27.09.2026 24:37
Title: Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs Source: http://arxiv.org/abs/2608.22657v1 Summary: This paper proposes a novel architecture for "Physical Agentic AI," demonstrating how LLMs can orchestrate a crew of robots. It represents a foundational breakthrough by integrating advanced language models into physical systems, enabling complex multi-r...
EP455: Stopping Coalition Pollution In AI Agents 27.09.2026 21:45
Title: Coalition-Aware Skill Reliability for Self-Evolving Agents Source: http://arxiv.org/abs/2608.22610v1 Summary: This research introduces a foundational framework for "self-evolving agents" with "coalition-aware skill reliability." It presents a novel approach for agents to learn, adapt, and assess their skills dynamically in collaborative settings, paving the way for more...
EP454: AI is eating the software stack 26.09.2026 20:07
Title: The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents Source: http://arxiv.org/abs/2608.20201v1 Summary: This paper proposes a fundamental architectural paradigm shift for software systems, moving from traditional three-tier designs to a structure centered on storage, models, and agents. This redefines the primitives for building future so...
EP453: PolicyGuide forces AI agents to follow workflows 26.09.2026 22:32
Title: PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsSource: http://arxiv.org/abs/2608.19861v1 Summary: This work introduces a novel agentic reasoning framework designed to guide LLM agents through entire workflows while ensuring continuous policy compliance. It represents a significant breakthrough in developing robust, reliable, and governable...
EP452: A Budget for Smarter AI Reasoning 25.09.2026 20:18
Title: Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models Source: http://arxiv.org/abs/2608.18884v1Summary: This paper presents a training-free, inference-time self-reflection protocol that enhances LLM reasoning through a prompt-level generate-critique-revise cycle. It introduces a cost-bounded early stopping mechanism using a confirmation senti...
EP451: Breaking the AI environment bottleneck with SPADE 25.09.2026 26:08
Title: SPADE: Self-Play in Adaptive Synthetic Executable Environments Source: http://arxiv.org/abs/2608.19197v1 Summary: This paper introduces a novel reinforcement learning framework for LLM agents where a single model plays dual roles as both environment designer and reasoning agent. By generating and interacting within its own synthetic executable environments, the framework provides a self-pla...
EP450: AI models grade each other to scale 24.09.2026 20:40
Title: Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Source: http://arxiv.org/abs/2608.17253v2 Summary:This paper presents a breakthrough in how unsupervised reasoning capabilities can emerge naturally within multi-agent reinforcement learning systems. It introduces a novel framework where diverse cohorts of agents learn to reason without explicit supervision, represe...
Lignende podkaster
Replaio er ikke podkastutgiver; programnavn, omslag og lyd tilhører opphavspersonene og distribueres via offentlige RSS-feeder.