Yun Wu

Learning GenAI via SOTA Papers

This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.

Husk å besøke podkastens nettsted og støtte skaperen: podcasters.spotify.com

Forfatter

Yun Wu

Kategori

Technology

Podkastens nettsted

podcasters.spotify.com

Siste episode

7. okt 2026

Hvor kan du lytte?

Podkaster i appen Replaio Radio Kommer snart

Podkaster kommer snart til appen. Installer nå, og bli den første som ser en helt ny tilnærming til podkaster

Last ned på Google Play Installer gratis Android nesten 10 mill. nedlastinger · 4,8 i vurdering iOS snart

Episoder

EP325: Why robots have too much brain 24.07.2026

Title: Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Source: http://arxiv.org/abs/2606.27755v1 Summary: This paper presents an architectural efficiency breakthrough for Vision-Language-Action (VLA) models by systematically analyzing block redundancy and block sensitivity via a novel 'GateProbe' metric. It reveals that VLA language backbones are highly redundant, demo...

EP324: JERP synchronizes AI rules and neural weights 23.07.2026

Title: Joint Learning of Experiential Rules and Policies for Large Language Model Agents Source: http://arxiv.org/abs/2606.27136v1 Summary: This work introduces JERP, a novel agentic framework that simultaneously updates an external pool of natural-language rules and the model's parametric policy from the same interaction trajectories. By keeping prompt-based rules synchronized with the evolvi...

EP323: Giving AI Einstein s visual imagination 23.07.2026

Title: Einstein World Models Source: http://arxiv.org/abs/2606.26969v1 Summary: This paper proposes a blueprint for LLM-based reasoning systems that integrates visual-temporal rollouts directly into the reasoning trace as inspectable hypotheses. By extending tool-calling into the domain of visual thought experiments, it introduces a novel framework for grounding complex physical and counterfactual...

EP322: Why cliff tokens break AI math 22.07.2026

Title: Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Source: http://arxiv.org/abs/2606.25524v1 Summary: This paper identifies 'cliff tokens' as the exact single-token triggers that cause large language models to diverge into reasoning failures during multi-step mathematical tasks. By introducing a taxonomy of these failures and a targeted preference...

EP321: Measuring AI intelligence in bits 22.07.2026

Title: Agentic System as Compressor: Quantifying System Intelligence in Bits Source: http://arxiv.org/abs/2606.25960v1 Summary: This paper introduces a novel theoretical framework that quantifies agentic system intelligence through the lens of compression efficiency, linking agent capabilities like tool-use and search directly to codelength reduction. It provides a foundational methodology for ana...

EP320: Universal AI is mathematically impossible 21.07.2026

Title: World Models in Pieces: Structural Certification for General Agents Source: http://arxiv.org/abs/2606.24842v1 Summary:This paper introduces structural certification, a novel transition-local mathematical framework that guarantees goal-conditioned performance for general agents operating with segmented, non-universal world models. By proving tight error bounds, it establishes a foundational...

EP319: How TRUSTMEM Fixes Broken AI Memory 21.07.2026

Title: TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory Source: http://arxiv.org/abs/2606.25161v1 Summary: This work is foundational as it presents a novel preference-guided reinforcement learning framework to optimize trustworthy memory consolidation for long-term LLM agents. By utilizing a Memory Transition Verifier to mitigate omission, corruption, and ha...

EP318: Open Data Recipes for AI Agents 20.07.2026

Title: OpenThoughts-Agent: Data Recipes for Agentic Models Source: http://arxiv.org/abs/2606.24855v1 Summary: This paper establishes the first open, systematically ablated data curation pipeline and training recipes designed to generalize language models across diverse agentic benchmarks. By demonstrating strong training data scaling properties and releasing a high-performing 32B model, it provide...

EP317: The Architecture Of Genuine Artificial Agency 20.07.2026

Title: Critique of Agent Model Source: http://arxiv.org/abs/2606.23991v1 Summary: This paper establishes a crucial conceptual boundary between workflow-scaffolded automation and true endogenous agency, introducing the novel Goal-Identity-Configurator (GIC) architecture. It provides a foundational blueprint for general-purpose agent models by integrating hierarchical goal decomposition, identity ev...

EP316: Teaching robotaxis the biological urge to survive 19.07.2026

Title: Active Inference as the Test-Time Scaling Law for Physical AI Agents Source: http://arxiv.org/abs/2606.22813v1 Summary: This paper introduces a novel test-time scaling law for physical agents grounded in active inference and free energy minimization to handle out-of-distribution environments. By updating policies dynamically at test-time through variational inference, it unlocks continuous...

EP315: Teaching robots to think like scientists 19.07.2026

Title: Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence Source: http://arxiv.org/abs/2606.22449v1 Summary: This paper proposes a novel agentic cognitive framework that shifts embodied AI from simple predictive world modeling to self-evolving epistemic intelligence. It introduces integrated loops for causal discovery, intervention-driven reasoning, an...

EP314: Why AI hacks its own geometry 18.07.2026

Title: All Routes Lead to Collapse Source: http://arxiv.org/abs/2606.22325v1 Summary: This paper presents a foundational geometric analysis demonstrating that representation collapse and attention sinks are inherent to content-based routing across diverse architectures, not just transformers. By showing this pathology persists in selective state-space models and recurrent mixers, it establishes cr...

EP313: How ARTS reasons through its own failures 18.07.2026

Title: Learning the ARTS of Search for Automated Discovery Source: http://arxiv.org/abs/2606.21891v1 Summary: This paper introduces ARTS, a novel agentic reasoning framework that uses LLMs to navigate search spaces by diagnosing execution failures and selecting hypotheses. It also presents a reasoning breakthrough by using test-time training to distill search tree knowledge directly into model wei...

EP312: BioMatrix translates English to 3D biology 17.07.2026

Title: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language Source: http://arxiv.org/abs/2606.22138v1 Summary: This paper presents BioMatrix, a novel decoder-only foundation model architecture that natively integrates text, structural, and sequence data into a shared token space. By eliminating external encoders, project...

EP311: Why AI Teams Hallucinate Together 17.07.2026

Title: Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems Source: http://arxiv.org/abs/2606.21666v1 Summary: This paper establishes a novel distributed systems framework for multi-agent LLMs by modeling hallucination as a consequence of context drift between independent agents. It introduces the Shared State Verification Protocol (SSVP) and Context Divergence Sco...

EP310: Why AI Breaks While Fixing Itself 16.07.2026

Title: Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning Source: http://arxiv.org/abs/2606.21724v1 Summary: This work introduces a novel test-time reasoning primitive by formulating LLM self-correction as an iterative denoising process across structured verify-judge-correct loops. By utilizing a binary judgment gate and role allocation, it achieves a sig...

EP309: AutoRAS builds self-healing AI agent networks 16.07.2026

Title: AutoRAS: Learning Robust Agentic Systems with Primitive Representations Source: http://arxiv.org/abs/2606.21445v1 Summary: This paper introduces a foundational framework for the automated design and optimization of agentic systems by representing workflows as sequences of symbolic primitives. By optimizing these configurations using safety signals and flow-based objectives, it shifts multi-...

EP308: Giving AI Agents Mathematical Muscle Memory 15.07.2026

Title: SoftSkill: Behavioral Compression for Contextual Adaptation Source: http://arxiv.org/abs/2606.20333v1 Summary: This paper introduces SoftSkill, a novel framework that compresses natural-language agent instructions into compact continuous context objects using soft-tuning on frozen LLM backbones. This represents a significant efficiency and architectural breakthrough for agentic systems by r...

EP307: AI agents now train physical robots autonomously 15.07.2026

Title: ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Source: http://arxiv.org/abs/2606.19980v1 Summary: This paper presents ENPIRE, a novel agentic self-improvement framework that automates real-world robot policy optimization through a closed physical feedback loop. By orchestrating environment resets, policy rollouts, and autonomous log analysis, it establishes a repeatable loo...

EP306: AIs that engineer their own pipelines 14.07.2026

Title: FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines Source: http://arxiv.org/abs/2606.19605v1 Summary: This paper introduces a novel optimization framework that autonomously diagnoses and refines both prompts and pipeline structures in multi-step LLM systems. By dynamically resolving bottlenecks through structural chain modifications, it provides a foundational method for...

EP305: Mathematical guardrails for autonomous AI agents 14.07.2026

Title: Deontic Policies for Runtime Governance of Agentic AI Systems Source: http://arxiv.org/abs/2606.19464v1 Summary: This paper proposes AgenticRei, a deontic policy framework that enables runtime governance, constraints, and deontic reasoning for autonomous multi-agent systems. By decoupling policy enforcement from the LLM via an external logic engine, it establishes a foundational security an...

EP304: MagicSim Bridges AI and Physics 13.07.2026

Title: MagicSim: A Unified Infrastructure for Executable Embodied Interaction Source: http://arxiv.org/abs/2606.17511v1 Summary: This paper is foundational because it provides a unified execution substrate and shared Markov Decision Process that bridges high-level agent planning with low-level physical control. By enabling deterministic multi-modal episode rollouts, evaluation, and annotation unde...

EP303: How MODE-RAG stops AI video lies 13.07.2026

Title: MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation Source: http://arxiv.org/abs/2606.17449v1 Summary: MODE-RAG proposes a multi-agent system driven by Variational Free Energy and Monte Carlo Tree Search to dynamically gate and guide interventions in multimodal RAG. It establishes a principled energy-based reasoning loop to quantify and mitigate c...

EP302: Transferable interaction patterns for web agents 12.07.2026

Title: Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns Source: http://arxiv.org/abs/2606.17645v1 Summary: This paper introduces SkillMigrator and Transferable Interaction Patterns, a novel framework that allows web agents to generalize interaction skills across diverse sites by matching layout structures rather than specific element references. This approach addresses a co...

EP301: VeriGraph Makes AI Data Analysis Verifiable 12.07.2026

Title: VeriGraph: Towards Verifiable Data-Analytic Agents Source: http://arxiv.org/abs/2606.16603v1 Summary: This paper introduces a novel neuro-symbolic reasoning framework that replaces linear text trajectories with explicit evidence-based directed acyclic graphs to ensure structural traceability. It provides a significant breakthrough in agentic verifiability by grounding natural-language claim...

Hør på podkasten Learning GenAI via SOTA Papers i Replaio

Radio og podkaster i én app - gratis og uten registrering. Installer i dag, og ikke gå glipp av lanseringen

Last ned på Google Play

Replaio er ikke podkastutgiver; programnavn, omslag og lyd tilhører opphavspersonene og distribueres via offentlige RSS-feeder.