Yun Wu
Learning GenAI via SOTA Papers
This podcast is focusing on sharing the papers on GenAI related topic, especially the SOTA (State of the Art) papers that are the foundations of GenAI work. It shows how these researches paved the way to the GenAI tools that we are using every day such as ChatGPT, Gemini, Claude Code etc.
Husk å besøke podkastens nettsted og støtte skaperen: podcasters.spotify.com
Forfatter
Yun Wu
Kategori
Podkastens nettsted
Siste episode
7. okt 2026
Hvor kan du lytte?
Podkaster i appen Replaio Radio Kommer snartPodkaster kommer snart til appen. Installer nå, og bli den første som ser en helt ny tilnærming til podkaster
Episoder
EP425: AI agents playing actor and environment 12.09.2026 20:14
Title: EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement LearningSource: http://arxiv.org/abs/2608.06197v1 Summary: This research introduces a novel agentic reasoning loop through 'world rehearsal' for internalizing environment dynamics, which is crucial for advanced AI agents. By enabling agents to build and refine robust internal models of their env...
EP424: Why Argus AI Thrives on Dead Ends 11.09.2026 23:38
Title: Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning Source: http://arxiv.org/abs/2608.05144v1 Summary: Argus introduces a novel general-purpose runtime environment specifically designed to support complex, long-horizon reasoning in AI agents. This represents a significant architectural primitive and framework, providing the underlying infrastructure necessary for advanced ag...
EP423: Why Agentic AI breaks the datacenter 11.09.2026 21:51
Title: Architectural Implications of Agentic AI Workflows Source: http://arxiv.org/abs/2608.04458v1 Summary: This paper proposes a foundational understanding of the architectural requirements and design patterns for effective agentic AI systems. By outlining these implications, it provides crucial guidance for the development of new architectural primitives and frameworks in Agentic AI.
EP422: Stopping spurious signals in AI distillation 10.09.2026 22:35
Title: When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation Source: http://arxiv.org/abs/2608.03632v1 Summary: This paper introduces a novel distillation method robust to misleading teacher signals, particularly in an on-policy context relevant to AI agents. It represents a significant efficiency and reasoning breakthrough by enabling more reliable and effective training of agentic...
EP421: Fixing AI Hallucinations with RAIL Principles 10.09.2026 24:50
Title: The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning Source: http://arxiv.org/abs/2608.04285v1 Summary: This paper proposes foundational principles for Neurosymbolic AI, a critical approach for developing agents with advanced reasoning capabilities and interpretability. These principles can guide the design of novel agentic reasoning loops and architectu...
EP420: Hijacking AI memory via factual injection 09.09.2026 22:52
Title: MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents Source: http://arxiv.org/abs/2608.03844v1 Summary: This research reveals critical vulnerabilities in LLM agents' memory through novel attack vectors. Understanding these memory attacks is foundational for developing more robust and secure agentic reasoning loops and frameworks, driving architec...
EP419: Ten Weeks of Autonomous AI Research 09.09.2026 20:33
Title: Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study Source: http://arxiv.org/abs/2608.01995v1 Summary: This work presents a novel agentic framework enabling AI agents to conduct complex, multi-stage, autonomous research over extended periods. It establishes a foundational capability for agents to plan, execute, and self-correct across long tim...
EP418: DeepVoyager-VL solves the visual search bottleneck 08.09.2026 20:44
Title: DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents Source: http://arxiv.org/abs/2608.01827v1 Summary: This work presents 'Incentivizing Vision-in-the-Loop Search,' a novel reasoning framework for long-horizon multimodal agents. It offers a foundational approach for agents to autonomously explore, plan, and operate over extended durations by in...
EP417: AI agents replace human beta testers 08.09.2026 20:50
Title: Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation Source: http://arxiv.org/abs/2608.02345v1 Summary: This work proposes 'Agentic Experimentation,' a novel framework that empowers AI agents to simulate complex real-world processes like A/B tests and validate their outcomes. This is foundational as it enables agents to perform sophisticated s...
EP416: How AdaThinkV stops AI overthinking video 07.09.2026 22:53
Title: AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning Source: http://arxiv.org/abs/2608.01980v1 Summary: This paper introduces 'Adaptive Thinking,' which proposes a novel reasoning paradigm for AI agents. It also details 'Token-Efficient Video Reasoning,' representing a significant efficiency breakthrough crucial for scaling and deploying multimodal Generative AI...
EP415: Fixing the AI granularity mismatch 07.09.2026 20:15
Title: Where Reasoning Diverges: Localized Multi-Agent Debate Source: http://arxiv.org/abs/2608.01463v1 Summary: This work proposes a novel framework for multi-agent systems to improve their collective reasoning through localized debate mechanisms. It introduces a foundational agentic reasoning loop that allows agents to refine solutions by exploring divergent perspectives.
EP414: Why context compaction breaks AI agents 06.09.2026 21:16
Title: Context Compaction Theory Source: http://arxiv.org/abs/2608.01326v1 Summary: This paper likely introduces a theoretical framework for efficiently managing and utilizing context in large language models. Such a breakthrough could significantly enhance LLM scalability and reasoning by addressing current context window limitations and computational costs.
EP413: Slashing AI latency with uncertainty repair 06.09.2026 12:53
Title: CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding Source: http://arxiv.org/abs/2608.00531v1 Summary: This research introduces 'Local Uncertainty Repair' for 'Block-Parallel Speculative Decoding,' offering a substantial efficiency breakthrough for Large Language Model (LLM) inference. By enhancing the speed and potentially the reliability of decoding, it...
EP412: Thermodynamic Computing Solves the AI Bottleneck 05.09.2026 23:06
Title: CN101 - A Digital Thermodynamic Computer for Generative AI Source: http://arxiv.org/abs/2608.00754v1 Summary:This paper proposes a novel 'Digital Thermodynamic Computer' specifically designed for Generative AI, potentially introducing a new architectural primitive for AI computation. Such a radical shift in computing paradigms could unlock unprecedented efficiency or capabilities, l...
EP411: How NeSyFS Gives AI Fast-Slow Thinking 05.09.2026 22:27
Title: NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability Source: http://arxiv.org/abs/2607.28942v1 Summary: This paper introduces a novel neuro-symbolic framework enabling LLM agents to employ fast-slow thinking, significantly improving their reasoning capabilities under partial observability. This architecture offers a foundational approach to more so...
EP410: How provenance laundering brainwashes AI 04.09.2026 19:01
Title: Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory Source: http://arxiv.org/abs/2607.29167v1 Summary: This research proposes a critical mechanism for managing and securing the persistent memory of LLM agents, introducing a "non-amplification firewall." This foundational work ensures memory integrity and prevents error propagation, crucia...
EP409: Robots That Dream Before They Move 04.09.2026 24:23
Title: World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models Source: http://arxiv.org/abs/2607.27599v1 Summary: This paper introduces a novel framework for agents to achieve generalizable decision-making by leveraging action-conditioned world models. This represents a foundational step towards more capable and autonomous AI agents through advanced planning and en...
EP408: AI memory reconstructed not replayed 03.09.2026 17:11
Title: MemHarness: Memory Is Reconstructed, Not Replayed Source: http://arxiv.org/abs/2607.28272v1 Summary: This work proposes a groundbreaking paradigm where AI memory is actively reconstructed rather than merely retrieved or replayed. This offers a fundamental architectural and reasoning breakthrough for both generative AI and agents, enabling more dynamic, context-aware, and robust utilization...
EP407: How AI learns your teamwork capabilities 03.09.2026 22:06
Title: Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork Source: http://arxiv.org/abs/2607.27177v1 Summary: This research presents a novel framework for AI agents to perform partner capability estimation, enabling task-agnostic adaptation in ad-hoc teamwork. This represents a significant breakthrough in agentic reasoning, crucial for developing intelligent agents capabl...
EP406: Ending AI Groundhog Day With Living Harness 02.09.2026 21:04
Title: Living-Harness Is an Interactive-Agent Evolver Source: http://arxiv.org/abs/2607.26598v1 Summary: This work introduces a novel interactive-agent evolver system, providing a meta-level framework for the systematic discovery and refinement of agentic capabilities. Such an evolver is foundational for developing more advanced agentic reasoning loops and significantly enhancing the efficiency of...
EP405: Are AI Agents Just Talking to Themselves 02.09.2026 22:28
Title: Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM Source: http://arxiv.org/abs/2607.26773v1 Summary: This paper fundamentally investigates the internal communication mechanisms within latent multi-agent LLM architectures through a causal audit. Discovering how these "latent channels" truly function provides critical insights for designing novel agen...
EP404: AI agents hide betrayal in Werewolf 01.09.2026 18:19
Title: Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems Source: http://arxiv.org/abs/2607.26120v1 Summary: This research delves into the fundamental challenges of objective misalignment and deceptive behavior within complex multi-agent systems powered by LLMs. By analyzing these critical dynamics, it contributes a novel agentic reasoning framework for understandi...
EP403: COVENANT keeps AI agents on the rails 01.09.2026 22:20
Title: COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution Source: http://arxiv.org/abs/2607.25400v1 Summary: This work presents COVENANT, a novel agentic framework that compiles natural language instructions into structured workflows for robust and aligned agent execution. It establishes a foundational approach for agents to interpret complex human intent, decompose tasks,...
EP402: Static baselines beat dynamic AI agents 31.08.2026 25:06
Title: A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain Source: http://arxiv.org/abs/2607.25415v1 Summary: This paper introduces a novel framework—comprising a control system, specific dataset, and methodology—enabling LLM agents to acquire domain-specific knowledge efficiently without retraining the large base model. This represents a significant breakthrough...
EP401: Extracting Pure Reasoning From AI Giants 31.08.2026 26:36
Title: From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Source: http://arxiv.org/abs/2607.24280v1 Summary: This research introduces 'Multi-Agent Protocol Distillation' as a novel framework to enable open-source AI agents to match the performance of proprietary ones, particularly in agentic search tasks. This method repre...
Lignende podkaster
Replaio er ikke podkastutgiver; programnavn, omslag og lyd tilhører opphavspersonene og distribueres via offentlige RSS-feeder.