duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2025.11.25 | 即时编译让记忆无损;AutoEnv自动挑环境提两成 25.11.2025

本期的 15 篇论文如下: [00:25] 🧠 General Agentic Memory Via Deep Research(通过深度研究的通用代理记忆) [00:52] 🧪 AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning(AutoEnv:用于跨环境智能体学习的自动化环境测量) [01:24] 🤖 Computer-Use Agents as Judges for Generative User Interface(以计算机使用代理作为生成式用户界面的评判者) [01:55] 🎨 DeCo: Frequency-Decoupled Pi...

2025.11.24 | 开源7B模型刷新多模态推理;GeoVista小模型精准地理定位 24.11.2025

本期的 15 篇论文如下: [00:21] 🧠 OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe(OpenMMReasoner:以开放通用方案推动多模态推理前沿) [01:04] 🌍 GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization(GeoVista:用于地理定位的Web增强智能视觉推理) [01:41] 🎯 SAM 3: Segment Anything with Concepts(SAM 3:基于概念的通用分割模型) [02:31] 📊...

【周末特辑】11月第4周最火AI论文 | Kandinsky 5.0开源全家桶;MiroThinker开源智能体 22.11.2025

本期的 5 篇论文如下: [00:41] TOP1(🔥171) | 🎨 Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation(Kandinsky 5.0:用于图像和视频生成的基础模型家族) [02:02] TOP2(🔥150) | 🚀 MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling(MiroThinker:通过模型、上下文与交互扩展,将开源研究智能体性能推向新边界) [04...

2025.11.21 | V-ReasonBench考视频模型推理;Step-Audio-R1让语音越“想”越强 21.11.2025

本期的 15 篇论文如下: [00:22] 📊 V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models(V-ReasonBench:面向视频生成模型的统一推理基准套件) [01:06] 🧠 Step-Audio-R1 Technical Report(Step-Audio-R1技术报告) [01:48] 🧭 Scaling Spatial Intelligence with Multimodal Foundation Models(通过多模态基础模型扩展空间智能) [02:18] 🎬 First Frame Is the Place to Go for Video Co...

2025.11.20 | 视频模型拍推理链,迷宫百发百中;无标注左右互搏,视觉模型自学跃升 20.11.2025

本期的 4 篇论文如下: [00:23] 🎬 Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks(通过视频进行推理:基于走迷宫任务对视频模型推理能力的首次评测) [01:17] 🔄 VisPlay: Self-Evolving Vision-Language Models from Images(VisPlay:基于无标注图像自我进化的视觉-语言模型) [01:54] 📚 ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters a...

2025.11.19 | 像素演员难推理;视觉误导测真章 19.11.2025

本期的 11 篇论文如下: [00:23] 🧠 Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark(世界模拟器会推理吗?Gen-ViRe生成式视觉推理基准) [01:03] 🕵 MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs(MVI-Bench:评估大型视觉语言模型对误导性视觉输入鲁棒性的综合基准) [01:49] 🎞 REVISOR: Beyond Textual Reflection, Towards Multim...

2025.11.18 | RL奥赛夺金;Uni-MoE 2.0全能跃升 18.11.2025

本期的 14 篇论文如下: [00:17] 🏅 P1: Mastering Physics Olympiads with Reinforcement Learning(用强化学习攻克物理奥赛) [00:56] 🌐 Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data(Uni-MoE 2.0 Omni:以语言为中心的万模态大模型,通过先进MoE、训练与数据实现规模跃升) [01:42] 🧩 Part-X-MLLM: Part-aware 3D Multimodal Large Language Model(Part-X-MLLM...

2025.11.17 | RoPE去噪救长文本;AI速筛离子液体 17.11.2025

本期的 13 篇论文如下: [00:24] 🧹 DoPE: Denoising Rotary Position Embedding(DoPE:面向旋转位置嵌入的去噪处理) [00:58] 🧪 AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery(AIonopedia:面向离子液体发现的LLM智能体多模态学习编排) [01:44] 🖼 UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation(UI2Code^N:面向测试时可扩展交互...

【周末特辑】11月第3周最火AI论文 | 3D游戏智能体开源方案;桌面AI少样本精准操控 15.11.2025

本期的 5 篇论文如下: [00:38] TOP1(🔥135) | 🌍 Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds(Lumine:在3D开放世界中打造通才智能体的开源方案) [02:47] TOP2(🔥97) | 🖥 Grounding Computer Use Agents on Human Demonstrations(基于人类演示的计算机使用智能体定位研究) [04:44] TOP3(🔥89) | 🧠 Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Abili...

2025.11.14 | UniVA四合一开源视频通才;Depth Anything 3单ViT通吃3D 14.11.2025

本期的 4 篇论文如下: [00:24] 🎬 UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist(UniVA:面向开源下一代视频通才的通用视频智能体) [00:59] 🌐 Depth Anything 3: Recovering the Visual Space from Any Views(Depth Anything 3:从任意视角恢复视觉空间) [01:50] 🔍 AlphaResearch: Accelerating New Algorithm Discovery with Language Models(AlphaResearch:用语言模型加速全新算...

2025.11.13 | 原神数据炼成7B通用AI;零训练轨迹秒变视频遥控器 13.11.2025

本期的 9 篇论文如下: [00:19] 🌍 Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds(Lumine:在3D开放世界中打造通才智能体的开源方案) [00:54] 🎬 Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising(Time-to-Move:无需训练的双时钟去噪运动控制视频生成) [01:31] ⚡ TiDAR: Think in Diffusion, Talk in Autoregression(TiDAR:扩散式思考,自回归式表...

2025.11.12 | 1.5B小模型反超671B大模型;多智能体质检聊天机器人 12.11.2025

本期的 9 篇论文如下: [00:24] 🧠 Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B(小模型大逻辑:多样性驱动优化唤醒VibeThinker-1.5B的大模型推理力) [00:59] 🤝 Adaptive Multi-Agent Response Refinement in Conversational Systems(对话系统中自适应多智能体响应精炼机制) [01:30] 🧩 Wasm: A Pipeline for Constructing Structured Arabic Interleav...

2025.11.11 | 小窗口勤总结刷新深度研究;先广撒网再啃难题激活代码竞赛 11.11.2025

本期的 13 篇论文如下: [00:25] 🧩 IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction(IterResearch:基于马尔可夫状态重构的长程智能体再思考) [01:16] 🏆 DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation(DRIVE:面向可验证奖励强化学习的竞赛级代码生成数据精选最佳实践) [02:03] 🔬 The Station: An Open-World...

2025.11.10 | DeepEyesV2小模型边看图边写代码;纯数据让AI长出立体眼 10.11.2025

本期的 7 篇论文如下: [00:21] 🧠 DeepEyesV2: Toward Agentic Multimodal Model(DeepEyesV2:迈向智能体多模态模型) [01:13] 🧭 Visual Spatial Tuning(视觉空间调优) [01:54] 🦹 Too Good to be Bad: On the Failure of LLMs to Role-Play Villains(过于完美以致无法邪恶:大语言模型反派角色扮演的失败) [02:27] 🧠 Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings...

【周末特辑】11月第2周最火AI论文 | 视频生成即推理;SVG草图变代码 08.11.2025

本期的 5 篇论文如下: [00:31] TOP1(🔥137) | 🎬 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm(用视频思考:视频生成作为统一多模态推理新范式) [02:43] TOP2(🔥95) | 🖼 VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation(VCode:以SVG为符号视觉表征的多模态代码评测基准) [05:12] TOP3(🔥90) | 🚀 Diffusion Language Models are Super Data Lear...

2025.11.07 | 视频推理新范式;图像互动促思维 07.11.2025

本期的 12 篇论文如下: [00:21] 🎬 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm(用视频思考:视频生成作为统一多模态推理新范式) [00:58] 🧠 V-Thinker: Interactive Thinking with Images(V-Thinker:与图像互动的思维推理) [01:39] 🧠 Scaling Agent Learning via Experience Synthesis(基于经验合成的智能体规模化强化学习) [02:23] 🧠 Cambrian-S: Towards Spatial Supersens...

2025.11.06 | 扩散模型省数据;音视频对口型 06.11.2025

本期的 9 篇论文如下: [00:17] 🚀 Diffusion Language Models are Super Data Learners(扩散语言模型是超级数据学习者) [01:06] 🎬 UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions(统一音视频生成的不对称跨模态交互方法) [01:42] 🧩 LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation(LEGO-Eval:面向具身3D环境合成...

2025.11.05 | 向量草图测代码;先画后想补视觉 05.11.2025

本期的 15 篇论文如下: [00:21] 🖼 VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation(VCode:以SVG为符号视觉表征的多模态代码评测基准) [01:12] 🧠 When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought(当可视化成为推理第一步:MIRA视觉思维链基准测试) [01:48] ⚖ When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Pr...

2025.11.04 | 超稀疏MoE激活万亿参数;视觉模型看图胜GNN 04.11.2025

本期的 15 篇论文如下: [00:23] 🧠 Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation(全激活赋能:将通用推理模型扩展到万亿参数的开放语言基座) [01:03] 👁 The Underappreciated Power of Vision Models for Graph Structural Understanding(被低估的视觉模型在图结构理解中的强大潜能) [01:38] 💡 UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausib...

2025.11.03 | OS-Sentinel实时守护手机操作安全;ThinkMorph让小模型边想边画 03.11.2025

本期的 15 篇论文如下: [00:21] 🛡 OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows(OS-Sentinel:在真实工作流中通过混合验证提升移动GUI代理安全性) [01:13] 🧠 ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning(ThinkMorph:多模态交错思维链中的涌现特性) [01:49] ⚔ INT v.s. FP: A Comprehensive Study of Fine-Grained Lo...

【月末特辑】10月最火AI论文 | 幼龙BDH稀疏可解释;迷你递归7兆碾压大模型 02.11.2025

本期的 10 篇论文如下: [00:30] TOP1(🔥522) | 🐣 The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain(幼龙破壳: Transformer 与大脑模型之间缺失的环节) [02:31] TOP2(🔥462) | 🧠 Less is More: Recursive Reasoning with Tiny Networks(小而精:用微型网络递归推理) [04:48] TOP3(🔥255) | 🌱 Agent Learning via Early Experience(基于早期经验的主体学习) [07:04] TOP4(🔥182)...

【周末特辑】11月第1周最火AI论文 | 循环模型省参强推理;Concerto 2D-3D自监督涨点 01.11.2025

本期的 5 篇论文如下: [00:35] TOP1(🔥174) | 🔄 Scaling Latent Reasoning via Looped Language Models(通过循环语言模型扩展潜在推理能力) [02:30] TOP2(🔥166) | 🎼 Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations(Concerto:2D-3D联合自监督学习涌现空间表征) [05:17] TOP3(🔥115) | 🧩 ReCode: Unify Plan and Action for Universal Granularity Control(ReCode:用递归代码统一规划...

2025.10.31 | Emu3.5统一预测时空;扩散提示驱动机器人 31.10.2025

本期的 15 篇论文如下: [00:26] 🌍 Emu3.5: Native Multimodal Models are World Learners(Emu3.5:原生多模态世界模型让AI看懂并预测未来) [01:04] 🤖 Exploring Conditions for Diffusion models in Robotic Control(探索扩散模型在机器人控制中的条件化策略) [01:42] 🎬 Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark(视频模型已准备好做零样本推理了吗?基于MME-CoF基...

2025.10.30 | 看图写码7B逆袭;视频思维RL破局 30.10.2025

本期的 15 篇论文如下: [00:22] 👁 JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence(JanusCoder:面向代码智能的基础视觉-编程接口) [01:00] 🧠 Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning(Video-Thinker:用强化学习点燃“视频思维”) [01:55] 🔄 ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization(ReForm:...

2025.10.29 | 通义深度研究报告;小模型折记忆胜671B巨模型 29.10.2025

本期的 10 篇论文如下: [00:23] 🔍 Tongyi DeepResearch Technical Report(通义深度研究报告:面向长程深度信息检索任务的智能体大模型) [01:00] 🧠 AgentFold: Long-Horizon Web Agents with Proactive Context Management(AgentFold:面向长程任务的主动式上下文管理智能体) [01:36] 🤖 RoboOmni: Proactive Robot Manipulation in Omni-modal Context(RoboOmni:全模态上下文下的主动机器人操作) [02:33] 🎮 Game-TARS:...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.