duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2025.11.25 | 即时编译让记忆无损;AutoEnv自动挑环境提两成 25.11.2025 10:01
本期的 15 篇论文如下: [00:25] 🧠 General Agentic Memory Via Deep Research(通过深度研究的通用代理记忆) [00:52] 🧪 AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning(AutoEnv:用于跨环境智能体学习的自动化环境测量) [01:24] 🤖 Computer-Use Agents as Judges for Generative User Interface(以计算机使用代理作为生成式用户界面的评判者) [01:55] 🎨 DeCo: Frequency-Decoupled Pi...
2025.11.24 | 开源7B模型刷新多模态推理;GeoVista小模型精准地理定位 24.11.2025 10:42
本期的 15 篇论文如下: [00:21] 🧠 OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe(OpenMMReasoner:以开放通用方案推动多模态推理前沿) [01:04] 🌍 GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization(GeoVista:用于地理定位的Web增强智能视觉推理) [01:41] 🎯 SAM 3: Segment Anything with Concepts(SAM 3:基于概念的通用分割模型) [02:31] 📊...
【周末特辑】11月第4周最火AI论文 | Kandinsky 5.0开源全家桶;MiroThinker开源智能体 22.11.2025 10:19
本期的 5 篇论文如下: [00:41] TOP1(🔥171) | 🎨 Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation(Kandinsky 5.0:用于图像和视频生成的基础模型家族) [02:02] TOP2(🔥150) | 🚀 MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling(MiroThinker:通过模型、上下文与交互扩展,将开源研究智能体性能推向新边界) [04...
2025.11.21 | V-ReasonBench考视频模型推理;Step-Audio-R1让语音越“想”越强 21.11.2025 9:54
本期的 15 篇论文如下: [00:22] 📊 V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models(V-ReasonBench:面向视频生成模型的统一推理基准套件) [01:06] 🧠 Step-Audio-R1 Technical Report(Step-Audio-R1技术报告) [01:48] 🧭 Scaling Spatial Intelligence with Multimodal Foundation Models(通过多模态基础模型扩展空间智能) [02:18] 🎬 First Frame Is the Place to Go for Video Co...
2025.11.20 | 视频模型拍推理链,迷宫百发百中;无标注左右互搏,视觉模型自学跃升 20.11.2025 3:36
本期的 4 篇论文如下: [00:23] 🎬 Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks(通过视频进行推理:基于走迷宫任务对视频模型推理能力的首次评测) [01:17] 🔄 VisPlay: Self-Evolving Vision-Language Models from Images(VisPlay:基于无标注图像自我进化的视觉-语言模型) [01:54] 📚 ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters a...
2025.11.19 | 像素演员难推理;视觉误导测真章 19.11.2025 8:19
本期的 11 篇论文如下: [00:23] 🧠 Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark(世界模拟器会推理吗?Gen-ViRe生成式视觉推理基准) [01:03] 🕵 MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs(MVI-Bench:评估大型视觉语言模型对误导性视觉输入鲁棒性的综合基准) [01:49] 🎞 REVISOR: Beyond Textual Reflection, Towards Multim...
2025.11.18 | RL奥赛夺金;Uni-MoE 2.0全能跃升 18.11.2025 10:08
本期的 14 篇论文如下: [00:17] 🏅 P1: Mastering Physics Olympiads with Reinforcement Learning(用强化学习攻克物理奥赛) [00:56] 🌐 Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data(Uni-MoE 2.0 Omni:以语言为中心的万模态大模型,通过先进MoE、训练与数据实现规模跃升) [01:42] 🧩 Part-X-MLLM: Part-aware 3D Multimodal Large Language Model(Part-X-MLLM...
2025.11.17 | RoPE去噪救长文本;AI速筛离子液体 17.11.2025 10:06
本期的 13 篇论文如下: [00:24] 🧹 DoPE: Denoising Rotary Position Embedding(DoPE:面向旋转位置嵌入的去噪处理) [00:58] 🧪 AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery(AIonopedia:面向离子液体发现的LLM智能体多模态学习编排) [01:44] 🖼 UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation(UI2Code^N:面向测试时可扩展交互...
【周末特辑】11月第3周最火AI论文 | 3D游戏智能体开源方案;桌面AI少样本精准操控 15.11.2025 11:34
本期的 5 篇论文如下: [00:38] TOP1(🔥135) | 🌍 Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds(Lumine:在3D开放世界中打造通才智能体的开源方案) [02:47] TOP2(🔥97) | 🖥 Grounding Computer Use Agents on Human Demonstrations(基于人类演示的计算机使用智能体定位研究) [04:44] TOP3(🔥89) | 🧠 Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Abili...
2025.11.14 | UniVA四合一开源视频通才;Depth Anything 3单ViT通吃3D 14.11.2025 3:25
本期的 4 篇论文如下: [00:24] 🎬 UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist(UniVA:面向开源下一代视频通才的通用视频智能体) [00:59] 🌐 Depth Anything 3: Recovering the Visual Space from Any Views(Depth Anything 3:从任意视角恢复视觉空间) [01:50] 🔍 AlphaResearch: Accelerating New Algorithm Discovery with Language Models(AlphaResearch:用语言模型加速全新算...
2025.11.13 | 原神数据炼成7B通用AI;零训练轨迹秒变视频遥控器 13.11.2025 6:28
本期的 9 篇论文如下: [00:19] 🌍 Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds(Lumine:在3D开放世界中打造通才智能体的开源方案) [00:54] 🎬 Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising(Time-to-Move:无需训练的双时钟去噪运动控制视频生成) [01:31] ⚡ TiDAR: Think in Diffusion, Talk in Autoregression(TiDAR:扩散式思考,自回归式表...
2025.11.12 | 1.5B小模型反超671B大模型;多智能体质检聊天机器人 12.11.2025 6:56
本期的 9 篇论文如下: [00:24] 🧠 Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B(小模型大逻辑:多样性驱动优化唤醒VibeThinker-1.5B的大模型推理力) [00:59] 🤝 Adaptive Multi-Agent Response Refinement in Conversational Systems(对话系统中自适应多智能体响应精炼机制) [01:30] 🧩 Wasm: A Pipeline for Constructing Structured Arabic Interleav...
2025.11.11 | 小窗口勤总结刷新深度研究;先广撒网再啃难题激活代码竞赛 11.11.2025 9:58
本期的 13 篇论文如下: [00:25] 🧩 IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction(IterResearch:基于马尔可夫状态重构的长程智能体再思考) [01:16] 🏆 DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation(DRIVE:面向可验证奖励强化学习的竞赛级代码生成数据精选最佳实践) [02:03] 🔬 The Station: An Open-World...
2025.11.10 | DeepEyesV2小模型边看图边写代码;纯数据让AI长出立体眼 10.11.2025 5:30
本期的 7 篇论文如下: [00:21] 🧠 DeepEyesV2: Toward Agentic Multimodal Model(DeepEyesV2:迈向智能体多模态模型) [01:13] 🧭 Visual Spatial Tuning(视觉空间调优) [01:54] 🦹 Too Good to be Bad: On the Failure of LLMs to Role-Play Villains(过于完美以致无法邪恶:大语言模型反派角色扮演的失败) [02:27] 🧠 Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings...
【周末特辑】11月第2周最火AI论文 | 视频生成即推理;SVG草图变代码 08.11.2025 12:07
本期的 5 篇论文如下: [00:31] TOP1(🔥137) | 🎬 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm(用视频思考:视频生成作为统一多模态推理新范式) [02:43] TOP2(🔥95) | 🖼 VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation(VCode:以SVG为符号视觉表征的多模态代码评测基准) [05:12] TOP3(🔥90) | 🚀 Diffusion Language Models are Super Data Lear...
2025.11.07 | 视频推理新范式;图像互动促思维 07.11.2025 8:29
本期的 12 篇论文如下: [00:21] 🎬 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm(用视频思考:视频生成作为统一多模态推理新范式) [00:58] 🧠 V-Thinker: Interactive Thinking with Images(V-Thinker:与图像互动的思维推理) [01:39] 🧠 Scaling Agent Learning via Experience Synthesis(基于经验合成的智能体规模化强化学习) [02:23] 🧠 Cambrian-S: Towards Spatial Supersens...
2025.11.06 | 扩散模型省数据;音视频对口型 06.11.2025 7:01
本期的 9 篇论文如下: [00:17] 🚀 Diffusion Language Models are Super Data Learners(扩散语言模型是超级数据学习者) [01:06] 🎬 UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions(统一音视频生成的不对称跨模态交互方法) [01:42] 🧩 LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation(LEGO-Eval:面向具身3D环境合成...
2025.11.05 | 向量草图测代码;先画后想补视觉 05.11.2025 11:31
本期的 15 篇论文如下: [00:21] 🖼 VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation(VCode:以SVG为符号视觉表征的多模态代码评测基准) [01:12] 🧠 When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought(当可视化成为推理第一步:MIRA视觉思维链基准测试) [01:48] ⚖ When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Pr...
2025.11.04 | 超稀疏MoE激活万亿参数;视觉模型看图胜GNN 04.11.2025 11:06
本期的 15 篇论文如下: [00:23] 🧠 Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation(全激活赋能:将通用推理模型扩展到万亿参数的开放语言基座) [01:03] 👁 The Underappreciated Power of Vision Models for Graph Structural Understanding(被低估的视觉模型在图结构理解中的强大潜能) [01:38] 💡 UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausib...
2025.11.03 | OS-Sentinel实时守护手机操作安全;ThinkMorph让小模型边想边画 03.11.2025 11:02
本期的 15 篇论文如下: [00:21] 🛡 OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows(OS-Sentinel:在真实工作流中通过混合验证提升移动GUI代理安全性) [01:13] 🧠 ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning(ThinkMorph:多模态交错思维链中的涌现特性) [01:49] ⚔ INT v.s. FP: A Comprehensive Study of Fine-Grained Lo...
【月末特辑】10月最火AI论文 | 幼龙BDH稀疏可解释;迷你递归7兆碾压大模型 02.11.2025 22:46
本期的 10 篇论文如下: [00:30] TOP1(🔥522) | 🐣 The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain(幼龙破壳: Transformer 与大脑模型之间缺失的环节) [02:31] TOP2(🔥462) | 🧠 Less is More: Recursive Reasoning with Tiny Networks(小而精:用微型网络递归推理) [04:48] TOP3(🔥255) | 🌱 Agent Learning via Early Experience(基于早期经验的主体学习) [07:04] TOP4(🔥182)...
【周末特辑】11月第1周最火AI论文 | 循环模型省参强推理;Concerto 2D-3D自监督涨点 01.11.2025 11:53
本期的 5 篇论文如下: [00:35] TOP1(🔥174) | 🔄 Scaling Latent Reasoning via Looped Language Models(通过循环语言模型扩展潜在推理能力) [02:30] TOP2(🔥166) | 🎼 Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations(Concerto:2D-3D联合自监督学习涌现空间表征) [05:17] TOP3(🔥115) | 🧩 ReCode: Unify Plan and Action for Universal Granularity Control(ReCode:用递归代码统一规划...
2025.10.31 | Emu3.5统一预测时空;扩散提示驱动机器人 31.10.2025 10:09
本期的 15 篇论文如下: [00:26] 🌍 Emu3.5: Native Multimodal Models are World Learners(Emu3.5:原生多模态世界模型让AI看懂并预测未来) [01:04] 🤖 Exploring Conditions for Diffusion models in Robotic Control(探索扩散模型在机器人控制中的条件化策略) [01:42] 🎬 Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark(视频模型已准备好做零样本推理了吗?基于MME-CoF基...
2025.10.30 | 看图写码7B逆袭;视频思维RL破局 30.10.2025 11:29
本期的 15 篇论文如下: [00:22] 👁 JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence(JanusCoder:面向代码智能的基础视觉-编程接口) [01:00] 🧠 Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning(Video-Thinker:用强化学习点燃“视频思维”) [01:55] 🔄 ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization(ReForm:...
2025.10.29 | 通义深度研究报告;小模型折记忆胜671B巨模型 29.10.2025 8:14
本期的 10 篇论文如下: [00:23] 🔍 Tongyi DeepResearch Technical Report(通义深度研究报告:面向长程深度信息检索任务的智能体大模型) [01:00] 🧠 AgentFold: Long-Horizon Web Agents with Proactive Context Management(AgentFold:面向长程任务的主动式上下文管理智能体) [01:36] 🤖 RoboOmni: Proactive Robot Manipulation in Omni-modal Context(RoboOmni:全模态上下文下的主动机器人操作) [02:33] 🎮 Game-TARS:...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.