duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2026.05.13 | 原生统一看画;边缘隐私记管 13.05.2026 13:06
【目录】 本期的 15 篇论文如下: [00:23] 🧠 SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture(SenseNova-U1: 基于NEO-unify架构统一多模态理解与生成) [01:10] 🔒 MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents(MemPrivacy:面向边缘-云智能体的隐私保护个性化记忆管理) [01:59] 🧠 $δ$-mem: Efficient Online Memory for Large Langu...
2026.05.12 | 数学家闭门出题考倒大模型;生图模型千字提示精准成画 12.05.2026 13:11
【目录】 本期的 15 篇论文如下: [00:25] 🧮 Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs(Soohak:由数学家策划的基准测试,用于评估大语言模型的研究级数学能力) [01:30] 🎨 Qwen-Image-2.0 Technical Report(Qwen-Image-2.0技术报告) [02:23] 🎥 CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models(CollabVR:基于视觉...
2026.05.11 | 音乐驱舞拆分专家;流匹配蒸馏全科状元 11.05.2026 13:03
【目录】 本期的 15 篇论文如下: [00:29] 💃 MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation(MACE-Dance:音乐驱动舞蹈视频生成的运动与外观级联专家模型) [01:07] 🎯 Flow-OPD: On-Policy Distillation for Flow Matching Models(Flow-OPD:面向流匹配模型的在线策略蒸馏) [01:58] 🎯 Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response...
【周末特辑】5月第2周最火AI论文 | MolmoAct2开源机器人大脑;长文狼人杀自练暗规则 10.05.2026 12:16
【目录】 本期的 5 篇论文如下: [00:33] TOP1(🔥266) | 🤖 MolmoAct2: Action Reasoning Models for Real-world Deployment(MolmoAct2:面向实际部署的動作推理模型) [03:10] TOP2(🔥145) | 🧠 From Context to Skills: Can Language Models Learn from Context Skillfully?(从上下文到技能:语言模型能否从上下文中巧妙学习?) [05:03] TOP3(🔥117) | 🎥 Stream-R1: Reliability-Perplexity Aware Reward Distillation for S...
2026.05.08 | 全局速写助长文;技能库让智能体进化 08.05.2026 12:58
【目录】 本期的 15 篇论文如下: [00:23] 🧠 MiA-Signature: Approximating Global Activation for Long-Context Understanding(MiA-签名:面向长上下文理解的全局激活近似方法) [01:32] 🧬 Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning(Skill1:通过强化学习实现技能增强型智能体的统一进化) [02:14] 🎯 MARBLE: Multi-Aspect Reward Balance for Diffusion RL(MARBLE:面向扩散强化学...
2026.05.07 | 奖励蒸馏让像素会“挑重点”;测试时扩展逐块稳长视频 07.05.2026 13:46
【目录】 本期的 15 篇论文如下: [00:24] 🎥 Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation(Stream-R1:面向流式视频生成的可靠性-困惑度感知奖励蒸馏) [01:27] 🎥 Stream-T1: Test-Time Scaling for Streaming Video Generation(Stream-T1:面向流式视频生成的测试时扩展) [02:06] 🔍 OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents(OpenSearch-VL...
2026.05.06 | ARIS自怼写论文;PRISM三段洗数据再RL 07.05.2026 12:29
【目录】 本期的 15 篇论文如下: [00:25] 🤖 ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration(自主研究:通过对抗性多智能体协作实现自动化科研) [00:59] 🎯 Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL(超越SFT到RL:通过黑盒在线策略蒸馏实现多模态强化学习的预对齐) [01:54] 🔍 OpenSeeker-v2: Pushing the Limits of Search Agents with Informative...
2026.05.05 | 开源MolmoAct2实战87%成功率;GPT上下文提炼技能再升级 05.05.2026 11:03
【目录】 本期的 14 篇论文如下: [00:21] 🤖 MolmoAct2: Action Reasoning Models for Real-world Deployment(MolmoAct2:面向实际部署的動作推理模型) [01:02] 🧠 From Context to Skills: Can Language Models Learn from Context Skillfully?(从上下文到技能:语言模型能否从上下文中巧妙学习?) [01:44] 🔁 Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling(重复...
【月末特辑】4月最火AI论文 | GrandCode登顶Codeforces;高频Prompt提效大模型 05.05.2026 21:07
【目录】 本期的 10 篇论文如下: [00:47] TOP1(🔥626) | 🏆 GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning(GrandCode:通过智能体强化学习在竞技编程中达到宗师级水平) [02:41] TOP2(🔥501) | 📈 Adam's Law: Textual Frequency Law on Large Language Models(亚当定律:大语言模型上的文本频率定律) [04:54] TOP3(🔥364) | 🔄 DataFlex: A Unified Framework for...
2026.05.04 | 统一扩散框架十五合一;多智能体搜索碾压单兵 04.05.2026 13:32
【目录】 本期的 15 篇论文如下: [00:23] 🎥 UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors(UniVidX:一种基于扩散先验的统一多模态框架用于多功能视频生成) [01:20] 🕸 Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction(Web2BigTable:一种用于互联网规模信息搜索与提取的双层多智能体大语言模型系统) [02:11] 🌍 Ma...
【周末特辑】5月第1周最火AI论文 | 潜空间套娃提分快;世界模型分级演化 03.05.2026 10:47
【目录】 本期的 5 篇论文如下: [00:35] TOP1(🔥241) | 🔄 Recursive Multi-Agent Systems(递归多智能体系统) [02:34] TOP2(🔥219) | 🌍 Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond(智能体世界建模:基础、能力、法则及其超越) [04:45] TOP3(🔥188) | 🧠 Heterogeneous Scientific Foundation Model Collaboration(异构科学基础模型协作) [06:31] TOP4(🔥116) | 🏢 From Skills to Talent: Orga...
2026.05.01 | Eywa让LLM牵手领域模型提效30%;视觉生成五级跃迁仍卡第三关 01.05.2026 13:42
【目录】 本期的 15 篇论文如下: [00:25] 🧠 Heterogeneous Scientific Foundation Model Collaboration(异构科学基础模型协作) [01:24] 🌍 Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling(新时代的视觉生成:从原子映射到智能体世界建模的演进) [02:04] 🧬 Co-Evolving Policy Distillation(共同演化策略蒸馏) [02:47] 🤖 ExoActor: Exocentric Video Generation as Gene...
2026.04.30 | GLM-5V一锅端训多模态;潜在蒸馏采样省样本 30.04.2026 9:43
【目录】 本期的 11 篇论文如下: [00:22] 🤖 GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents(GLM-5V-Turbo:迈向多模态智能体的原生基础模型) [01:26] 🔬 Large Language Models Explore by Latent Distilling(大型语言模型通过潜在蒸馏进行探索) [02:16] 🌊 Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models(扭转潮流:面向扩散大语言模型的跨架构蒸馏) [...
2026.04.29 | 递归多智能体套娃提速;数据编程Git式自改进 29.04.2026 12:37
【目录】 本期的 15 篇论文如下: [00:25] 🔄 Recursive Multi-Agent Systems(递归多智能体系统) [01:01] 🔧 Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora(数据编程:面向自改进大语言模型从原始语料库进行测试驱动数据工程) [01:55] 📊 DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios(DV-World:在真实世界场景中评估数据可视化智能体的基准...
2026.04.28 | 强化学习逼出几何一致视频;AI公司乐高式组队降本提效 28.04.2026 15:12
【目录】 本期的 15 篇论文如下: [00:24] 🌍 World-R1: Reinforcing 3D Constraints for Text-to-Video Generation(世界-R1:通过强化学习为文本到视频生成注入3D约束) [01:29] 🏢 From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company(从技能到人才:将异构智能体组织为现实世界公司) [02:26] 🧠 ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D...
2026.04.27 | 坐标系统摄世界模型;扩散重建提速临床CT 27.04.2026 9:21
【目录】 本期的 11 篇论文如下: [00:31] 🌍 Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond(智能体世界建模:基础、能力、法则及其超越) [01:24] 🩻 DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction(DiffNR:基于扩散增强的神经表示优化用于稀疏视角三维断层重建) [02:10] 🛡 LLM Safety From Within: Detecting Harmful Content with...
【周末特辑】4月第4周最火AI论文 | Tstars-Tryon登顶虚拟试穿;LLaDA2.0-Uni统一多模态生成 26.04.2026 11:11
【目录】 本期的 5 篇论文如下: 00:31 TOP1(🔥244) | 👗 Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items(Tstars-Tryon 1.0:面向多样化时尚商品的鲁棒且逼真的虚拟试穿系统) 02:42 TOP2(🔥229) | 🔮 LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model(LLaDA2.0-Uni:基于扩散大语言模型统一多模态理解与生成) 05:07 TOP3(🔥154) | 🤖...
2026.04.24 | LLaTiSA四级闯关教模型读时序;WorldMark统一基准测视频世界模型 24.04.2026 13:26
【目录】 本期的 15 篇论文如下: 00:23 📈 LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics(LLaTiSA:从视觉感知到语义的难度分层时间序列推理) 01:11 🎮 WorldMark: A Unified Benchmark Suite for Interactive Video World Models(WorldMark:交互式视频世界模型的统一基准套件) 01:54 🤖 UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Lear...
2026.04.23 | LLaDA2.0统一多模态;未来经验外挂RL 23.04.2026 12:19
【目录】 本期的 15 篇论文如下: 00:28 🔮 LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model(LLaDA2.0-Uni:基于扩散大语言模型统一多模态理解与生成) 01:17 🔮 Near-Future Policy Optimization(近未来策略优化) 02:07 🤖 DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data(DR-Venus:仅用1万条开放数据迈向前沿边缘规模深度研究...
2026.04.22 | 虚拟试衣3.9秒高清生成;协同生成HOI视频物理一致 22.04.2026 12:48
【目录】 本期的 15 篇论文如下: 00:23 👗 Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items(Tstars-Tryon 1.0:面向多样化时尚商品的鲁棒且逼真的虚拟试穿系统) 01:05 🎬 CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation(CoInteract:通过空间结构化协同生成实现物理一致的人-物交互视频合成) 01:58 🤖 AgentSPEX: A...
2026.04.21 | 一步听懂句子出图;单步潜码搞定驾驶推理 21.04.2026 13:37
【目录】 本期的 15 篇论文如下: 00:24 🚀 Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation(从类别标签到文本:通过判别性文本表征扩展一步图像生成) 01:08 🚗 OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation(OneVL:基于视觉语言解释的单步潜在推理与规划) 01:54 🤖 Agent-World: Scaling Real-World Environment Synthesis for...
2026.04.20 | DPM零训画质糖;两位翻转毁模型 20.04.2026 13:16
【目录】 本期的 15 篇论文如下: 00:20 🔍 Elucidating the SNR-t Bias of Diffusion Probabilistic Models(阐明扩散概率模型的信噪比-时间步偏差) 01:00 💥 Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips(无需数据或优化的最大脑损伤:通过符号位翻转破坏神经网络) 01:45 🧠 PersonaVLM: Long-Term Personalized Multimodal LLMs(PersonaVLM:面向长期个性化的多模态...
【周末特辑】4月第3周最火AI论文 | 单图识3D升维;记忆防崩溃 18.04.2026 13:18
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗www.xiaoyuzhoufm.com 【目录】 本期的 5 篇论文如下: 00:41 TOP1(🔥237) | 🔍 WildDet3D: Scaling Promptable 3D Detection in the Wild(WildDet3D:可扩展的野外可提示三维检测) 02:54 TOP2(🔥135) | 🧠 The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping(过去并未过去:基于记忆增强的动态奖励塑形) 05:12 TOP3(🔥134) | 🤖 Cla...
2026.04.17 | HY-World2.0统一生成与重建;DR³-Eval建可复现研究基准 17.04.2026 13:04
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗www.xiaoyuzhoufm.com 【目录】 本期的 15 篇论文如下: 00:31 🌍 HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds(HY-World 2.0:用于重建、生成和模拟3D世界的多模态世界模型) 01:24 🔬 DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation(DR³-Eval:迈向现实且...
2026.04.16 | Seedance 2.0一统多模态生成;RationalRewards让奖励模型讲理 16.04.2026 13:24
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗www.xiaoyuzhoufm.com 【目录】 本期的 15 篇论文如下: 00:31 🎬 Seedance 2.0: Advancing Video Generation for World Complexity(Seedance 2.0:面向世界复杂性的视频生成技术进展) 01:21 🧠 RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time(RationalRewards:推理奖励在训练和测试时均能提升视觉生...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.