duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2026.03.18 | 验证求精代理破局;工业代码模型一次过 18.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:29] 🤖 MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification(MiroThinker-1.7与H1:通过验证迈向重型研究智能体) [01:10] 🏭 InCoder-32B: Code Foundation Model for Industrial Scenarios(InCoder-32B:面向工业场...

2026.03.17 | AI学会科学审美;开源数据打破搜索垄断 18.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:29] 🧠 AI Can Learn Scientific Taste(AI可以学习科学品味) [01:13] 🔍 OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data(OpenSeeker:通过完全开源训练数据实现前沿搜索代理的民主化) [02:06] 🏢...

2026.03.16 | LMEB填补长记忆评测盲区;Cheers解耦语义与细节实现多模态统一 16.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:28] 🧠 LMEB: Long-horizon Memory Embedding Benchmark(LMEB:长时程记忆嵌入基准) [01:12] 🔄 Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation(Cheers:通过解...

【周末特辑】3月第3周最火AI论文 | 几何强化3D编辑;LLM视觉编码轻量飞跃 15.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 5 篇论文如下: [00:50] TOP1(🔥136) | 🎨 Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing(几何引导的强化学习用于多视角一致的3D场景编辑) [02:57] TOP2(🔥104) | 🐧 Penguin-VL: Exploring the Efficiency Limits of VLM w...

2026.03.13 | 流式空间记忆2B小模型逆袭;AI“蛮力”翻页不敌人类策略 14.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:32] 🧠 Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training(Spatial-TTT:基于测试时训练的流式视觉空间智能) [01:17] 🤔 Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Do...

2026.03.12 | 边聊边训智能体;GPU秒解亿级K均 12.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:29] 🤖 OpenClaw-RL: Train Any Agent Simply by Talking(OpenClaw-RL:通过对话训练任意智能体) [01:17] ⚡ Flash-KMeans: Fast and Memory-Efficient Exact K-Means(Flash-KMeans:快速且内存高效的精确K-Means算法) [02:01] 👁 MA-EgoQA:...

2026.03.11 | 几何强化3D编辑;掩码扩散多模态 12.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:32] 🎨 Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing(几何引导的强化学习用于多视角一致的3D场景编辑) [01:11] 🔄 Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Dis...

2026.03.10 | 长故事一致性漏洞扫描;零人工3D空间智能标注 10.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:32] 📖 Lost in Stories: Consistency Bugs in Long Story Generation by LLMs(迷失于故事:大语言模型生成长篇故事中的一致性错误) [01:16] 🧠 Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence(Holi-Spatial:...

2026.03.09 | LLM做视觉编码器;BandPO剪得更聪明 09.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:34] 🐧 Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders(Penguin-VL:探索基于LLM视觉编码器的VLM效率极限) [01:16] 🚀 BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bound...

【周末特辑】3月第2周最火AI论文 | 统一编码器跨域点云;异构模型协作省样本 08.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 5 篇论文如下: [00:40] TOP1(🔥143) | 🧩 Utonia: Toward One Encoder for All Point Clouds(Utonia:迈向适用于所有点云的统一编码器) [03:19] TOP2(🔥141) | 🤝 Heterogeneous Agent Collaborative Reinforcement Learning(异构智能体协作强化学习) [05:46] T...

【月末特辑】2月最火AI论文 | VBVR百万视频炼视觉推理;OPUS同频优化器省算力 07.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 10 篇论文如下: [00:44] TOP1(🔥508) | 🧠 A Very Big Video Reasoning Suite(一个超大规模视频推理套件) [03:11] TOP2(🔥343) | 🚀 OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration(OPUS:迈...

2026.03.06 | MOOSE-Star打破科学发现训练壁垒;DARE让LLM秒变严谨统计助手 06.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:32] 🚀 MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier(MOOSE-Star:通过打破复杂性壁垒解锁科学发现的可处理训练) [01:50] 📊 DARE: Aligning LLM Agents with the R Statistical E...

2026.03.05 | Helios无限续写长视频;异构模型协同减半刷题 05.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:33] 🎬 Helios: Real Real-Time Long Video Generation Model(Helios:实时长视频生成模型) [01:12] 🤝 Heterogeneous Agent Collaborative Reinforcement Learning(异构智能体协作强化学习) [01:56] 🧠 T2S-Bench & Structure-of-Thought:...

2026.03.04 | 统一模型“对齐税”拖累理解;通用点云编码器一锅端多场景 04.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:32] 🔍 UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?(UniG2U-Bench:统一模型是否推动了多模态理解的发展?) [01:40] 🧩 Utonia: Toward One Encoder for All Point Clouds(Utonia:迈向适用于所有点云的统一编码器)...

2026.03.03 | 自适应扩展省算力;令牌秒变动效 03.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:30] ⚡ From Scale to Speed: Adaptive Test-Time Scaling for Image Editing(从规模到速度:图像编辑的自适应测试时扩展) [01:16] 🎨 OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens(OmniLottie:通过参数化Lot...

2026.03.02 | dLLM统一扩散框架;SpatialScore让AI读懂空间 02.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:29] 🛠 dLLM: Simple Diffusion Language Modeling(dLLM:简单的扩散语言建模) [01:15] 🧠 Enhancing Spatial Understanding in Image Generation via Reward Modeling(通过奖励建模增强图像生成中的空间理解) [02:11] 🌍 Recovered in Trans...

【周末特辑】3月第1周最火AI论文 | VBVR 百万级视频基准刷新推理极限;SAGE 自信早停让模型省话又精准 01.03.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 5 篇论文如下: [00:39] TOP1(🔥491) | 🧠 A Very Big Video Reasoning Suite(一个超大规模视频推理套件) [02:33] TOP2(🔥246) | 💭 Does Your Reasoning Model Implicitly Know When to Stop Thinking?(你的推理模型是否隐含地知道何时停止思考?) [04:48] TOP3...

2026.02.27 | 诊断补课反超72B;三一致性考趴世界模型 27.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:31] 🔍 From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models(从盲点到增益:诊断驱动的迭代训练用于大型多模态模型) [01:16] 🌍 The Trinity of Consistency as a Defining Principle for General...

2026.02.26 | 分子图生成首破99%化学有效性;DreamID-Omni把多人脸音色混剪错配率砍到8% 26.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:31] ⚗ MolHIT: Advancing Molecular-Graph Generation with Hierarchical Discrete Diffusion Models(MolHIT:基于分层离散扩散模型推进分子图生成) [01:08] 🎭 DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video...

2026.02.25 | 数据工程赋能小模型;轻量重排刷新长文本SOTA 25.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:29] 🖥 On Data Engineering for Scaling LLM Terminal Capabilities(论扩展大型语言模型终端能力的数据工程) [01:20] 🧠 Query-focused and Memory-aware Reranker for Long Context Processing(面向长文本处理的查询聚焦与记忆感知重排序器...

2026.02.24 | VBVR百万视频补推理教材;VLANeXt十二配方炼成VLA 24.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 14 篇论文如下: [00:31] 🧠 A Very Big Video Reasoning Suite(一个超大规模视频推理套件) [01:16] 🧪 VLANeXt: Recipes for Building Strong VLA Models(VLANeXt:构建强大视觉-语言-动作模型的实践指南) [02:06] 🧭 ManCAR: Manifold-Constrained Latent Reas...

2026.02.23 | VESPO防抖离线RL;推理模型学会“点到为止” 23.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 10 篇论文如下: [00:40] ⚖ VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training(VESPO:用于稳定离策略LLM训练的变分序列级软策略优化) [01:45] 💭 Does Your Reasoning Model Implicitly Know When to Stop Thinkin...

【周末特辑】2月第4周最火AI论文 | 少即是够;FAC靶向补特征;噪声基准SQuTR 22.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 5 篇论文如下: [00:45] TOP1(🔥219) | 🧠 Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs(少即是够:在大型语言模型特征空间中合成多样化数据) [03:23] TOP2(🔥140) | 🔊 SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieva...

2026.02.20 | 砍95%注意力画质反升;边压缩边生成FID 1.4 20.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:31] ⚡ SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning(SpargeAttention2:通过混合Top-k+Top-p掩码与蒸馏微调实现可训练的稀疏注意力) [01:27] 🧠 Unified Latents (UL): How t...

2026.02.19 | 可学习路由+量化加速视频扩散;残差追踪让人形90%抓取 19.02.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 14 篇论文如下: [00:30] ⚡ SLA2: Sparse-Linear Attention with Learnable Routing and QAT(SLA2:具有可学习路由和量化感知训练的稀疏线性注意力) [01:16] 🤖 Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation(面向开放...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.