duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2026.01.21 | AI修Bug统一打分;MLLM未来预测仍易盲猜 21.01.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:31] 🤖 Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey(基于大语言模型的软件工程问题解决:进展、前沿与全面综述) [01:15] 🔮 FutureOmni: Evaluating Future Forecasting from Omn...

2026.01.20 | 沙盒测通才是真后端;分叉合并少字多想 20.01.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 8 篇论文如下: [00:30] ⚙ ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development(ABC-Bench:面向真实世界开发的智能体后端编码基准测试) [01:15] 🧠 Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge(多路思考:基于词元级...

2026.01.19 | GRPO回报纠偏助啃难题;毒苹果AI未用已扰市 19.01.2026

【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗 https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:33] ⚖ Your Group-Relative Advantage Is Biased(你的组相对优势存在偏差) [01:20] 🍎 The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents(毒苹果效应:通过AI代理技术扩展对中...

【周末特辑】1月第3周最火AI论文 | VideoDR测模型搜证漂移;BabyVision曝视觉短板 17.01.2026

本期的 5 篇论文如下: [00:29] TOP1(🔥201) | 🔍 Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning(观察、推理与搜索:面向智能体视频推理的开放网络视频深度研究基准) [02:45] TOP2(🔥179) | 👶 BabyVision: Visual Reasoning Beyond Language(BabyVision:超越语言的视觉推理) [05:00] TOP3(🔥158) | 🗺 Thinking with Map: Reinforced Parallel Map-Augmente...

2026.01.16 | 10B模型逆袭千亿巨头;AI一眼读出城市功能 16.01.2026

本期的 15 篇论文如下: [00:20] 🚀 STEP3-VL-10B Technical Report(STEP3-VL-10B 技术报告) [01:01] 🏙 Urban Socio-Semantic Segmentation with Vision-Language Reasoning(基于视觉语言推理的城市社会语义分割) [01:42] 💡 Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs(奖励罕见:面向LLM创造性问题解决的独特性感知强化学习) [02:33] 🤖 Collaborative Multi-Agent Test-Time Reinforc...

2026.01.15 | 算法自进化夺冠;LLM远瞻省token 15.01.2026

本期的 15 篇论文如下: [00:20] 🧬 Controlled Self-Evolution for Algorithmic Code Optimization(用于算法代码优化的受控自进化方法) [00:52] 🧠 MAXS: Meta-Adaptive Exploration with LLM Agents(MAXS:基于大语言模型智能体的元自适应探索) [01:27] 🧠 $A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation(A³-Bench:通过锚点与吸引子激活基准测试记忆驱动的科学推理)...

2026.01.14 | 合成数据喂出低资源学霸;AI自演多轮对话更靠谱 14.01.2026

本期的 15 篇论文如下: [00:20] 🌍 Solar Open Technical Report(Solar Open 技术报告) [00:54] 🤖 User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale(面向用户的大规模多轮对话生成与工具使用) [01:39] 🧠 MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences(MemGovern:通过从受治理的人类经验中学习来增强代码代理) [02:11] 🖱 ShowUI-$π$: Flow-based Generative...

2026.01.13 | VideoDR让模型边搜边推理;BabyVision揭视觉短板 13.01.2026

本期的 15 篇论文如下: [00:20] 🔍 Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning(观察、推理与搜索:面向智能体视频推理的开放网络视频深度研究基准) [01:01] 👶 BabyVision: Visual Reasoning Beyond Language(BabyVision:超越语言的视觉推理) [01:45] 🚀 PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning(PaCoRe:通...

2026.01.12 | 地图AI强化寻位;多模态Lean形式化 12.01.2026

本期的 15 篇论文如下: [00:20] 🗺 Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization(借助地图思考:用于地理定位的强化并行地图增强智能体) [01:03] 🧠 MMFormalizer: Multimodal Autoformalization in the Wild(MMFormalizer:面向真实世界的多模态自动形式化方法) [01:38] 🧬 The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning(思维分子结构...

【周末特辑】1月第2周最火AI论文 | GDPO分灶吃饭稳优化;NeoVerse单目视频建4D 11.01.2026

本期的 5 篇论文如下: [00:39] TOP1(🔥126) | 📈 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization(GDPO:面向多奖励强化学习优化的组奖励解耦归一化策略优化) [02:31] TOP2(🔥108) | 🌍 NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos(NeoVerse:利用野外单目视频增强4D世界模型) [04:40] TOP3(🔥107) | 🤖 Youtu-Agent: Scaling Agent Productiv...

2026.01.09 | GDPO解耦奖励优化多任务;可学习乘数解锁矩阵尺度 09.01.2026

本期的 15 篇论文如下: [00:21] 📈 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization(GDPO:面向多奖励强化学习优化的组奖励解耦归一化策略优化) [01:05] ⚖ Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers(可学习的乘数:释放语言模型矩阵层的尺度) [01:33] 🌙 RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low...

2026.01.08 | 熵加权微调保旧学;演化技能网络不断进阶 08.01.2026

本期的 15 篇论文如下: [00:21] ⚖ Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting(熵自适应微调:解决置信冲突以缓解遗忘) [01:15] 🧠 Evolving Programmatic Skill Networks(演化式程序化技能网络) [01:51] 🧠 Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning(Atlas:面向多领域复杂推理的异构模型与工具编排框架) [02:31] 📊 Benchmark^...

2026.01.07 | 无限深度任意采样;端到端语音转录分离 07.01.2026

本期的 15 篇论文如下: [00:25] 🔍 InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields(InfiniDepth:基于神经隐式场的任意分辨率与细粒度深度估计) [01:07] 🎙 MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization(MOSS转录与说话人分离:带说话人归属和时间戳的准确转录) [01:46] 🔬 SciEvalKit: An Open-source Evaluation Toolkit for Scientific...

2026.01.06 | K-EXAONE MoE;NextFlow统一序列建模多模态 06.01.2026

本期的 15 篇论文如下: [00:21] 🧠 K-EXAONE Technical Report(K-EXAONE技术报告) [00:56] 🚀 NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation(NextFlow:统一序列建模激活多模态理解与生成) [01:36] 🎭 DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer(DreamID-V:通过扩散Transformer弥合图像到视频的鸿沟以实现高保真...

2026.01.05 | Agent流水线提速;4D建模平民化 05.01.2026

本期的 12 篇论文如下: [00:22] 🤖 Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization(Youtu-Agent:通过自动化生成与混合策略优化扩展智能体生产力) [00:52] 🌍 NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos(NeoVerse:利用野外单目视频增强4D世界模型) [01:27] 🤖 Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural...

【周末特辑】1月第1周最火AI论文 | mHC 稳梯度;思维景观 RAG 读长文 03.01.2026

本期的 5 篇论文如下: [00:33] TOP1(🔥132) | 🧠 mHC: Manifold-Constrained Hyper-Connections(mHC:流形约束的超连接) [02:32] TOP2(🔥100) | 🧠 Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding(面向提升长文本理解的思维景观感知检索增强生成) [04:45] TOP3(🔥94) | 🎬 InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion...

2026.01.02 | 语义密度压缩;扩散边画边想 02.01.2026

本期的 3 篇论文如下: [00:19] 🧠 Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space(动态大型概念模型:自适应语义空间中的潜在推理) [00:56] 🧠 DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models(DiffThinker:基于扩散模型的生成式多模态推理) [01:27] 🔄 On the Role of Discreteness in Diffusion LLMs(论离散性在扩散语言模型中的作用) 【关注我们】 您还...

【月末特辑】12月最火AI论文 | 代码智能全链路落地;开源模型推理代理双突破 01.01.2026

本期的 10 篇论文如下: [00:29] TOP1(🔥279) | 🧠 From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence(从代码基础模型到智能体与应用:代码智能实用指南) [02:22] TOP2(🔥242) | 🚀 DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models(DeepSeek-V3.2:推动开放大型语言模型前沿) [04:45] TOP3(🔥217) | 🚀 Z-Image: An Efficient Image Generation Foundatio...

2026.01.01 | 小模型也能原生外挂;30B-MoE智体逼近大模型 01.01.2026

本期的 15 篇论文如下: [00:22] 🚀 Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models(Youtu-LLM:解锁轻量级大语言模型的原生智能体潜力) [01:00] 🤖 Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem(任其流动:摇滚乐上的智能体构建,在开放智能体学习生态系统中建立ROME模型) [01:52] 🧠 mHC: Manifold-Con...

2025.12.31 | 粗模精雕UltraShape;涂鸦编辑DreamOmni3 31.12.2025

本期的 6 篇论文如下: [00:24] 🧊 UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement(UltraShape 1.0:通过可扩展几何精化的高保真3D形状生成) [01:00] 🎨 DreamOmni3: Scribble-based Editing and Generation(DreamOmni3:基于涂鸦的编辑与生成) [01:34] 🧠 End-to-End Test-Time Training for Long Context(面向长上下文的端到端测试时训练) [02:18] 🔬 Evaluating Parameter Effici...

2025.12.30 | ERC耦合路由与专家;LiveTalk实时视频对话 30.12.2025

本期的 15 篇论文如下: [00:24] 🔗 Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss(通过辅助损失耦合专家混合模型中的专家与路由器) [01:07] 🎬 LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation(LiveTalk:通过改进的策略内蒸馏实现实时多模态交互式视频扩散) [01:55] 🌍 Yume-1.5: A Text-Controlled Interactive World Generation Model(Yu...

2025.12.29 | 鸟瞰式检索提效小模型;4D扩散一键插入逼真物体 29.12.2025

本期的 13 篇论文如下: [00:27] 🧠 Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding(面向提升长文本理解的思维景观感知检索增强生成) [01:07] 🎬 InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion(InsertAnywhere:连接4D场景几何与扩散模型以实现逼真的视频对象插入) [01:46] 🤖 MAI-UI Technical Report: Real-World Cent...

【周末特辑】12月第5周最火AI论文 | DataFlow炼数工厂上线;AI科学家跑不完闭环 27.12.2025

本期的 5 篇论文如下: [00:42] TOP1(🔥188) | ⚙ DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI(DataFlow:面向数据为中心AI时代的统一数据准备与工作流自动化LLM驱动框架) [02:34] TOP2(🔥105) | 🔬 Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows(通过科学家对齐的工作流程探究大语言模型的科学通用智能) [0...

2025.12.26 | 暗号token涨点视觉推理;3D便签本让视频长脑子 26.12.2025

本期的 6 篇论文如下: [00:19] 🧠 Latent Implicit Visual Reasoning(潜在隐式视觉推理) [00:56] 🎬 Spatia: Video Generation with Updatable Spatial Memory(Spatia:基于可更新空间记忆的视频生成) [01:36] 🧠 Schoenfeld's Anatomy of Mathematical Reasoning by Language Models(基于舍恩菲尔德理论的语言模型数学推理解剖) [02:11] 🔍 How Much 3D Do Video Foundation Models Encode?(视频基础模型编码了多少3D信息...

2025.12.25 | 四维动态理解刷新VLM;单卡200倍速生成高清视频 25.12.2025

本期的 14 篇论文如下: [00:20] 🧠 Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models(学习在四维空间中推理:视觉语言模型的动态空间理解) [01:11] ⚡ TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times(TurboDiffusion:将视频扩散模型加速100-200倍) [01:52] 🧭 T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation(T2AV-Compass:迈...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.