duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2026.01.21 | AI修Bug统一打分;MLLM未来预测仍易盲猜 21.01.2026 13:06
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:31] 🤖 Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey(基于大语言模型的软件工程问题解决:进展、前沿与全面综述) [01:15] 🔮 FutureOmni: Evaluating Future Forecasting from Omn...
2026.01.20 | 沙盒测通才是真后端;分叉合并少字多想 20.01.2026 7:06
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 8 篇论文如下: [00:30] ⚙ ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development(ABC-Bench:面向真实世界开发的智能体后端编码基准测试) [01:15] 🧠 Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge(多路思考:基于词元级...
2026.01.19 | GRPO回报纠偏助啃难题;毒苹果AI未用已扰市 19.01.2026 14:35
【赞助商】 通勤路上就听AI每周谈。AI每周谈,每周带你回顾上周AI大事 传送门 🔗 https://www.xiaoyuzhoufm.com/podcast/688a34636f5a275f1cba40fd 【目录】 本期的 15 篇论文如下: [00:33] ⚖ Your Group-Relative Advantage Is Biased(你的组相对优势存在偏差) [01:20] 🍎 The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents(毒苹果效应:通过AI代理技术扩展对中...
【周末特辑】1月第3周最火AI论文 | VideoDR测模型搜证漂移;BabyVision曝视觉短板 17.01.2026 11:55
本期的 5 篇论文如下: [00:29] TOP1(🔥201) | 🔍 Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning(观察、推理与搜索:面向智能体视频推理的开放网络视频深度研究基准) [02:45] TOP2(🔥179) | 👶 BabyVision: Visual Reasoning Beyond Language(BabyVision:超越语言的视觉推理) [05:00] TOP3(🔥158) | 🗺 Thinking with Map: Reinforced Parallel Map-Augmente...
2026.01.16 | 10B模型逆袭千亿巨头;AI一眼读出城市功能 16.01.2026 11:15
本期的 15 篇论文如下: [00:20] 🚀 STEP3-VL-10B Technical Report(STEP3-VL-10B 技术报告) [01:01] 🏙 Urban Socio-Semantic Segmentation with Vision-Language Reasoning(基于视觉语言推理的城市社会语义分割) [01:42] 💡 Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs(奖励罕见:面向LLM创造性问题解决的独特性感知强化学习) [02:33] 🤖 Collaborative Multi-Agent Test-Time Reinforc...
2026.01.15 | 算法自进化夺冠;LLM远瞻省token 15.01.2026 11:13
本期的 15 篇论文如下: [00:20] 🧬 Controlled Self-Evolution for Algorithmic Code Optimization(用于算法代码优化的受控自进化方法) [00:52] 🧠 MAXS: Meta-Adaptive Exploration with LLM Agents(MAXS:基于大语言模型智能体的元自适应探索) [01:27] 🧠 $A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation(A³-Bench:通过锚点与吸引子激活基准测试记忆驱动的科学推理)...
2026.01.14 | 合成数据喂出低资源学霸;AI自演多轮对话更靠谱 14.01.2026 10:25
本期的 15 篇论文如下: [00:20] 🌍 Solar Open Technical Report(Solar Open 技术报告) [00:54] 🤖 User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale(面向用户的大规模多轮对话生成与工具使用) [01:39] 🧠 MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences(MemGovern:通过从受治理的人类经验中学习来增强代码代理) [02:11] 🖱 ShowUI-$π$: Flow-based Generative...
2026.01.13 | VideoDR让模型边搜边推理;BabyVision揭视觉短板 13.01.2026 10:51
本期的 15 篇论文如下: [00:20] 🔍 Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning(观察、推理与搜索:面向智能体视频推理的开放网络视频深度研究基准) [01:01] 👶 BabyVision: Visual Reasoning Beyond Language(BabyVision:超越语言的视觉推理) [01:45] 🚀 PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning(PaCoRe:通...
2026.01.12 | 地图AI强化寻位;多模态Lean形式化 12.01.2026 11:07
本期的 15 篇论文如下: [00:20] 🗺 Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization(借助地图思考:用于地理定位的强化并行地图增强智能体) [01:03] 🧠 MMFormalizer: Multimodal Autoformalization in the Wild(MMFormalizer:面向真实世界的多模态自动形式化方法) [01:38] 🧬 The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning(思维分子结构...
【周末特辑】1月第2周最火AI论文 | GDPO分灶吃饭稳优化;NeoVerse单目视频建4D 11.01.2026 12:17
本期的 5 篇论文如下: [00:39] TOP1(🔥126) | 📈 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization(GDPO:面向多奖励强化学习优化的组奖励解耦归一化策略优化) [02:31] TOP2(🔥108) | 🌍 NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos(NeoVerse:利用野外单目视频增强4D世界模型) [04:40] TOP3(🔥107) | 🤖 Youtu-Agent: Scaling Agent Productiv...
2026.01.09 | GDPO解耦奖励优化多任务;可学习乘数解锁矩阵尺度 09.01.2026 10:07
本期的 15 篇论文如下: [00:21] 📈 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization(GDPO:面向多奖励强化学习优化的组奖励解耦归一化策略优化) [01:05] ⚖ Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers(可学习的乘数:释放语言模型矩阵层的尺度) [01:33] 🌙 RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low...
2026.01.08 | 熵加权微调保旧学;演化技能网络不断进阶 08.01.2026 11:18
本期的 15 篇论文如下: [00:21] ⚖ Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting(熵自适应微调:解决置信冲突以缓解遗忘) [01:15] 🧠 Evolving Programmatic Skill Networks(演化式程序化技能网络) [01:51] 🧠 Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning(Atlas:面向多领域复杂推理的异构模型与工具编排框架) [02:31] 📊 Benchmark^...
2026.01.07 | 无限深度任意采样;端到端语音转录分离 07.01.2026 11:36
本期的 15 篇论文如下: [00:25] 🔍 InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields(InfiniDepth:基于神经隐式场的任意分辨率与细粒度深度估计) [01:07] 🎙 MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization(MOSS转录与说话人分离:带说话人归属和时间戳的准确转录) [01:46] 🔬 SciEvalKit: An Open-source Evaluation Toolkit for Scientific...
2026.01.06 | K-EXAONE MoE;NextFlow统一序列建模多模态 06.01.2026 10:42
本期的 15 篇论文如下: [00:21] 🧠 K-EXAONE Technical Report(K-EXAONE技术报告) [00:56] 🚀 NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation(NextFlow:统一序列建模激活多模态理解与生成) [01:36] 🎭 DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer(DreamID-V:通过扩散Transformer弥合图像到视频的鸿沟以实现高保真...
2026.01.05 | Agent流水线提速;4D建模平民化 05.01.2026 7:36
本期的 12 篇论文如下: [00:22] 🤖 Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization(Youtu-Agent:通过自动化生成与混合策略优化扩展智能体生产力) [00:52] 🌍 NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos(NeoVerse:利用野外单目视频增强4D世界模型) [01:27] 🤖 Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural...
【周末特辑】1月第1周最火AI论文 | mHC 稳梯度;思维景观 RAG 读长文 03.01.2026 11:56
本期的 5 篇论文如下: [00:33] TOP1(🔥132) | 🧠 mHC: Manifold-Constrained Hyper-Connections(mHC:流形约束的超连接) [02:32] TOP2(🔥100) | 🧠 Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding(面向提升长文本理解的思维景观感知检索增强生成) [04:45] TOP3(🔥94) | 🎬 InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion...
2026.01.02 | 语义密度压缩;扩散边画边想 02.01.2026 2:33
本期的 3 篇论文如下: [00:19] 🧠 Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space(动态大型概念模型:自适应语义空间中的潜在推理) [00:56] 🧠 DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models(DiffThinker:基于扩散模型的生成式多模态推理) [01:27] 🔄 On the Role of Discreteness in Diffusion LLMs(论离散性在扩散语言模型中的作用) 【关注我们】 您还...
【月末特辑】12月最火AI论文 | 代码智能全链路落地;开源模型推理代理双突破 01.01.2026 23:06
本期的 10 篇论文如下: [00:29] TOP1(🔥279) | 🧠 From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence(从代码基础模型到智能体与应用:代码智能实用指南) [02:22] TOP2(🔥242) | 🚀 DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models(DeepSeek-V3.2:推动开放大型语言模型前沿) [04:45] TOP3(🔥217) | 🚀 Z-Image: An Efficient Image Generation Foundatio...
2026.01.01 | 小模型也能原生外挂;30B-MoE智体逼近大模型 01.01.2026 10:51
本期的 15 篇论文如下: [00:22] 🚀 Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models(Youtu-LLM:解锁轻量级大语言模型的原生智能体潜力) [01:00] 🤖 Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem(任其流动:摇滚乐上的智能体构建,在开放智能体学习生态系统中建立ROME模型) [01:52] 🧠 mHC: Manifold-Con...
2025.12.31 | 粗模精雕UltraShape;涂鸦编辑DreamOmni3 31.12.2025 4:32
本期的 6 篇论文如下: [00:24] 🧊 UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement(UltraShape 1.0:通过可扩展几何精化的高保真3D形状生成) [01:00] 🎨 DreamOmni3: Scribble-based Editing and Generation(DreamOmni3:基于涂鸦的编辑与生成) [01:34] 🧠 End-to-End Test-Time Training for Long Context(面向长上下文的端到端测试时训练) [02:18] 🔬 Evaluating Parameter Effici...
2025.12.30 | ERC耦合路由与专家;LiveTalk实时视频对话 30.12.2025 10:39
本期的 15 篇论文如下: [00:24] 🔗 Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss(通过辅助损失耦合专家混合模型中的专家与路由器) [01:07] 🎬 LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation(LiveTalk:通过改进的策略内蒸馏实现实时多模态交互式视频扩散) [01:55] 🌍 Yume-1.5: A Text-Controlled Interactive World Generation Model(Yu...
2025.12.29 | 鸟瞰式检索提效小模型;4D扩散一键插入逼真物体 29.12.2025 9:43
本期的 13 篇论文如下: [00:27] 🧠 Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding(面向提升长文本理解的思维景观感知检索增强生成) [01:07] 🎬 InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion(InsertAnywhere:连接4D场景几何与扩散模型以实现逼真的视频对象插入) [01:46] 🤖 MAI-UI Technical Report: Real-World Cent...
【周末特辑】12月第5周最火AI论文 | DataFlow炼数工厂上线;AI科学家跑不完闭环 27.12.2025 12:36
本期的 5 篇论文如下: [00:42] TOP1(🔥188) | ⚙ DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI(DataFlow:面向数据为中心AI时代的统一数据准备与工作流自动化LLM驱动框架) [02:34] TOP2(🔥105) | 🔬 Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows(通过科学家对齐的工作流程探究大语言模型的科学通用智能) [0...
2025.12.26 | 暗号token涨点视觉推理;3D便签本让视频长脑子 26.12.2025 4:37
本期的 6 篇论文如下: [00:19] 🧠 Latent Implicit Visual Reasoning(潜在隐式视觉推理) [00:56] 🎬 Spatia: Video Generation with Updatable Spatial Memory(Spatia:基于可更新空间记忆的视频生成) [01:36] 🧠 Schoenfeld's Anatomy of Mathematical Reasoning by Language Models(基于舍恩菲尔德理论的语言模型数学推理解剖) [02:11] 🔍 How Much 3D Do Video Foundation Models Encode?(视频基础模型编码了多少3D信息...
2025.12.25 | 四维动态理解刷新VLM;单卡200倍速生成高清视频 25.12.2025 10:24
本期的 14 篇论文如下: [00:20] 🧠 Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models(学习在四维空间中推理:视觉语言模型的动态空间理解) [01:11] ⚡ TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times(TurboDiffusion:将视频扩散模型加速100-200倍) [01:52] 🧭 T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation(T2AV-Compass:迈...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.