duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2025.09.30 | SLA稀疏注意力砍算力;StableToken抗噪不训模 30.09.2025

本期的 15 篇论文如下: [00:22] ⚡ SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention(SLA:通过可微调稀疏线性注意力突破扩散Transformer的稀疏性极限) [01:05] 🗣 StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs(StableToken:一种面向韧性SpeechLLM的噪声鲁棒语义语音分词器) [01:54] 🎮 Multiplayer Nash Preference Optimization(多玩家纳什...

2025.09.29 | 实时长视频边聊边播;分位数基线稳控推理熵 29.09.2025

本期的 15 篇论文如下: [00:20] 🎬 LongLive: Real-time Interactive Long Video Generation(LongLive:实时交互式长视频生成框架) [00:56] 🎯 Quantile Advantage Estimation for Entropy-Safe Reasoning(用于熵安全推理的分位数优势估计) [01:34] 📄 MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing(MinerU2.5:面向高效高分辨率文档解析的解耦视觉-语言模型) [02:11] 🧠...

【周末特辑】9月第5周最火AI论文 | Qwen3-Omni开源称王; 锁定视觉训解码,Baseer刷新阿文OCR; 27.09.2025

本期的 5 篇论文如下: [00:38] TOP1(🔥116) | 📜 Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR(Baseer:面向阿拉伯文档OCR的视觉-语言模型) [02:43] TOP2(🔥113) | 🌐 Qwen3-Omni Technical Report(Qwen3-Omni技术报告:首个无性能损耗的全模态大模型) [05:23] TOP3(🔥112) | 🗺 RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation(RPG:用于统一可扩展代码库生成的仓...

2025.09.26 | SciReasoner八项全能;MMR1模糊区炼出开源多模态 26.09.2025

本期的 15 篇论文如下: [00:20] 🔬 SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines(SciReasoner:跨学科夯实科学推理基石) [01:00] 🧠 MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources(MMR1:基于方差感知采样与开放资源的多模态推理增强) [01:41] 📈 VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models(VCRL:面向大语...

2025.09.25 | 视频模型零样本全能;隐式思维链省token提效 25.09.2025

本期的 10 篇论文如下: [00:22] 🎥 Video models are zero-shot learners and reasoners(视频模型是零样本学习者与推理者) [01:09] 🧠 SIM-CoT: Supervised Implicit Chain-of-Thought(SIM-CoT:基于监督式隐式思维链的高效推理) [01:55] 🪶 EmbeddingGemma: Powerful and Lightweight Text Representations(EmbeddingGemma:强大而轻量的文本表征模型) [02:29] 🗣 Advancing Speech Understanding in Speech-Aware Language...

2025.09.24 | 阿语OCR刷新指标;无标注RL涨分 24.09.2025

本期的 15 篇论文如下: [00:24] 📜 Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR(Baseer:面向阿拉伯文档OCR的视觉-语言模型) [00:58] 🚀 Reinforcement Learning on Pre-Training Data(基于预训练数据的强化学习) [01:37] 👁 Do You Need Proprioceptive States in Visuomotor Policies?(无需本体感觉状态的视觉-运动策略是否可行?) [02:36] 🚀 MiniCPM-V 4.5: Cooking Efficient MLLMs via Arch...

2025.09.23 | 少78条示范让AI飙73.5%;免掩膜视频插主体超Pika 23.09.2025

本期的 15 篇论文如下: [00:21] 🚀 LIMI: Less is More for Agency(LIMI:少即是多,打造AI智能体) [00:55] 🎬 OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models(无需掩膜的视频任意主体插入:基于扩散Transformer模型) [01:28] 🧩 OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System(OnePiece:面向工业级级联排序系统的上下文工程与推...

2025.09.22 | 有向图驱动代码生成;双通道视觉统一模型 22.09.2025

本期的 13 篇论文如下: [00:25] 🗺 RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation(RPG:用于统一可扩展代码库生成的仓库规划图) [01:00] 🌉 MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer(MANZANO:基于混合视觉词元器的简洁可扩展统一多模态模型) [01:42] 🧩 Latent Zoning Network: A Unified Principle for Generative Modeling, Representa...

【周末特辑】9月第4周最火AI论文 | OmniWorld打造4D数据工厂;WebWeaver让AI边搜边写 20.09.2025

本期的 5 篇论文如下: [00:43] TOP1(🔥95) | 🌍 OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling(OmniWorld:面向4D世界建模的多领域多模态大规模数据集) [02:51] TOP2(🔥93) | 🔍 WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research(WebWeaver:面向开放型深度研究的动态提纲式网络证据结构化框架) [05:09] TOP3(🔥91) | 🤖 Scaling Agents via Cont...

2025.09.19 | 跨平台GUI模型刷榜;FlowRL分布匹配提推理 19.09.2025

本期的 15 篇论文如下: [00:26] 🖥 ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data(ScaleCUA:基于跨平台数据的开源计算机智能体规模化方案) [01:01] 🌊 FlowRL: Matching Reward Distributions for LLM Reasoning(FlowRL:通过流匹配奖励分布提升大语言模型推理能力) [01:57] 🧭 Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Delibration(跨越边界推理:借助...

2025.09.18 | FP8压缩+翻译微调低成本炼阿语大模型;2B-8B小模型洗数据硬刚GPT-4o 18.09.2025

本期的 14 篇论文如下: [00:19] 🐪 Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale(Hala技术报告:规模化构建阿拉伯语为中心的指令与翻译模型) [00:56] 🚀 SAIL-VL2 Technical Report(SAIL-VL2技术报告) [01:42] 🌐 PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era(全景视界:具身AI时代的360°视觉崛起) [02:33] 🎓 GenExam: A Multidisciplinary Text-...

2025.09.17 | WebWeaver框架提升可信长文报告;Agentic预训练扩展智能体系统 17.09.2025

本期的 11 篇论文如下: [00:27] 🔍 WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research(WebWeaver:面向开放型深度研究的动态提纲式网络证据结构化框架) [01:08] 🤖 Scaling Agents via Continual Pre-training(基于持续预训练扩展智能体系统规模的研究) [01:52] ⛵ WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Lea...

2025.09.16 | OmniWorld建4D数据底座;UI-S1半在线驯界面代理 16.09.2025

本期的 14 篇论文如下: [00:24] 🌍 OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling(OmniWorld:面向4D世界建模的多领域多模态大规模数据集) [01:12] 🤖 UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning(UI-S1:基于半在线强化学习的图形界面自动化新进展) [01:51] 🏠 InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts(InternScen...

2025.09.15 | 数据集升级测互动;模型大小非长程瓶颈 15.09.2025

本期的 14 篇论文如下: [00:25] 📚 IntrEx: A Dataset for Modeling Engagement in Educational Conversations(IntrEx:面向教育对话中参与度建模的数据集) [01:02] 📏 The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs(“收益递减的幻觉”:衡量大语言模型的长时程执行能力) [01:54] 🧩 X-Part: high fidelity and structure coherent shape decomposition(X-Part:高保真且结构一致的三维形...

【周末特辑】9月第3周最火AI论文 | 群智RL提速大模型;小VLA零预训练控机械 14.09.2025

本期的 5 篇论文如下: [00:40] TOP1(🔥455) | 🤝 Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing(共享即关爱:基于集体RL经验共享的高效大模型后训练) [03:19] TOP2(🔥163) | 🤖 VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model(VLA-Adapter:面向小型视觉-语言-动作模型的有效范式) [05:44] TOP3(🔥156) | 🤔 Why Language Models Hallucinate(语...

2025.09.12 | HuMo多模态控人视频;SimpleVLA-RL强化升效 12.09.2025

本期的 15 篇论文如下: [00:27] 🎭 HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning(HuMo:通过协同多模态条件控制实现以人为中心的视频生成) [01:18] 🤖 SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning(SimpleVLA-RL:通过强化学习实现VLA训练规模化) [02:02] 🗣 EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs(EchoX:...

2025.09.11 | 强化学习提升推理能力;奖励缩放优化视觉生成 11.09.2025

本期的 10 篇论文如下: [00:24] 🧠 A Survey of Reinforcement Learning for Large Reasoning Models(大型推理模型的强化学习综述) [00:45] 🔄 RewardDance: Reward Scaling in Visual Generation(RewardDance:视觉生成中的奖励缩放) [01:08] 🌐 3D and 4D World Modeling: A Survey(3D和4D世界建模:一项综述) [01:41] 🤖 AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinfor...

2025.09.10 | 强化学习并行思维;视觉搜索推理扩展 10.09.2025

本期的 14 篇论文如下: [00:22] 🧠 Parallel-R1: Towards Parallel Thinking via Reinforcement Learning(Parallel-R1: 通过强化学习实现并行思维) [00:50] 🔍 Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search(Mini-o3:扩展视觉搜索中的推理模式与交互轮次) [01:15] 👁 Visual Representation Alignment for Multimodal Large Language Models(多模态大语言模型的视觉表征对齐) [01:54]...

2025.09.09 | REER提升推理性能;WebExplorer训练智能体 09.09.2025

本期的 15 篇论文如下: [00:21] 💡 Reverse-Engineered Reasoning for Open-Ended Generation(面向开放式生成的逆向工程推理) [00:47] 🌐 WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents(WebExplorer:探索与演进,用于训练长周期网络智能体) [01:17] 🚀 Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models(革新扩散大语言模型的强化学习框架) [01:38] 🤔 Doe...

2025.09.08 | 语言模型幻觉源于预训练;大模型图形编程性能提升 08.09.2025

本期的 12 篇论文如下: [00:24] 🤔 Why Language Models Hallucinate(语言模型为何产生幻觉) [00:47] 🎨 Symbolic Graphics Programming with Large Language Models(使用大型语言模型进行符号化图形编程) [01:17] ⚡ Set Block Decoding is a Language Model Inference Accelerator(集合块解码:一种语言模型推理加速器) [01:43] 🎼 WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning(WildScore:多模...

【周末特辑】9月第2周最火AI论文 | LLM智能体RL综述;AI代码安全基准 06.09.2025

本期的 5 篇论文如下: [00:35] TOP1(🔥139) | 🤖 The Landscape of Agentic Reinforcement Learning for LLMs: A Survey(面向大语言模型的智能体强化学习全景:一项综述) [01:52] TOP2(🔥133) | 🔒 A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code(A.S.E:一个用于评估AI生成代码安全的仓库级基准) [02:57] TOP3(🔥127) | 🤖 A Survey of Scientific Large Language Models: From Data Fo...

2025.09.05 | 大型语言模型语义理解弱;图像编辑模型提升几何估计 05.09.2025

本期的 13 篇论文如下: [00:22] 🤔 Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth(废话学:用深度解读无意义内容挑战大型语言模型) [00:47] 📐 From Editor to Dense Geometry Estimator(从编辑模型到密集几何估计器) [01:08] 🧠 Towards a Unified View of Large Language Model Post-Training(迈向大语言模型后训练的统一视角) [01:39] 🔄 Inverse IFEval: Can LLMs Unlearn Stubborn Training...

2025.09.04 | 机器人任务规划高效;数据推理能力提升 04.09.2025

本期的 5 篇论文如下: [00:24] 🤖 Robix: A Unified Model for Robot Interaction, Reasoning and Planning(Robix:一个用于机器人交互、推理和规划的统一模型) [00:54] 🔍 Open Data Synthesis For Deep Research(面向深度研究的开放数据合成) [01:30] 🧠 LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations(LMEnt:一套分析语言模型从预训练数据到表示的知识套件) [02...

2025.09.03 | 智能体RL提升大模型自主性;SimpleTIR解多轮工具推理 03.09.2025

本期的 15 篇论文如下: [00:19] 🤖 The Landscape of Agentic Reinforcement Learning for LLMs: A Survey(面向大语言模型的智能体强化学习全景:一项综述) [00:40] 🚀 SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning(SimpleTIR:面向多轮工具集成推理的端到端强化学习) [01:12] 🤖 UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning(UI-T...

2025.09.02 | PVPO优化推理性能;T2R-bench暴露模型短板 02.09.2025

本期的 6 篇论文如下: [00:23] 🧠 PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning(PVPO:基于预估值策略优化的智能体推理方法) [00:49] 📊 T2R-bench: A Benchmark for Generating Article-Level Reports from Real World Industrial Tables(T2R-bench:一个用于从真实世界工业表格生成文章级报告的基准测试) [01:18] 🔍 No Label Left Behind: A Unified Surface Defect Detection Model for a...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.