duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2025.05.12 | 波兰语模型优化;高效参数利用 12.05.2025

本期的 7 篇论文如下: [00:23] 🇵 Bielik v3 Small: Technical Report(Bielik v3 Small:技术报告) [01:07] 🇵 Bielik 11B v2 Technical Report(Bielik 11B v2 技术报告) [01:42] 🤖 UniVLA: Learning to Act Anywhere with Task-centric Latent Actions(UniVLA:通过任务中心潜在动作学习在任意环境行动) [02:30] 🎨 G-FOCUS: Towards a Robust Method for Assessing UI Design Persuasiveness(G-FOCUS:迈向评估用户界面设...

【周末特辑】5月第2周最火AI论文 | 零数据自博弈推理;多模态长推理模型综述 10.05.2025

本期的 5 篇论文如下: [00:42] TOP1(🔥93) | 🚀 Absolute Zero: Reinforced Self-play Reasoning with Zero Data(绝对零度:基于零数据的强化自博弈推理) [02:38] TOP2(🔥91) | 🧠 Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models(感知、推理、思考与规划:大型多模态推理模型综述) [04:44] TOP3(🔥83) | 🧠 Unified Multimodal Chain-of-Thought Reward Model through Reinforcement F...

2025.05.09 | 多模态推理模型发展综述;通用智能评估框架提出 09.05.2025

本期的 15 篇论文如下: [00:22] 🧠 Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models(感知、推理、思考与规划:大型多模态推理模型综述) [00:57] 🤖 On Path to Multimodal Generalist: General-Level and General-Bench(迈向多模态通用智能:通用水平与通用基准) [01:40] 🤖 Flow-GRPO: Training Flow Matching Models via Online RL(Flow-GRPO:通过在线强化学习训练Flow Matching模...

2025.05.08 | 多模态模型整合潜力大;零搜索提升LLMs效率。 08.05.2025

本期的 14 篇论文如下: [00:21] 💡 Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities(统一多模态理解与生成模型:进展、挑战与机遇) [01:02] 🤖 ZeroSearch: Incentivize the Search Capability of LLMs without Searching(零搜索:无需搜索即可激励大型语言模型的搜索能力) [01:50] 🤔 Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Mode...

2025.05.07 | 多模态思维链提升模型性能;零数据自博弈强化推理能力。 07.05.2025

本期的 14 篇论文如下: [00:24] 🧠 Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning(基于强化微调的统一多模态思维链奖励模型) [01:10] 🤖 Absolute Zero: Reinforced Self-play Reasoning with Zero Data(绝对零度:零数据下的强化自博弈推理) [01:52] 🤸 FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios(FlexiAct:面向异构场景的灵活动作控制) [02:33] 🚀...

2025.05.06 | Voila实现低延迟全双工对话;RM-R1提升大模型推理奖励。 06.05.2025

本期的 15 篇论文如下: [00:22] 🤖 Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play(Voila:用于实时自主交互和语音角色扮演的语音-语言基础模型) [01:09] 🤔 RM-R1: Reward Modeling as Reasoning(RM-R1:将奖励建模视为推理) [01:52] 🧠 Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers(野外Grokking:用于Transforme...

2025.05.05 | PixelHacker提升图像修复质量;分层记忆增强图像编辑可控性。 05.05.2025

本期的 8 篇论文如下: [00:21] 🖼 PixelHacker: Image Inpainting with Structural and Semantic Consistency(PixelHacker:基于结构和语义一致性的图像修复) [01:01] 🎨 Improving Editability in Image Generation with Layer-wise Memory(通过分层记忆提升图像生成的可编辑性) [01:35] 🤖 Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts(超越一刀切:用于高效自然语言生成评...

【周末特辑】5月第1周最火AI论文 | 相机运动理解显著提升;单样本强化学习提升推理能力。 03.05.2025

本期的 5 篇论文如下: [00:43] TOP1(🔥149) | 🎥 Towards Understanding Camera Motions in Any Video(迈向理解任意视频中的相机运动) [03:05] TOP2(🔥74) | 🧠 Reinforcement Learning for Reasoning in Large Language Models with One Training Example(单样本强化学习赋能大语言模型推理) [05:48] TOP3(🔥54) | 🎭 The Leaderboard Illusion(排行榜的幻觉) [07:58] TOP4(🔥51) | 🔍 UniversalRAG: Retrieval-Augmented...

2025.05.02 | 交互式视频生成技术探讨;DeepCritic提升大模型评判能力。 02.05.2025

本期的 8 篇论文如下: [00:28] 🎮 A Survey of Interactive Generative Video(交互式生成视频综述) [01:05] 🧐 DeepCritic: Deliberate Critique with Large Language Models(DeepCritic: 基于大语言模型的审慎评判) [01:38] 🖼 T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT(T2I-R1:通过协作式语义级和令牌级思维链强化图像生成) [02:15] 👄 KeySync: A Robust Approach f...

2025.05.01 | 阿拉伯语变音难题新解;深度推理模型能力增强 01.05.2025

本期的 14 篇论文如下: [00:21] 🗣 Sadeed: Advancing Arabic Diacritization Through Small Language Model(Sadeed:通过小型语言模型推进阿拉伯语变音) [01:05] 🔎 WebThinker: Empowering Large Reasoning Models with Deep Research Capability(WebThinker:利用深度研究能力增强大型推理模型) [01:43] 🧮 Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math(Phi-4-Mini-Reasoning...

2025.04.30 | 多模态检索增强生成;单样本强化学习提升推理。 30.04.2025

本期的 12 篇论文如下: [00:24] 🔍 UniversalRAG: Retrieval-Augmented Generation over Multiple Corpora with Diverse Modalities and Granularities(通用RAG:基于多模态、多粒度异构语料库的检索增强生成) [01:06] 🧠 Reinforcement Learning for Reasoning in Large Language Models with One Training Example(单样本强化学习赋能大语言模型推理) [01:52] 🧠 ReasonIR: Training Retrievers for Reasoning Tasks(Reaso...

2025.04.29 | RepText提升多语言文本渲染;LLM改进手机GUI自动化。 29.04.2025

本期的 11 篇论文如下: [00:23] ✍ RepText: Rendering Visual Text via Replicating(RepText:通过复制渲染视觉文本) [01:02] 📱 LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects(LLM驱动的手机GUI代理:进展与展望) [01:44] 🔐 CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges(CipherBank:通过密码学挑战探索大型语言模型推理能力的边...

2025.04.28 | 视频相机运动理解提升;多模态推理模型优化 28.04.2025

本期的 11 篇论文如下: [00:22] 🎥 Towards Understanding Camera Motions in Any Video(迈向理解任意视频中的相机运动) [01:04] 🧠 Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning(Skywork R1V2:用于推理的多模态混合强化学习) [01:49] 💡 BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs(BitNet v2:用于1-bit LLM的具有哈达玛变换的原生4-bit激活) [02:28]...

【周末特辑】4月第4周最火AI论文 | 阿拉伯语模型扩展成功;强化学习提升有限。 26.04.2025

本期的 5 篇论文如下: [00:33] TOP1(🔥108) | 💡 Kuwain 1.5B: An Arabic SLM via Language Injection(Kuwain 1.5B:一种基于语言注入的阿拉伯语SLM) [02:43] TOP2(🔥98) | 🤔 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?(强化学习真的能激励大语言模型产生超越基础模型的推理能力吗?) [04:58] TOP3(🔥78) | 🤖 TTRL: Test-Time Reinforcement Learning(测试时强化...

2025.04.25 | 开源模型超越闭源;新型评估指标提升生成质量。 25.04.2025

本期的 15 篇论文如下: [00:24] 🖼 Step1X-Edit: A Practical Framework for General Image Editing(Step1X-Edit:一个通用的图像编辑实用框架) [01:05] 🖼 RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation(RefVNLI:面向主体驱动的文本到图像生成的可扩展评估) [01:48] 🤖 Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning(Paper2Code:从机器学习科学...

2025.04.24 | 视觉推理评估新基准;高保真人脸替换技术 24.04.2025

本期的 14 篇论文如下: [00:23] 👁 VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models(VisuLogic:一个用于评估多模态大型语言模型中视觉推理能力的基准) [01:08] 🎭 DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning(DreamID:基于Triplet ID Group Learning的高保真快速扩散人脸替换) [01:46] 🌐 Trillion 7B Technical Report(...

2025.04.23 | 阿拉伯语性能提升;推理任务性能显著提高。 23.04.2025

本期的 15 篇论文如下: [00:22] 💡 Kuwain 1.5B: An Arabic SLM via Language Injection(Kuwain 1.5B:一种基于语言注入的阿拉伯语SLM) [00:58] 🤖 TTRL: Test-Time Reinforcement Learning(测试时强化学习) [01:40] 🌍 The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks(从2000+多语种评测基准中汲取的惨痛教训) [02:23] 🖼 Describe Anything: Detailed Localized Image and Video Captioning(描述一切:细...

2025.04.22 | LUFFY提升推理性能;FlowReasoner增强系统适应性。 22.04.2025

本期的 15 篇论文如下: [00:25] 🧠 Learning to Reason under Off-Policy Guidance(离线策略指导下的推理学习) [01:00] 🤖 FlowReasoner: Reinforcing Query-Level Meta-Agents(FlowReasoner:强化查询级别元代理) [01:40] 🦅 Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models(Eagle 2.5:提升前沿视觉-语言模型长文本后训练性能) [02:22] 🧰 ToolRL: Reward is All Tool Learning Nee...

2025.04.21 | 强化学习未提升新推理能力;MIG优化指令微调数据选择。 21.04.2025

本期的 9 篇论文如下: [00:22] 🤔 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?(强化学习真的能激励大语言模型产生超越基础模型的推理能力吗?) [00:59] 🧠 MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space(MIG:通过最大化语义空间中的信息增益实现指令微调的自动数据选择) [01:41] 🤔 Could Thinking Mult...

【周末特辑】4月第3周最火AI论文 | 多模态模型InternVL3创新预训练;Seaweed-7B高效视频生成。 19.04.2025

本期的 5 篇论文如下: [00:52] TOP1(🔥223) | 🖼 InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models(InternVL3:探索开源多模态模型的高级训练和测试时方案) [03:22] TOP2(🔥117) | 🎬 Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model(Seaweed-7B:一种经济高效的视频生成基础模型训练方法) [05:40] TOP3(🔥112) | 🏠 PRIMA.CPP: Speeding Up 70B-...

2025.04.18 | CLIMB提升领域模型表现;反蒸馏采样防止模型被盗用。 18.04.2025

本期的 15 篇论文如下: [00:23] 🗂 CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training(CLIMB:基于聚类的迭代数据混合引导预训练方法) [01:03] 🧪 Antidistillation Sampling(反蒸馏采样) [01:41] 🤝 A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis(小型LLM的策略协调框架在数据合成方面与大型LLM相媲美) [02:26] 🎬 Packing Input...

2025.04.17 | ColorBench测试VLM颜色理解;BitNet提升计算效率。 17.04.2025

本期的 11 篇论文如下: [00:27] 🎨 ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness(ColorBench:视觉语言模型能否看到并理解多彩世界?一个关于颜色感知、推理和鲁棒性的综合基准) [01:09] 💡 BitNet b1.58 2B4T Technical Report(BitNet b1.58 2B4T 技术报告) [01:50] 🎨 Cobra: Efficient Line Art COlorization with BRoAder R...

2025.04.16 | Genius提升LLM推理能力;xVerify高效验证推理模型。 16.04.2025

本期的 15 篇论文如下: [00:22] 🧠 Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning(Genius:一种用于高级推理的通用且纯粹的无监督自训练框架) [01:06] ✅ xVerify: Efficient Answer Verifier for Reasoning Model Evaluations(xVerify:用于推理模型评估的高效答案验证器) [01:52] 🖼 Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding(Pixel-SAIL:用...

2025.04.15 | 多模态模型性能提升;低资源推理加速优化 15.04.2025

本期的 15 篇论文如下: [00:23] 🖼 InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models(InternVL3:探索开源多模态模型的高级训练和测试时方案) [01:03] 🏠 PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters(PRIMA.CPP: 加速低资源家用集群上700亿参数规模大语言模型的推理) [01:46] 🖼 FUSION: Fully Integration of Vision-Language R...

2025.04.14 | 经济高效视频生成;自回归图像生成扩展。 14.04.2025

本期的 13 篇论文如下: [00:24] 🎬 Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model(Seaweed-7B:一种经济高效的视频生成基础模型训练方法) [01:00] 🖼 GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation(GigaTok:将视觉标记器扩展到30亿参数以进行自回归图像生成) [01:42] 🎮 MineWorld: a Real-Time and Open-Source Interactive World Model o...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.