duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2025.05.12 | 波兰语模型优化;高效参数利用 12.05.2025 5:40
本期的 7 篇论文如下: [00:23] 🇵 Bielik v3 Small: Technical Report(Bielik v3 Small:技术报告) [01:07] 🇵 Bielik 11B v2 Technical Report(Bielik 11B v2 技术报告) [01:42] 🤖 UniVLA: Learning to Act Anywhere with Task-centric Latent Actions(UniVLA:通过任务中心潜在动作学习在任意环境行动) [02:30] 🎨 G-FOCUS: Towards a Robust Method for Assessing UI Design Persuasiveness(G-FOCUS:迈向评估用户界面设...
【周末特辑】5月第2周最火AI论文 | 零数据自博弈推理;多模态长推理模型综述 10.05.2025 11:10
本期的 5 篇论文如下: [00:42] TOP1(🔥93) | 🚀 Absolute Zero: Reinforced Self-play Reasoning with Zero Data(绝对零度:基于零数据的强化自博弈推理) [02:38] TOP2(🔥91) | 🧠 Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models(感知、推理、思考与规划:大型多模态推理模型综述) [04:44] TOP3(🔥83) | 🧠 Unified Multimodal Chain-of-Thought Reward Model through Reinforcement F...
2025.05.09 | 多模态推理模型发展综述;通用智能评估框架提出 09.05.2025 10:53
本期的 15 篇论文如下: [00:22] 🧠 Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models(感知、推理、思考与规划:大型多模态推理模型综述) [00:57] 🤖 On Path to Multimodal Generalist: General-Level and General-Bench(迈向多模态通用智能:通用水平与通用基准) [01:40] 🤖 Flow-GRPO: Training Flow Matching Models via Online RL(Flow-GRPO:通过在线强化学习训练Flow Matching模...
2025.05.08 | 多模态模型整合潜力大;零搜索提升LLMs效率。 08.05.2025 10:32
本期的 14 篇论文如下: [00:21] 💡 Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities(统一多模态理解与生成模型:进展、挑战与机遇) [01:02] 🤖 ZeroSearch: Incentivize the Search Capability of LLMs without Searching(零搜索:无需搜索即可激励大型语言模型的搜索能力) [01:50] 🤔 Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Mode...
2025.05.07 | 多模态思维链提升模型性能;零数据自博弈强化推理能力。 07.05.2025 10:25
本期的 14 篇论文如下: [00:24] 🧠 Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning(基于强化微调的统一多模态思维链奖励模型) [01:10] 🤖 Absolute Zero: Reinforced Self-play Reasoning with Zero Data(绝对零度:零数据下的强化自博弈推理) [01:52] 🤸 FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios(FlexiAct:面向异构场景的灵活动作控制) [02:33] 🚀...
2025.05.06 | Voila实现低延迟全双工对话;RM-R1提升大模型推理奖励。 06.05.2025 11:14
本期的 15 篇论文如下: [00:22] 🤖 Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play(Voila:用于实时自主交互和语音角色扮演的语音-语言基础模型) [01:09] 🤔 RM-R1: Reward Modeling as Reasoning(RM-R1:将奖励建模视为推理) [01:52] 🧠 Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers(野外Grokking:用于Transforme...
2025.05.05 | PixelHacker提升图像修复质量;分层记忆增强图像编辑可控性。 05.05.2025 6:04
本期的 8 篇论文如下: [00:21] 🖼 PixelHacker: Image Inpainting with Structural and Semantic Consistency(PixelHacker:基于结构和语义一致性的图像修复) [01:01] 🎨 Improving Editability in Image Generation with Layer-wise Memory(通过分层记忆提升图像生成的可编辑性) [01:35] 🤖 Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts(超越一刀切:用于高效自然语言生成评...
【周末特辑】5月第1周最火AI论文 | 相机运动理解显著提升;单样本强化学习提升推理能力。 03.05.2025 13:07
本期的 5 篇论文如下: [00:43] TOP1(🔥149) | 🎥 Towards Understanding Camera Motions in Any Video(迈向理解任意视频中的相机运动) [03:05] TOP2(🔥74) | 🧠 Reinforcement Learning for Reasoning in Large Language Models with One Training Example(单样本强化学习赋能大语言模型推理) [05:48] TOP3(🔥54) | 🎭 The Leaderboard Illusion(排行榜的幻觉) [07:58] TOP4(🔥51) | 🔍 UniversalRAG: Retrieval-Augmented...
2025.05.02 | 交互式视频生成技术探讨;DeepCritic提升大模型评判能力。 02.05.2025 6:18
本期的 8 篇论文如下: [00:28] 🎮 A Survey of Interactive Generative Video(交互式生成视频综述) [01:05] 🧐 DeepCritic: Deliberate Critique with Large Language Models(DeepCritic: 基于大语言模型的审慎评判) [01:38] 🖼 T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT(T2I-R1:通过协作式语义级和令牌级思维链强化图像生成) [02:15] 👄 KeySync: A Robust Approach f...
2025.05.01 | 阿拉伯语变音难题新解;深度推理模型能力增强 01.05.2025 9:56
本期的 14 篇论文如下: [00:21] 🗣 Sadeed: Advancing Arabic Diacritization Through Small Language Model(Sadeed:通过小型语言模型推进阿拉伯语变音) [01:05] 🔎 WebThinker: Empowering Large Reasoning Models with Deep Research Capability(WebThinker:利用深度研究能力增强大型推理模型) [01:43] 🧮 Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math(Phi-4-Mini-Reasoning...
2025.04.30 | 多模态检索增强生成;单样本强化学习提升推理。 30.04.2025 8:59
本期的 12 篇论文如下: [00:24] 🔍 UniversalRAG: Retrieval-Augmented Generation over Multiple Corpora with Diverse Modalities and Granularities(通用RAG:基于多模态、多粒度异构语料库的检索增强生成) [01:06] 🧠 Reinforcement Learning for Reasoning in Large Language Models with One Training Example(单样本强化学习赋能大语言模型推理) [01:52] 🧠 ReasonIR: Training Retrievers for Reasoning Tasks(Reaso...
2025.04.29 | RepText提升多语言文本渲染;LLM改进手机GUI自动化。 29.04.2025 8:41
本期的 11 篇论文如下: [00:23] ✍ RepText: Rendering Visual Text via Replicating(RepText:通过复制渲染视觉文本) [01:02] 📱 LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects(LLM驱动的手机GUI代理:进展与展望) [01:44] 🔐 CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges(CipherBank:通过密码学挑战探索大型语言模型推理能力的边...
2025.04.28 | 视频相机运动理解提升;多模态推理模型优化 28.04.2025 8:00
本期的 11 篇论文如下: [00:22] 🎥 Towards Understanding Camera Motions in Any Video(迈向理解任意视频中的相机运动) [01:04] 🧠 Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning(Skywork R1V2:用于推理的多模态混合强化学习) [01:49] 💡 BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs(BitNet v2:用于1-bit LLM的具有哈达玛变换的原生4-bit激活) [02:28]...
【周末特辑】4月第4周最火AI论文 | 阿拉伯语模型扩展成功;强化学习提升有限。 26.04.2025 11:57
本期的 5 篇论文如下: [00:33] TOP1(🔥108) | 💡 Kuwain 1.5B: An Arabic SLM via Language Injection(Kuwain 1.5B:一种基于语言注入的阿拉伯语SLM) [02:43] TOP2(🔥98) | 🤔 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?(强化学习真的能激励大语言模型产生超越基础模型的推理能力吗?) [04:58] TOP3(🔥78) | 🤖 TTRL: Test-Time Reinforcement Learning(测试时强化...
2025.04.25 | 开源模型超越闭源;新型评估指标提升生成质量。 25.04.2025 10:52
本期的 15 篇论文如下: [00:24] 🖼 Step1X-Edit: A Practical Framework for General Image Editing(Step1X-Edit:一个通用的图像编辑实用框架) [01:05] 🖼 RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation(RefVNLI:面向主体驱动的文本到图像生成的可扩展评估) [01:48] 🤖 Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning(Paper2Code:从机器学习科学...
2025.04.24 | 视觉推理评估新基准;高保真人脸替换技术 24.04.2025 10:25
本期的 14 篇论文如下: [00:23] 👁 VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models(VisuLogic:一个用于评估多模态大型语言模型中视觉推理能力的基准) [01:08] 🎭 DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning(DreamID:基于Triplet ID Group Learning的高保真快速扩散人脸替换) [01:46] 🌐 Trillion 7B Technical Report(...
2025.04.23 | 阿拉伯语性能提升;推理任务性能显著提高。 23.04.2025 10:53
本期的 15 篇论文如下: [00:22] 💡 Kuwain 1.5B: An Arabic SLM via Language Injection(Kuwain 1.5B:一种基于语言注入的阿拉伯语SLM) [00:58] 🤖 TTRL: Test-Time Reinforcement Learning(测试时强化学习) [01:40] 🌍 The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks(从2000+多语种评测基准中汲取的惨痛教训) [02:23] 🖼 Describe Anything: Detailed Localized Image and Video Captioning(描述一切:细...
2025.04.22 | LUFFY提升推理性能;FlowReasoner增强系统适应性。 22.04.2025 10:53
本期的 15 篇论文如下: [00:25] 🧠 Learning to Reason under Off-Policy Guidance(离线策略指导下的推理学习) [01:00] 🤖 FlowReasoner: Reinforcing Query-Level Meta-Agents(FlowReasoner:强化查询级别元代理) [01:40] 🦅 Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models(Eagle 2.5:提升前沿视觉-语言模型长文本后训练性能) [02:22] 🧰 ToolRL: Reward is All Tool Learning Nee...
2025.04.21 | 强化学习未提升新推理能力;MIG优化指令微调数据选择。 21.04.2025 6:50
本期的 9 篇论文如下: [00:22] 🤔 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?(强化学习真的能激励大语言模型产生超越基础模型的推理能力吗?) [00:59] 🧠 MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space(MIG:通过最大化语义空间中的信息增益实现指令微调的自动数据选择) [01:41] 🤔 Could Thinking Mult...
【周末特辑】4月第3周最火AI论文 | 多模态模型InternVL3创新预训练;Seaweed-7B高效视频生成。 19.04.2025 13:07
本期的 5 篇论文如下: [00:52] TOP1(🔥223) | 🖼 InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models(InternVL3:探索开源多模态模型的高级训练和测试时方案) [03:22] TOP2(🔥117) | 🎬 Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model(Seaweed-7B:一种经济高效的视频生成基础模型训练方法) [05:40] TOP3(🔥112) | 🏠 PRIMA.CPP: Speeding Up 70B-...
2025.04.18 | CLIMB提升领域模型表现;反蒸馏采样防止模型被盗用。 18.04.2025 10:49
本期的 15 篇论文如下: [00:23] 🗂 CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training(CLIMB:基于聚类的迭代数据混合引导预训练方法) [01:03] 🧪 Antidistillation Sampling(反蒸馏采样) [01:41] 🤝 A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis(小型LLM的策略协调框架在数据合成方面与大型LLM相媲美) [02:26] 🎬 Packing Input...
2025.04.17 | ColorBench测试VLM颜色理解;BitNet提升计算效率。 17.04.2025 8:21
本期的 11 篇论文如下: [00:27] 🎨 ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness(ColorBench:视觉语言模型能否看到并理解多彩世界?一个关于颜色感知、推理和鲁棒性的综合基准) [01:09] 💡 BitNet b1.58 2B4T Technical Report(BitNet b1.58 2B4T 技术报告) [01:50] 🎨 Cobra: Efficient Line Art COlorization with BRoAder R...
2025.04.16 | Genius提升LLM推理能力;xVerify高效验证推理模型。 16.04.2025 11:34
本期的 15 篇论文如下: [00:22] 🧠 Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning(Genius:一种用于高级推理的通用且纯粹的无监督自训练框架) [01:06] ✅ xVerify: Efficient Answer Verifier for Reasoning Model Evaluations(xVerify:用于推理模型评估的高效答案验证器) [01:52] 🖼 Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding(Pixel-SAIL:用...
2025.04.15 | 多模态模型性能提升;低资源推理加速优化 15.04.2025 11:25
本期的 15 篇论文如下: [00:23] 🖼 InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models(InternVL3:探索开源多模态模型的高级训练和测试时方案) [01:03] 🏠 PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters(PRIMA.CPP: 加速低资源家用集群上700亿参数规模大语言模型的推理) [01:46] 🖼 FUSION: Fully Integration of Vision-Language R...
2025.04.14 | 经济高效视频生成;自回归图像生成扩展。 14.04.2025 9:40
本期的 13 篇论文如下: [00:24] 🎬 Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model(Seaweed-7B:一种经济高效的视频生成基础模型训练方法) [01:00] 🖼 GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation(GigaTok:将视觉标记器扩展到30亿参数以进行自回归图像生成) [01:42] 🎮 MineWorld: a Real-Time and Open-Source Interactive World Model o...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.