duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

【周末特辑】1月第2周最火AI论文 | MiniMax-01扩展长上下文处理,数学推理PRM提升过程监督。 18.01.2025

本期的 5 篇论文如下: [00:35] TOP1(🔥258) | ⚡ MiniMax-01: Scaling Foundation Models with Lightning Attention(MiniMax-01:基于闪电注意力机制扩展基础模型) [02:52] TOP2(🔥77) | 📊 The Lessons of Developing Process Reward Models in Mathematical Reasoning(数学推理中过程奖励模型开发的经验教训) [05:06] TOP3(🔥66) | 🧠 Tensor Product Attention Is All You Need(张量积注意力机制是关键) [06:49] TOP4(🔥...

2025.01.17 | OmniThink提升机器写作深度与新颖性,扩散模型推理扩展提升生成质量。 18.01.2025

本期的 12 篇论文如下: [00:26] 🧠 OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking(OmniThink:通过思考扩展机器写作的知识边界) [01:06] 🔍 Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps(扩散模型推理时扩展:超越去噪步骤的扩展) [01:37] 🩺 Exploring the Inquiry-Diagnosis Relationship with Advanced Patient Simulators(探索高级患者模拟器中的问...

2025.01.16 | MMDocIR推动多模态检索标准化,CityDreamer4D创新4D城市生成模型。 16.01.2025

本期的 9 篇论文如下: [00:25] 📊 MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents(MMDocIR:长文档多模态检索的基准测试) [01:06] 🏙 CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities(CityDreamer4D:无界4D城市的组合生成模型) [01:49] 🎥 RepVideo: Rethinking Cross-Layer Representation for Video Generation(RepVideo:重新思考视频生成中的跨层表示) [02:30] 📚 Towards Be...

2025.01.15 | MiniMax-01扩展基础模型处理长上下文,填充符在T2I模型中影响图像生成。 15.01.2025

本期的 15 篇论文如下: [00:23] ⚡ MiniMax-01: Scaling Foundation Models with Lightning Attention(MiniMax-01:基于闪电注意力机制扩展基础模型) [01:04] 🖼 Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models(填充符:T2I模型中填充符的机制分析) [01:44] 🎨 MangaNinja: Line Art Colorization with Precise Reference Following(MangaNinja:基于精确参考跟随的线稿上色) [02:21] 🧬 A Multi-Mo...

2025.01.14 | 数学推理提升,内存开销减少 14.01.2025

本期的 11 篇论文如下: [00:24] 📊 The Lessons of Developing Process Reward Models in Mathematical Reasoning(数学推理中过程奖励模型开发的经验教训) [01:10] 🧠 Tensor Product Attention Is All You Need(张量积注意力机制是关键) [01:53] 🤖 $\text{Transformer}^2$: Self-adaptive LLMs(Transformer²:自适应大型语言模型) [02:34] 🎥 VideoAuteur: Towards Long Narrative Video Generation(视频导演:面向长篇...

2025.01.13 | OmniManip实现通用机器人操作,VideoRAG提升视频检索生成性能。 13.01.2025

本期的 10 篇论文如下: [00:24] 🤖 OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints(OmniManip:通过以对象为中心的交互原语作为空间约束实现通用机器人操作) [01:02] 🎥 VideoRAG: Retrieval-Augmented Generation over Video Corpus(VideoRAG:基于视频语料库的检索增强生成) [01:38] 🎥 OVO-Bench: How Far is Your Video-LLMs from Real-World Onlin...

【周末特辑】1月第1周最火AI论文 | 小型模型超越大型模型,REINFORCE++简化对齐方法 11.01.2025

本期的 5 篇论文如下: [00:39] TOP1(🔥173) | 🧠 rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking(rStar-Math:小型语言模型通过自我进化的深度思考掌握数学推理) [03:03] TOP2(🔥71) | 🚀 REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models(REINFORCE++:一种简单高效的大语言模型对齐方法) [05:17] TOP3(🔥63) | 🧠 Towards System 2 Reasoning in LLM...

2025.01.10 每日AI论文 | GAN训练简化性能提升,视频自回归预训练竞争力显著。 10.01.2025

本期的 7 篇论文如下: [00:23] 🧠 The GAN is dead; long live the GAN! A Modern GAN Baseline(GAN已死;GAN万岁!一个现代的GAN基线) [01:02] 🎥 An Empirical Study of Autoregressive Pre-training from Videos(视频自回归预训练的实证研究) [01:49] 🚗 Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives(视觉语言模型是否准备好用于自动驾驶?从可靠性...

2025.01.09 每日AI论文 | 小型模型自我进化超越GPT-3,多模态模型提升数学推理能力。 09.01.2025

本期的 11 篇论文如下: [00:25] 🧠 rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking(rStar-Math:小型语言模型通过自我进化的深度思考掌握数学推理) [01:06] 🧠 URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics(URSA:理解与验证多模态数学中的思维链推理) [01:45] 🧠 Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Ch...

2025.01.08 每日AI论文 | REINFORCE++提升大模型对齐效率,MotionBench优化视频运动理解 08.01.2025

本期的 11 篇论文如下: [00:24] 🚀 REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models(REINFORCE++:一种简单高效的大语言模型对齐方法) [01:00] 🎥 MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models(MotionBench:用于评估和改进视觉语言模型细粒度视频运动理解的基准) [01:40] 🔍 Sa2VA: Marrying SAM2 with LLaVA for Dense...

2025.01.07 每日AI论文 | STAR提升视频超分辨率时空一致性,BoostStep增强大模型数学推理能力。 07.01.2025

本期的 16 篇论文如下: [00:24] 🎥 STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution(STAR:基于文本到视频模型的空间-时间增强用于现实世界视频超分辨率) [01:06] 🧮 BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning(BoostStep:通过改进单步推理提升大语言模型的数学能力) [01:44] 🤖 Dispider: Enabling...

2025.01.06 每日AI论文 | EnerVerse提升机器人操作规划能力,VITA-1.5优化实时视觉语音交互。 06.01.2025

本期的 8 篇论文如下: [00:24] 🤖 EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation(EnerVerse:面向机器人操作的具身未来空间构想) [00:58] 🤖 VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction(VITA-1.5:迈向GPT-4o级别的实时视觉与语音交互) [01:33] 🤔 Virgo: A Preliminary Exploration on Reproducing o1-like MLLM(Virgo:关于复现o1类多模态大语言模型的初步探索...

【月末特辑】12月最火AI论文 | Qwen2.5提升大语言模型性能,阿波罗优化视频理解效率。 05.01.2025

本期的 10 篇论文如下: [00:31] TOP1(🔥335) | 🤖 Qwen2.5 Technical Report(Qwen2.5技术报告) [02:44] TOP2(🔥136) | 🎥 Apollo: An Exploration of Video Understanding in Large Multimodal Models(阿波罗:大型多模态模型中的视频理解探索) [05:01] TOP3(🔥123) | 🚀 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling(通过模型、数据和测试时扩展提升开源多...

【周末特辑】12月第5周最火AI论文 | 提升医学推理能力,自动化GUI轨迹构建。 04.01.2025

本期的 5 篇论文如下: [00:35] TOP1(🔥83) | 🧠 HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs(华佗GPT-o1:迈向医学复杂推理的大语言模型) [02:49] TOP2(🔥65) | 🤖 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis(OS-Genesis:通过逆向任务合成自动化GUI代理轨迹构建) [04:50] TOP3(🔥63) | 🎨 1.58-bit FLUX(1.58位FLUX:首个成功量化最先进文本生成图像模型的方法...

2025.01.03 每日AI论文 | 多模态教科书提升视觉语言模型性能,VideoAnydoor实现高保真视频对象插入 03.01.2025

本期的 17 篇论文如下: [00:24] 📚 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining(2.5年课堂:用于视觉-语言预训练的多模态教科书) [01:02] 🎥 VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control(VideoAnydoor:高保真视频对象插入与精确运动控制) [01:39] 🎥 VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM(VideoRefer...

2025.01.02 每日AI论文 | 自动化GUI代理轨迹构建,优化推理任务语言模型。 02.01.2025

本期的 2 篇论文如下: [00:26] 🤖 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis(OS-Genesis:通过逆向任务合成自动化GUI代理轨迹构建) [01:10] 🧠 Xmodel-2 Technical Report(Xmodel-2技术报告) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递 在小宇宙查看该单集文稿

2024.12.31 每日AI论文 | 解释性指令提升视觉任务泛化,多模态模型优化医学影像泛化。 31.12.2024

本期的 10 篇论文如下: [00:25] 🔍 Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization(解释性指令:迈向统一视觉任务理解与零样本泛化) [01:13] 🧠 On the Compositional Generalization of Multimodal LLMs for Medical Imaging(多模态大语言模型在医学影像中的组合泛化研究) [02:02] ⚙ Efficiently Serving LLM Reasoning Programs with Certaindex(高效服务LLM推理程...

2024.12.30 每日AI论文 | 华佗GPT-o1提升医学推理,Orient Anything精准估计物体方向。 30.12.2024

本期的 8 篇论文如下: [00:30] 🧠 HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs(华佗GPT-o1:迈向医学复杂推理的大语言模型) [01:16] 🧭 Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models(定向万物:从渲染3D模型中学习鲁棒的物体方向估计) [02:03] 🔍 Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment(任务...

【周末特辑】12月第4周最火AI论文 | 鲁棒微调提升大模型抗噪能力,并行生成加速视觉模型效率。 28.12.2024

本期的 5 篇论文如下: [00:37] TOP1(🔥78) | 🛡 RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response(RobustFT:在噪声响应下的大语言模型的鲁棒监督微调) [02:57] TOP2(🔥47) | ⚡ Parallelized Autoregressive Visual Generation(并行自回归视觉生成) [05:16] TOP3(🔥38) | 🔄 B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners(B-STaR:监控...

2024.12.27 每日AI论文 | YuLan-Mini提升数据效率,Gist Token优化上下文压缩。 27.12.2024

本期的 4 篇论文如下: [00:26] 🧠 YuLan-Mini: An Open Data-efficient Language Model(YuLan-Mini:一个开放的数据高效语言模型) [01:05] 🔍 A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression(银弹还是全注意力妥协?基于Gist Token的上下文压缩全面研究) [01:49] 🤖 Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequ...

2024.12.26 每日AI论文 | Token预算优化推理,Video-Panda提升视频处理效率。 26.12.2024

本期的 4 篇论文如下: [00:27] 💡 Token-Budget-Aware LLM Reasoning(基于Token预算的大语言模型推理) [01:07] 🎥 Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models(Video-Panda:无编码器视频语言模型的高效参数对齐方法) [01:49] 🧠 Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search(Mulberry:通过集体蒙特卡洛树搜索赋予ML...

2024.12.25 每日AI论文 | 提升三维场景理解,填补深度信息缺失。 25.12.2024

本期的 9 篇论文如下: [00:26] 🧠 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding(3DGraphLLM:结合语义图与大型语言模型进行三维场景理解) [01:11] 🖼 DepthLab: From Partial to Complete(DepthLab:从部分到完整) [01:54] 📊 Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization(傅里叶位置嵌入:增强注意力机制的周期性扩展...

2024.12.24 每日AI论文 | 探索与利用平衡,噪声数据处理提升。 24.12.2024

本期的 16 篇论文如下: [00:24] 🔄 B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners(B-STaR:监控和平衡自学习推理器中的探索与利用) [01:04] 🛡 RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response(RobustFT:在噪声响应下的大语言模型的鲁棒监督微调) [01:43] 🧠 Diving into Self-Evolving Training for Multimodal Reasoning(深入自进化...

2024.12.23 每日AI论文 | 加速视觉生成,优化多步推理 23.12.2024

本期的 10 篇论文如下: [00:22] ⚡ Parallelized Autoregressive Visual Generation(并行自回归视觉生成) [01:05] 🧠 Offline Reinforcement Learning for LLM Multi-Step Reasoning(基于离线强化学习的大语言模型多步推理) [01:43] 🔑 SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation(SCOPE:优化长上下文生成中的键值缓存压缩) [02:30] 🚀 CLEAR: Conv-Like Linearization Revs Pre-Trained D...

【周末特辑】12月第3周最火AI论文 | Qwen2.5提升LLMs性能,阿波罗优化视频理解。 21.12.2024

本期的 5 篇论文如下: [00:40] TOP1(🔥252) | 🤖 Qwen2.5 Technical Report(Qwen2.5技术报告) [02:31] TOP2(🔥127) | 🎥 Apollo: An Exploration of Video Understanding in Large Multimodal Models(阿波罗:大型多模态模型中的视频理解探索) [04:30] TOP3(🔥86) | 🚀 Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference(更智能、更...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.