duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
【周末特辑】1月第2周最火AI论文 | MiniMax-01扩展长上下文处理,数学推理PRM提升过程监督。 18.01.2025 11:35
本期的 5 篇论文如下: [00:35] TOP1(🔥258) | ⚡ MiniMax-01: Scaling Foundation Models with Lightning Attention(MiniMax-01:基于闪电注意力机制扩展基础模型) [02:52] TOP2(🔥77) | 📊 The Lessons of Developing Process Reward Models in Mathematical Reasoning(数学推理中过程奖励模型开发的经验教训) [05:06] TOP3(🔥66) | 🧠 Tensor Product Attention Is All You Need(张量积注意力机制是关键) [06:49] TOP4(🔥...
2025.01.17 | OmniThink提升机器写作深度与新颖性,扩散模型推理扩展提升生成质量。 18.01.2025 8:23
本期的 12 篇论文如下: [00:26] 🧠 OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking(OmniThink:通过思考扩展机器写作的知识边界) [01:06] 🔍 Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps(扩散模型推理时扩展:超越去噪步骤的扩展) [01:37] 🩺 Exploring the Inquiry-Diagnosis Relationship with Advanced Patient Simulators(探索高级患者模拟器中的问...
2025.01.16 | MMDocIR推动多模态检索标准化,CityDreamer4D创新4D城市生成模型。 16.01.2025 6:40
本期的 9 篇论文如下: [00:25] 📊 MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents(MMDocIR:长文档多模态检索的基准测试) [01:06] 🏙 CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities(CityDreamer4D:无界4D城市的组合生成模型) [01:49] 🎥 RepVideo: Rethinking Cross-Layer Representation for Video Generation(RepVideo:重新思考视频生成中的跨层表示) [02:30] 📚 Towards Be...
2025.01.15 | MiniMax-01扩展基础模型处理长上下文,填充符在T2I模型中影响图像生成。 15.01.2025 10:21
本期的 15 篇论文如下: [00:23] ⚡ MiniMax-01: Scaling Foundation Models with Lightning Attention(MiniMax-01:基于闪电注意力机制扩展基础模型) [01:04] 🖼 Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models(填充符:T2I模型中填充符的机制分析) [01:44] 🎨 MangaNinja: Line Art Colorization with Precise Reference Following(MangaNinja:基于精确参考跟随的线稿上色) [02:21] 🧬 A Multi-Mo...
2025.01.14 | 数学推理提升,内存开销减少 14.01.2025 9:01
本期的 11 篇论文如下: [00:24] 📊 The Lessons of Developing Process Reward Models in Mathematical Reasoning(数学推理中过程奖励模型开发的经验教训) [01:10] 🧠 Tensor Product Attention Is All You Need(张量积注意力机制是关键) [01:53] 🤖 $\text{Transformer}^2$: Self-adaptive LLMs(Transformer²:自适应大型语言模型) [02:34] 🎥 VideoAuteur: Towards Long Narrative Video Generation(视频导演:面向长篇...
2025.01.13 | OmniManip实现通用机器人操作,VideoRAG提升视频检索生成性能。 13.01.2025 7:01
本期的 10 篇论文如下: [00:24] 🤖 OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints(OmniManip:通过以对象为中心的交互原语作为空间约束实现通用机器人操作) [01:02] 🎥 VideoRAG: Retrieval-Augmented Generation over Video Corpus(VideoRAG:基于视频语料库的检索增强生成) [01:38] 🎥 OVO-Bench: How Far is Your Video-LLMs from Real-World Onlin...
【周末特辑】1月第1周最火AI论文 | 小型模型超越大型模型,REINFORCE++简化对齐方法 11.01.2025 12:08
本期的 5 篇论文如下: [00:39] TOP1(🔥173) | 🧠 rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking(rStar-Math:小型语言模型通过自我进化的深度思考掌握数学推理) [03:03] TOP2(🔥71) | 🚀 REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models(REINFORCE++:一种简单高效的大语言模型对齐方法) [05:17] TOP3(🔥63) | 🧠 Towards System 2 Reasoning in LLM...
2025.01.10 每日AI论文 | GAN训练简化性能提升,视频自回归预训练竞争力显著。 10.01.2025 5:25
本期的 7 篇论文如下: [00:23] 🧠 The GAN is dead; long live the GAN! A Modern GAN Baseline(GAN已死;GAN万岁!一个现代的GAN基线) [01:02] 🎥 An Empirical Study of Autoregressive Pre-training from Videos(视频自回归预训练的实证研究) [01:49] 🚗 Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives(视觉语言模型是否准备好用于自动驾驶?从可靠性...
2025.01.09 每日AI论文 | 小型模型自我进化超越GPT-3,多模态模型提升数学推理能力。 09.01.2025 7:54
本期的 11 篇论文如下: [00:25] 🧠 rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking(rStar-Math:小型语言模型通过自我进化的深度思考掌握数学推理) [01:06] 🧠 URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics(URSA:理解与验证多模态数学中的思维链推理) [01:45] 🧠 Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Ch...
2025.01.08 每日AI论文 | REINFORCE++提升大模型对齐效率,MotionBench优化视频运动理解 08.01.2025 7:58
本期的 11 篇论文如下: [00:24] 🚀 REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models(REINFORCE++:一种简单高效的大语言模型对齐方法) [01:00] 🎥 MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models(MotionBench:用于评估和改进视觉语言模型细粒度视频运动理解的基准) [01:40] 🔍 Sa2VA: Marrying SAM2 with LLaVA for Dense...
2025.01.07 每日AI论文 | STAR提升视频超分辨率时空一致性,BoostStep增强大模型数学推理能力。 07.01.2025 11:02
本期的 16 篇论文如下: [00:24] 🎥 STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution(STAR:基于文本到视频模型的空间-时间增强用于现实世界视频超分辨率) [01:06] 🧮 BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning(BoostStep:通过改进单步推理提升大语言模型的数学能力) [01:44] 🤖 Dispider: Enabling...
2025.01.06 每日AI论文 | EnerVerse提升机器人操作规划能力,VITA-1.5优化实时视觉语音交互。 06.01.2025 5:51
本期的 8 篇论文如下: [00:24] 🤖 EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation(EnerVerse:面向机器人操作的具身未来空间构想) [00:58] 🤖 VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction(VITA-1.5:迈向GPT-4o级别的实时视觉与语音交互) [01:33] 🤔 Virgo: A Preliminary Exploration on Reproducing o1-like MLLM(Virgo:关于复现o1类多模态大语言模型的初步探索...
【月末特辑】12月最火AI论文 | Qwen2.5提升大语言模型性能,阿波罗优化视频理解效率。 05.01.2025 23:51
本期的 10 篇论文如下: [00:31] TOP1(🔥335) | 🤖 Qwen2.5 Technical Report(Qwen2.5技术报告) [02:44] TOP2(🔥136) | 🎥 Apollo: An Exploration of Video Understanding in Large Multimodal Models(阿波罗:大型多模态模型中的视频理解探索) [05:01] TOP3(🔥123) | 🚀 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling(通过模型、数据和测试时扩展提升开源多...
【周末特辑】12月第5周最火AI论文 | 提升医学推理能力,自动化GUI轨迹构建。 04.01.2025 11:50
本期的 5 篇论文如下: [00:35] TOP1(🔥83) | 🧠 HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs(华佗GPT-o1:迈向医学复杂推理的大语言模型) [02:49] TOP2(🔥65) | 🤖 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis(OS-Genesis:通过逆向任务合成自动化GUI代理轨迹构建) [04:50] TOP3(🔥63) | 🎨 1.58-bit FLUX(1.58位FLUX:首个成功量化最先进文本生成图像模型的方法...
2025.01.03 每日AI论文 | 多模态教科书提升视觉语言模型性能,VideoAnydoor实现高保真视频对象插入 03.01.2025 11:27
本期的 17 篇论文如下: [00:24] 📚 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining(2.5年课堂:用于视觉-语言预训练的多模态教科书) [01:02] 🎥 VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control(VideoAnydoor:高保真视频对象插入与精确运动控制) [01:39] 🎥 VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM(VideoRefer...
2025.01.02 每日AI论文 | 自动化GUI代理轨迹构建,优化推理任务语言模型。 02.01.2025 2:21
本期的 2 篇论文如下: [00:26] 🤖 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis(OS-Genesis:通过逆向任务合成自动化GUI代理轨迹构建) [01:10] 🧠 Xmodel-2 Technical Report(Xmodel-2技术报告) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递 在小宇宙查看该单集文稿
2024.12.31 每日AI论文 | 解释性指令提升视觉任务泛化,多模态模型优化医学影像泛化。 31.12.2024 7:57
本期的 10 篇论文如下: [00:25] 🔍 Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization(解释性指令:迈向统一视觉任务理解与零样本泛化) [01:13] 🧠 On the Compositional Generalization of Multimodal LLMs for Medical Imaging(多模态大语言模型在医学影像中的组合泛化研究) [02:02] ⚙ Efficiently Serving LLM Reasoning Programs with Certaindex(高效服务LLM推理程...
2024.12.30 每日AI论文 | 华佗GPT-o1提升医学推理,Orient Anything精准估计物体方向。 30.12.2024 6:58
本期的 8 篇论文如下: [00:30] 🧠 HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs(华佗GPT-o1:迈向医学复杂推理的大语言模型) [01:16] 🧭 Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models(定向万物:从渲染3D模型中学习鲁棒的物体方向估计) [02:03] 🔍 Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment(任务...
【周末特辑】12月第4周最火AI论文 | 鲁棒微调提升大模型抗噪能力,并行生成加速视觉模型效率。 28.12.2024 12:32
本期的 5 篇论文如下: [00:37] TOP1(🔥78) | 🛡 RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response(RobustFT:在噪声响应下的大语言模型的鲁棒监督微调) [02:57] TOP2(🔥47) | ⚡ Parallelized Autoregressive Visual Generation(并行自回归视觉生成) [05:16] TOP3(🔥38) | 🔄 B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners(B-STaR:监控...
2024.12.27 每日AI论文 | YuLan-Mini提升数据效率,Gist Token优化上下文压缩。 27.12.2024 3:49
本期的 4 篇论文如下: [00:26] 🧠 YuLan-Mini: An Open Data-efficient Language Model(YuLan-Mini:一个开放的数据高效语言模型) [01:05] 🔍 A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression(银弹还是全注意力妥协?基于Gist Token的上下文压缩全面研究) [01:49] 🤖 Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequ...
2024.12.26 每日AI论文 | Token预算优化推理,Video-Panda提升视频处理效率。 26.12.2024 3:49
本期的 4 篇论文如下: [00:27] 💡 Token-Budget-Aware LLM Reasoning(基于Token预算的大语言模型推理) [01:07] 🎥 Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models(Video-Panda:无编码器视频语言模型的高效参数对齐方法) [01:49] 🧠 Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search(Mulberry:通过集体蒙特卡洛树搜索赋予ML...
2024.12.25 每日AI论文 | 提升三维场景理解,填补深度信息缺失。 25.12.2024 7:09
本期的 9 篇论文如下: [00:26] 🧠 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding(3DGraphLLM:结合语义图与大型语言模型进行三维场景理解) [01:11] 🖼 DepthLab: From Partial to Complete(DepthLab:从部分到完整) [01:54] 📊 Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization(傅里叶位置嵌入:增强注意力机制的周期性扩展...
2024.12.24 每日AI论文 | 探索与利用平衡,噪声数据处理提升。 24.12.2024 12:11
本期的 16 篇论文如下: [00:24] 🔄 B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners(B-STaR:监控和平衡自学习推理器中的探索与利用) [01:04] 🛡 RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response(RobustFT:在噪声响应下的大语言模型的鲁棒监督微调) [01:43] 🧠 Diving into Self-Evolving Training for Multimodal Reasoning(深入自进化...
2024.12.23 每日AI论文 | 加速视觉生成,优化多步推理 23.12.2024 8:23
本期的 10 篇论文如下: [00:22] ⚡ Parallelized Autoregressive Visual Generation(并行自回归视觉生成) [01:05] 🧠 Offline Reinforcement Learning for LLM Multi-Step Reasoning(基于离线强化学习的大语言模型多步推理) [01:43] 🔑 SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation(SCOPE:优化长上下文生成中的键值缓存压缩) [02:30] 🚀 CLEAR: Conv-Like Linearization Revs Pre-Trained D...
【周末特辑】12月第3周最火AI论文 | Qwen2.5提升LLMs性能,阿波罗优化视频理解。 21.12.2024 11:29
本期的 5 篇论文如下: [00:40] TOP1(🔥252) | 🤖 Qwen2.5 Technical Report(Qwen2.5技术报告) [02:31] TOP2(🔥127) | 🎥 Apollo: An Exploration of Video Understanding in Large Multimodal Models(阿波罗:大型多模态模型中的视频理解探索) [04:30] TOP3(🔥86) | 🚀 Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference(更智能、更...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.