duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2025.07.07 | GPT-4o在语义任务中表现良好;潜在空间模拟精度高。 07.07.2025 3:30
本期的 4 篇论文如下: [00:27] 🖼 How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks(GPT-4o的视觉理解能力如何?在标准计算机视觉任务上评估多模态基础模型) [01:09] 🌌 Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation(迷失于潜在空间:用于物理模拟的潜在扩散模型实证研究) [01:45] 🇮 Eka-Eval : A Compre...
【周末特辑】7月第1周最火AI论文 | 多模态推理模型提升;短视频理解领先。 06.07.2025 12:49
本期的 5 篇论文如下: [00:35] TOP1(🔥165) | 🧠 GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning(GLM-4.1V-Thinking:基于可扩展强化学习的通用多模态推理) [02:53] TOP2(🔥108) | 🎬 Kwai Keye-VL Technical Report(Kwai Keye-VL 技术报告) [05:17] TOP3(🔥67) | 🎨 LongAnimation: Long Animation Generation with Dynamic Global-Local Memory(LongAnimation:基于动...
【月末特辑】6月最火AI论文 | LLM通过自我反思提升性能;MiniMax-M1高效扩展测试计算。 05.07.2025 24:05
本期的 10 篇论文如下: [00:37] TOP1(🔥258) | 💡 Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning(反思、重试、奖励:通过强化学习实现LLM的自我提升) [02:51] TOP2(🔥249) | 💡 MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention(MiniMax-M1:利用闪电注意力高效扩展测试时计算) [05:24] TOP3(🔥240) | 🤖 Reinforcement Pre-Training(强化预训练) [07:54] TOP4(🔥1...
2025.07.04 | WebSailor提升LLM推理能力;LangScene-X优化3D场景重建。 04.07.2025 11:31
本期的 15 篇论文如下: [00:22] 🧭 WebSailor: Navigating Super-human Reasoning for Web Agent(WebSailor:为Web Agent导航超人推理) [00:59] 🖼 LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion(LangScene-X:通过TriMap视频扩散重建可泛化的3D语言嵌入场景) [01:44] 🧬 IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure P...
2025.07.03 | 多模态模型提升短视频理解;动画生成保持颜色一致。 04.07.2025 6:58
本期的 9 篇论文如下: [00:21] 🎬 Kwai Keye-VL Technical Report(Kwai Keye-VL 技术报告) [01:02] 🎨 LongAnimation: Long Animation Generation with Dynamic Global-Local Memory(LongAnimation:基于动态全局-局部记忆的长期动画生成) [01:50] 👁 Depth Anything at Any Condition(任意条件下的深度感知) [02:28] 🤖 A Survey on Vision-Language-Action Models: An Action Tokenization Perspective(视觉-语言-动作模...
2025.07.02 | 多模态推理提升;双向嵌入优化 02.07.2025 8:49
本期的 12 篇论文如下: [00:23] 💡 GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning(GLM-4.1V-Thinking:基于可扩展强化学习的通用多模态推理) [01:00] 🖼 MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings(MoCa:模态感知持续预训练提升双向多模态嵌入效果) [01:35] 🔬 SciArena: An Open Evaluation Platform for Foundati...
2025.07.01 | 多模态生成领先;视频扩散效率提升 01.07.2025 11:05
本期的 15 篇论文如下: [00:21] 🖼 Ovis-U1 Technical Report(Ovis-U1 技术报告) [00:58] 🎬 VMoBA: Mixture-of-Block Attention for Video Diffusion Models(VMoBA:用于视频扩散模型的混合块注意力机制) [01:36] ✍ Calligrapher: Freestyle Text Image Customization(书法家:自由风格的文本图像定制) [02:21] 🖼 Listener-Rewarded Thinking in VLMs for Image Preferences(图像偏好:视觉语言模型中基于监听者奖励的思考...
2025.06.30 | 3D视觉编辑;视频令牌压缩 01.07.2025 10:47
本期的 14 篇论文如下: [00:26] 🎨 BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing(BlenderFusion:基于3D的视觉编辑和生成式合成) [00:59] ✂ LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs(LLaVA-Scissor:基于语义连通分量的视频LLM令牌压缩) [01:42] 🖼 XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation...
【周末特辑】6月第5周最火AI论文 | 拖拽式大模型提升效率;法线光照恢复高精度。 28.06.2025 12:25
本期的 5 篇论文如下: [00:42] TOP1(🔥107) | 🧲 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights(拖拽式大语言模型:零样本提示到权重) [02:39] TOP2(🔥80) | 💡 Light of Normals: Unified Feature Representation for Universal Photometric Stereo(法线光照:用于通用光度立体的统一特征表示) [04:59] TOP3(🔥79) | 🖼 Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding(...
2025.06.27 | 强化学习提升搜索效率;记忆增强生成逼真驾驶场景。 28.06.2025 10:58
本期的 15 篇论文如下: [00:25] 🔍 MMSearch-R1: Incentivizing LMMs to Search(MMSearch-R1:激励大型多模态模型进行搜索) [00:59] 🚗 MADrive: Memory-Augmented Driving Scene Modeling(MADrive:基于记忆增强的驾驶场景建模) [01:43] 🤖 WorldVLA: Towards Autoregressive Action World Model(WorldVLA:面向自回归动作世界模型) [02:23] 💡 Where to find Grokking in LLM Pretraining? Monitor Memorization-to-Gener...
2025.06.26 | 高质量多模态模型;4比特量化提升性能 26.06.2025 10:30
本期的 14 篇论文如下: [00:23] 🖼 ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation(ShareGPT-4o-Image:通过GPT-4o级别的图像生成能力对齐多模态模型) [01:05] 🛡 Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models(面向稳健4比特量化的异常值安全预训练大语言模型) [01:49] 🎨 Inverse-and-Edit: Effective and Fast Image Editing by Cycle Consistenc...
2025.06.25 | AnimaX提升3D非生物体动画效果;Matrix-Game优化游戏世界模型。 26.06.2025 11:10
本期的 15 篇论文如下: [00:25] 🤖 AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models(AnimaX:利用联合视频-姿态扩散模型为3D非生物体赋予动画效果) [01:11] 🎮 Matrix-Game: Interactive World Foundation Model(矩阵游戏:交互式世界基础模型) [01:50] 🧠 GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning(GRPO-CARE:一致性感知的多模态推理强化学习) [02:...
2025.06.24 | 法线光照新方法提升细节;多模态生成模型表现优异。 25.06.2025 10:44
本期的 15 篇论文如下: [00:24] 💡 Light of Normals: Unified Feature Representation for Universal Photometric Stereo(法线光照:用于通用光度立体的统一特征表示) [01:00] 🎨 OmniGen2: Exploration to Advanced Multimodal Generation(OmniGen2:迈向更高级的多模态生成探索) [01:39] ✍ LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning(LongWriter-Zero:通过强化学习掌握超长文本...
2025.06.23 | DnD降低计算开销;视觉引导提升RAG性能。 23.06.2025 9:02
本期的 12 篇论文如下: [00:23] 🧲 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights(拖拽式大语言模型:零样本提示到权重) [01:04] 🖼 Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding(视觉引导分块:增强RAG的多模态文档理解方案) [01:49] 🔀 PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models(PAROAtt...
【周末特辑】6月第4周最火AI论文 | 高效扩展推理能力;多模态金融评估基准。 21.06.2025 11:55
本期的 5 篇论文如下: [00:36] TOP1(🔥216) | 💡 MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention(MiniMax-M1:利用闪电注意力高效扩展测试时计算) [02:44] TOP2(🔥82) | 📊 MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation(MultiFinBen:一个多语言、多模态和难度感知的金融领域大语言模型评估基准) [05:32] TOP3(🔥64) | 🔬 Scientis...
2025.06.20 | 强化学习提升跨领域推理;语音情感检测基准精细化。 20.06.2025 3:45
本期的 4 篇论文如下: [00:24] 🧠 Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective(跨领域视角下重新审视强化学习在大型语言模型推理中的应用) [01:00] 🗣 EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection(EmoNet-Voice:一个用于语音情感检测的精细化、专家验证的基准) [01:55] 🎵 SonicVerse: Multi-Task Learning for Music Feature-Informe...
2025.06.19 | SEKAI数据集提升视频生成;原型推理增强LLM泛化能力。 19.06.2025 11:18
本期的 15 篇论文如下: [00:22] 🌍 Sekai: A Video Dataset towards World Exploration(Sekai:一个面向世界探索的视频数据集) [01:02] 💡 ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs(原型推理:作为大型语言模型中通用推理基础的原型) [01:43] 💡 GenRecal: Generation after Recalibration from Large to Small Vision-Language Models(GenRecal:从大型到小型视觉-语言模型的重...
2025.06.18 | MultiFinBen揭示金融模型局限;测试时计算提升LLM Agent性能。 18.06.2025 10:52
本期的 15 篇论文如下: [00:23] 📊 MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation(MultiFinBen:一个多语言、多模态和难度感知的金融领域大语言模型评估基准) [01:03] 🤖 Scaling Test-time Compute for LLM Agents(扩展LLM Agent的测试时计算) [01:38] 🎼 CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following(CMI-Bench:一个评估...
2025.06.17 | MiniMax-M1提升推理性能;多模态模型认知测试创新。 17.06.2025 10:53
本期的 15 篇论文如下: [00:22] 💡 MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention(MiniMax-M1:利用闪电注意力高效扩展测试时计算) [01:00] 🔬 Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning(科学家的首次考试:通过感知、理解和推理来探索多模态大型语言模型的认知能力) [01:47] 🧐 DeepResearch Bench: A Comprehensive Benc...
2025.06.16 | 跨模态合成新视角图像;策略依从型智能体抗攻击 17.06.2025 11:27
本期的 15 篇论文如下: [00:23] 🖼 Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation(基于跨模态注意力提炼的对齐新视角图像与几何体合成) [01:02] 🛡 Effective Red-Teaming of Policy-Adherent Agents(有效对抗策略依从型智能体) [01:39] 🔄 The Diffusion Duality(扩散二元性) [02:20] 🤖 LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?...
【周末特辑】6月第3周最火AI论文 | 强化预训练提升语言模型推理能力;多语种分类器改善问答系统可信度。 15.06.2025 12:36
本期的 5 篇论文如下: [00:43] TOP1(🔥199) | 🤖 Reinforcement Pre-Training(强化预训练) [03:06] TOP2(🔥124) | 🕰 Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA(明日依旧为真吗?多语种常青问题分类以提升可信赖的问答系统) [05:07] TOP3(🔥105) | 🧠 Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models(自信即全部:基于语言模型的...
2025.06.13 | 医学推理模型新范式;自动化构建软件工程数据集 14.06.2025 11:36
本期的 15 篇论文如下: [00:22] 🩺 ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning(ReasonMed:一个用于推进医学推理的37万多智能体生成数据集) [01:12] 🏭 SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks(SWE-Factory:你的问题解决训练数据和评估基准自动化工厂) [01:55] 🖼 Text-Aware Image Restoration with Diffusion Models(...
2025.06.12 | 自信微调提升模型表现;视频生成模型高效优化。 12.06.2025 9:44
本期的 13 篇论文如下: [00:23] 🧠 Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models(自信即全部:基于语言模型的小样本强化学习微调) [01:07] 🎬 Seedance 1.0: Exploring the Boundaries of Video Generation Models(Seedance 1.0:探索视频生成模型的边界) [01:50] 🥽 PlayerOne: Egocentric World Simulator(PlayerOne:以自我为中心的真实世界模拟器) [02:30] 🎬 Autoregressive Adversarial...
2025.06.11 | LLM存在地缘政治偏见;RuleReasoner提升推理效率。 11.06.2025 11:06
本期的 15 篇论文如下: [00:22] 🌍 Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models(LLM中的地缘政治偏见:在当代语言模型中,哪些是“好”国家,哪些是“坏”国家?) [01:09] 🤖 RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling(RuleReasoner:基于领域感知动态采样的强化规则推理) [01:48] 🖼 Autoregressive Semantic...
2025.06.10 | 强化学习改进语言模型;医学多模态模型提升推理能力。 10.06.2025 11:16
本期的 15 篇论文如下: [00:21] 🤖 Reinforcement Pre-Training(强化预训练) [01:01] 🩺 Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning(灵枢:用于统一多模态医学理解与推理的通用基础模型) [01:42] 📱 MiniCPM4: Ultra-Efficient LLMs on End Devices(MiniCPM4:终端设备上的超高效大型语言模型) [02:30] 🛡 Saffron-1: Towards an Inference Scaling Paradigm for...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.