duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

【周末特辑】3月第2周最火AI论文 | 稀疏自编码器提升文本检测,自动化ICD编码提高医疗效率。 15.03.2025

本期的 5 篇论文如下: [00:44] TOP1(🔥208) | 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [03:15] TOP2(🔥122) | 🇷 RuCCoD: Towards Automated ICD Coding in Russian(RuCCoD:面向俄语自动化的ICD编码研究) [05:35] TOP3(🔥104) | 🌐 Unified Reward Model for Multimodal Understanding and Generation(多模态理解和生成的统一奖励模型...

2025.03.14 | CoSTA*优化多轮编辑效率,无声品牌攻击揭示扩散模型脆弱性。 14.03.2025

本期的 15 篇论文如下: [00:25] 🖼 CoSTA$\ast$: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing(CoSTA*:面向多轮图像编辑的成本敏感工具路径代理) [01:03] 🎭 Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models(无声品牌攻击:针对文本到图像扩散模型的无触发数据投毒攻击) [01:45] 🌍 World Modeling Makes a Better Planner: Dual Preference Optimization fo...

2025.03.13 | 降低视频扩散模型计算需求,提升多视角视频生成质量。 13.03.2025

本期的 15 篇论文如下: [00:20] 🎥 TPDiff: Temporal Pyramid Video Diffusion Model(TPDiff:时间金字塔视频扩散模型) [00:58] 🎥 Reangle-A-Video: 4D Video Generation as Video-to-Video Translation(Reangle-A-Video:将4D视频生成作为视频到视频的转换) [01:42] 🧠 Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models(块扩散:在自回归与扩散语言模型之间插值) [02:18] 🎯 Reward...

2025.03.12 | 东南亚数据集创新构建,大模态模型推理能力显著提升 12.03.2025

本期的 15 篇论文如下: [00:23] 🌏 Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia(众包、爬取还是生成?创建东南亚视觉语言数据集SEA-VL) [01:04] 🧠 LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL(LMM-R1:通过两阶段基于规则的强化学习赋予3B参数大模态模型强大的推理能力) [01:43] 🎵 YuE: Scaling Ope...

2025.03.11 | 稀疏自编码器提升文本检测,SEAP优化语言模型效率 11.03.2025

本期的 11 篇论文如下: [00:25] 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [01:00] 🧠 SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models(SEAP:无训练的稀疏专家激活剪枝解锁大规模语言模型的脑力) [01:43] 🧠 MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scal...

2025.03.10 | 多模态任务新框架,俄语ICD编码提升。 10.03.2025

本期的 20 篇论文如下: [00:19] 🌐 Unified Reward Model for Multimodal Understanding and Generation(多模态理解和生成的统一奖励模型) [01:04] 🇷 RuCCoD: Towards Automated ICD Coding in Russian(RuCCoD:面向俄语自动化的ICD编码研究) [01:41] 🌍 EuroBERT: Scaling Multilingual Encoders for European Languages(EuroBERT:扩展欧洲语言的多语言编码器) [02:28] 🗣 S2S-Arena, Evaluating Speech2Speech Protocols...

【周末特辑】3月第1周最火AI论文 | 多模态模型音频安全评估,集成工具提升推理效率。 08.03.2025

本期的 5 篇论文如下: [00:35] TOP1(🔥64) | 🧠 Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs(Phi-4-Mini技术报告:通过LoRA混合的多模态语言模型实现紧凑且强大的性能) [02:30] TOP2(🔥58) | 🛠 START: Self-taught Reasoner with Tools(自教工具集成推理器) [04:36] TOP3(🔥57) | 🧠 Visual-RFT: Visual Reinforcement Fine-Tuning(视觉强化微调:视觉强化微调) [...

2025.03.07 | 提升推理效率,AI助手优化生活。 07.03.2025

本期的 18 篇论文如下: [00:21] 🛠 START: Self-taught Reasoner with Tools(自教工具集成推理器) [01:03] 👓 EgoLife: Towards Egocentric Life Assistant(EgoLife:面向自我中心的生活助手) [01:39] 📞 LLM as a Broken Telephone: Iterative Generation Distorts Information(大型语言模型作为失真传话:迭代生成对信息的影响) [02:14] 🧠 LINGOLY-TOO: Disentangling Memorisation from Reasoning with Linguistic Templ...

2025.03.06 | 开源多语言模型Babel表现优异,多模态嵌入模型ABC提升控制能力。 06.03.2025

本期的 17 篇论文如下: [00:24] 🌍 Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers(巴别塔:服务于全球90%以上人口的开源多语言大型语言模型) [01:11] 🧠 ABC: Achieving Better Control of Multimodal Embeddings using VLMs(ABC:使用视觉语言模型实现多模态嵌入的更好控制) [01:47] 🩺 Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions(...

2025.03.05 | MPO提升LLM规划效率,Mask-DPO增强事实性对齐。 05.03.2025

本期的 18 篇论文如下: [00:21] 🚀 MPO: Boosting LLM Agents with Meta Plan Optimization(MPO:通过元计划优化提升LLM代理) [00:59] 🤖 Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs(Mask-DPO:大语言模型的可泛化细粒度事实性对齐) [01:43] 🧩 LADDER: Self-Improving LLMs Through Recursive Problem Decomposition(LADDER:通过递归问题分解实现自我改进的LLMs) [02:26] 📚 Wikipedia in the E...

2025.03.04 | 强化视觉推理,提升3D重建质量。 04.03.2025

本期的 20 篇论文如下: [00:21] 🧠 Visual-RFT: Visual Reinforcement Fine-Tuning(视觉强化微调:视觉强化微调) [01:05] 🌐 Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models(Difix3D+:通过单步扩散模型改进三维重建) [01:43] 🧠 Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs(Phi-4-Mini技术报告:通过LoRA混合的多模态语言模型实现紧...

2025.03.03 | 工程设计效率提升,推理任务成本降低。 03.03.2025

本期的 10 篇论文如下: [00:20] 🌲 DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking(深度解决方案:通过基于树的探索与双点思维提升复杂工程解决方案设计) [00:55] ✍ Chain of Draft: Thinking Faster by Writing Less(草稿链:通过减少书写提高思考速度) [01:39] 🧠 ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoni...

【月末特辑】2月最火AI论文 | 以数据为中心的小型语言模型训练;人类动画新框架。 02.03.2025

本期的 10 篇论文如下: [00:39] TOP1(🔥196) | 🤖 SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model(SmolLM2:当小型模型走向大型化——以数据为中心的小型语言模型训练) [02:32] TOP2(🔥183) | 🎥 OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models(OmniHuman-1:重新思考一阶段条件人类动画模型的扩展) [05:02] TOP3(🔥182) | 🦜 The Stochastic Par...

【周末特辑】2月第4周最火AI论文 | 标点符号影响LLM记忆,SurveyX提升问卷质量。 01.03.2025

本期的 5 篇论文如下: [00:50] TOP1(🔥152) | 🔍 LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers(LLM显微镜:揭示标点符号在Transformer上下文记忆中的隐藏作用) [03:08] TOP2(🔥89) | 📚 SurveyX: Academic Survey Automation via Large Language Models(SurveyX 基于大型语言模型的学术调查自动化系统) [05:42] TOP3(🔥65) | 🎥 VideoGrain: Modulating Space-Time Attenti...

2025.02.28 | 自我校正提升数学推理,强化学习优化医疗推理。 28.02.2025

本期的 19 篇论文如下: [00:23] 🧠 Self-rewarding correction for mathematical reasoning(自我奖励的数学推理校正) [01:03] 🧠 MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning(MedVLM-R1:通过强化学习激励视觉语言模型的医疗推理能力) [01:53] 🧠 R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts(R2-T2:测试时重路由在多模态...

2025.02.27 | Kanana提升韩英双语效率,GHOST 2.0实现高保真头部转移。 27.02.2025

本期的 18 篇论文如下: [00:23] 🌐 Kanana: Compute-efficient Bilingual Language Models(Kanana:计算高效的双语语言模型) [00:54] 👤 GHOST 2.0: generative high-fidelity one shot transfer of heads(GHOST 2.0:生成高保真一次性头部转移) [01:43] 🎥 TheoremExplainAgent: Towards Multimodal Explanations for LLM Theorem Understanding(定理解释代理:面向大语言模型定理理解的多模态解释) [02:21] 🤖 Agentic Re...

2025.02.26 | OmniAlign-V提升多模态模型对齐,SpargeAttn加速注意力计算 26.02.2025

本期的 14 篇论文如下: [00:23] 🤖 OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference(OmniAlign-V:迈向多模态大语言模型与人类偏好增强对齐) [01:06] ⚡ SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference(SpargeAttn:准确稀疏注意力加速任意模型推理) [01:53] 🖼 KV-Edit: Training-Free Image Editing for Precise Background Preservation(KV-编辑:无需训练的图像编辑...

2025.02.25 | 长上下文优化创新,视觉扩散高效通用。 25.02.2025

本期的 20 篇论文如下: [00:27] 📖 Thus Spake Long-Context Large Language Model(长上下文大语言模型如是说) [01:09] 🌈 DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks(用于视觉感知任务的通用扩散模型) [01:48] 🚀 Slamming: Training a Speech Language Model on One GPU in a Day(撞击:在一天内使用单个GPU训练语音语言模型) [02:32] 🎥 VideoGrain: Modulating Space-Time Attention for Mu...

2025.02.24 | 高效学术调查生成,标点符号关键作用 24.02.2025

本期的 20 篇论文如下: [00:23] 📚 SurveyX: Academic Survey Automation via Large Language Models(基于大型语言模型的学术调查自动化) [01:10] 🔍 LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers(LLM显微镜:揭示标点符号在Transformer上下文记忆中的隐藏作用) [01:50] 🚗 MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction(MaskGWM:结合视...

【周末特辑】2月第3周最火AI论文 | MLGym推动AI代理评估,Qwen2.5-VL提升多模态表现。 22.02.2025

本期的 5 篇论文如下: [00:42] TOP1(🔥138) | 🧠 MLGym: A New Framework and Benchmark for Advancing AI Research Agents(MLGym:推进AI研究代理的新框架与基准) [03:30] TOP2(🔥131) | 🌐 Qwen2.5-VL Technical Report(Qwen2.5-VL 技术报告) [06:56] TOP3(🔥130) | ⚡ Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention(原生稀疏注意力:硬件对齐与原生可训练的稀疏注意力) [09:20] T...

2025.02.21 | AI代理评估新框架,LLM学科表现差异显著。 21.02.2025

本期的 20 篇论文如下: [00:26] 🧠 MLGym: A New Framework and Benchmark for Advancing AI Research Agents(MLGym:推进AI研究代理的新框架与基准) [01:18] 📚 SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines(SuperGPQA:扩展LLM评估至285个研究生学科) [02:04] 🌐 SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features(SigLIP...

2025.02.20 | 提升视觉感知,强化自动驾驶安全。 20.02.2025

本期的 20 篇论文如下: [00:24] 🌐 Qwen2.5-VL Technical Report(Qwen2.5-VL 技术报告) [01:10] 🚗 RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning(RAD:基于大规模3DGS强化学习的端到端驾驶策略训练) [01:50] 🎶 SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation(SongGen:用于文本到歌曲生成的单阶段自回归Transformer) [02:38] 🧠 Mo...

2025.02.19 | 数据高效语音处理,嵌入空间压缩创新。 19.02.2025

本期的 20 篇论文如下: [00:25] 🎙 Soundwave: Less is More for Speech-Text Alignment in LLMs(声波:减少数据需求,优化语音与文本对齐在LLMs中的应用) [01:05] 🔍 Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity(将1568个Token压缩到一个向量并再次解压:探索嵌入空间容量的极限) [01:48] 🌊 Continuous Diffusion Model for Language Modeling(连续扩散...

2025.02.18 | 稀疏注意力提升效率,机器人起身策略优化。 18.02.2025

本期的 29 篇论文如下: [00:23] ⚡ Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention(原生稀疏注意力:硬件对齐与原生可训练的稀疏注意力) [01:10] 🤖 Learning Getting-Up Policies for Real-World Humanoid Robots(学习真实世界人形机器人起身策略) [01:55] 🧠 ReLearn: Unlearning via Learning for Large Language Models(ReLearn:通过学习实现大型语言模型的遗忘) [02:35] 💻 SWE...

2025.02.17 | RAS加速扩散变换器,视频生成提升质量 17.02.2025

本期的 21 篇论文如下: [00:22] 🌐 Region-Adaptive Sampling for Diffusion Transformers(区域自适应采样扩散变换器) [01:05] 🎥 Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model(步进视频生成技术报告:视频基础模型的实践、挑战与未来) [01:48] 🌊 Large Language Diffusion Models(大规模语言扩散模型) [02:31] 🧠 ZeroBench: An Impossible Visual Benchmark for C...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.