duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
【周末特辑】3月第2周最火AI论文 | 稀疏自编码器提升文本检测,自动化ICD编码提高医疗效率。 15.03.2025 12:50
本期的 5 篇论文如下: [00:44] TOP1(🔥208) | 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [03:15] TOP2(🔥122) | 🇷 RuCCoD: Towards Automated ICD Coding in Russian(RuCCoD:面向俄语自动化的ICD编码研究) [05:35] TOP3(🔥104) | 🌐 Unified Reward Model for Multimodal Understanding and Generation(多模态理解和生成的统一奖励模型...
2025.03.14 | CoSTA*优化多轮编辑效率,无声品牌攻击揭示扩散模型脆弱性。 14.03.2025 11:10
本期的 15 篇论文如下: [00:25] 🖼 CoSTA$\ast$: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing(CoSTA*:面向多轮图像编辑的成本敏感工具路径代理) [01:03] 🎭 Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models(无声品牌攻击:针对文本到图像扩散模型的无触发数据投毒攻击) [01:45] 🌍 World Modeling Makes a Better Planner: Dual Preference Optimization fo...
2025.03.13 | 降低视频扩散模型计算需求,提升多视角视频生成质量。 13.03.2025 10:50
本期的 15 篇论文如下: [00:20] 🎥 TPDiff: Temporal Pyramid Video Diffusion Model(TPDiff:时间金字塔视频扩散模型) [00:58] 🎥 Reangle-A-Video: 4D Video Generation as Video-to-Video Translation(Reangle-A-Video:将4D视频生成作为视频到视频的转换) [01:42] 🧠 Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models(块扩散:在自回归与扩散语言模型之间插值) [02:18] 🎯 Reward...
2025.03.12 | 东南亚数据集创新构建,大模态模型推理能力显著提升 12.03.2025 11:01
本期的 15 篇论文如下: [00:23] 🌏 Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia(众包、爬取还是生成?创建东南亚视觉语言数据集SEA-VL) [01:04] 🧠 LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL(LMM-R1:通过两阶段基于规则的强化学习赋予3B参数大模态模型强大的推理能力) [01:43] 🎵 YuE: Scaling Ope...
2025.03.11 | 稀疏自编码器提升文本检测,SEAP优化语言模型效率 11.03.2025 8:12
本期的 11 篇论文如下: [00:25] 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [01:00] 🧠 SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models(SEAP:无训练的稀疏专家激活剪枝解锁大规模语言模型的脑力) [01:43] 🧠 MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scal...
2025.03.10 | 多模态任务新框架,俄语ICD编码提升。 10.03.2025 14:47
本期的 20 篇论文如下: [00:19] 🌐 Unified Reward Model for Multimodal Understanding and Generation(多模态理解和生成的统一奖励模型) [01:04] 🇷 RuCCoD: Towards Automated ICD Coding in Russian(RuCCoD:面向俄语自动化的ICD编码研究) [01:41] 🌍 EuroBERT: Scaling Multilingual Encoders for European Languages(EuroBERT:扩展欧洲语言的多语言编码器) [02:28] 🗣 S2S-Arena, Evaluating Speech2Speech Protocols...
【周末特辑】3月第1周最火AI论文 | 多模态模型音频安全评估,集成工具提升推理效率。 08.03.2025 11:42
本期的 5 篇论文如下: [00:35] TOP1(🔥64) | 🧠 Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs(Phi-4-Mini技术报告:通过LoRA混合的多模态语言模型实现紧凑且强大的性能) [02:30] TOP2(🔥58) | 🛠 START: Self-taught Reasoner with Tools(自教工具集成推理器) [04:36] TOP3(🔥57) | 🧠 Visual-RFT: Visual Reinforcement Fine-Tuning(视觉强化微调:视觉强化微调) [...
2025.03.07 | 提升推理效率,AI助手优化生活。 07.03.2025 13:01
本期的 18 篇论文如下: [00:21] 🛠 START: Self-taught Reasoner with Tools(自教工具集成推理器) [01:03] 👓 EgoLife: Towards Egocentric Life Assistant(EgoLife:面向自我中心的生活助手) [01:39] 📞 LLM as a Broken Telephone: Iterative Generation Distorts Information(大型语言模型作为失真传话:迭代生成对信息的影响) [02:14] 🧠 LINGOLY-TOO: Disentangling Memorisation from Reasoning with Linguistic Templ...
2025.03.06 | 开源多语言模型Babel表现优异,多模态嵌入模型ABC提升控制能力。 06.03.2025 12:27
本期的 17 篇论文如下: [00:24] 🌍 Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers(巴别塔:服务于全球90%以上人口的开源多语言大型语言模型) [01:11] 🧠 ABC: Achieving Better Control of Multimodal Embeddings using VLMs(ABC:使用视觉语言模型实现多模态嵌入的更好控制) [01:47] 🩺 Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions(...
2025.03.05 | MPO提升LLM规划效率,Mask-DPO增强事实性对齐。 05.03.2025 13:23
本期的 18 篇论文如下: [00:21] 🚀 MPO: Boosting LLM Agents with Meta Plan Optimization(MPO:通过元计划优化提升LLM代理) [00:59] 🤖 Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs(Mask-DPO:大语言模型的可泛化细粒度事实性对齐) [01:43] 🧩 LADDER: Self-Improving LLMs Through Recursive Problem Decomposition(LADDER:通过递归问题分解实现自我改进的LLMs) [02:26] 📚 Wikipedia in the E...
2025.03.04 | 强化视觉推理,提升3D重建质量。 04.03.2025 14:17
本期的 20 篇论文如下: [00:21] 🧠 Visual-RFT: Visual Reinforcement Fine-Tuning(视觉强化微调:视觉强化微调) [01:05] 🌐 Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models(Difix3D+:通过单步扩散模型改进三维重建) [01:43] 🧠 Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs(Phi-4-Mini技术报告:通过LoRA混合的多模态语言模型实现紧...
2025.03.03 | 工程设计效率提升,推理任务成本降低。 03.03.2025 7:33
本期的 10 篇论文如下: [00:20] 🌲 DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking(深度解决方案:通过基于树的探索与双点思维提升复杂工程解决方案设计) [00:55] ✍ Chain of Draft: Thinking Faster by Writing Less(草稿链:通过减少书写提高思考速度) [01:39] 🧠 ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoni...
【月末特辑】2月最火AI论文 | 以数据为中心的小型语言模型训练;人类动画新框架。 02.03.2025 23:14
本期的 10 篇论文如下: [00:39] TOP1(🔥196) | 🤖 SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model(SmolLM2:当小型模型走向大型化——以数据为中心的小型语言模型训练) [02:32] TOP2(🔥183) | 🎥 OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models(OmniHuman-1:重新思考一阶段条件人类动画模型的扩展) [05:02] TOP3(🔥182) | 🦜 The Stochastic Par...
【周末特辑】2月第4周最火AI论文 | 标点符号影响LLM记忆,SurveyX提升问卷质量。 01.03.2025 12:13
本期的 5 篇论文如下: [00:50] TOP1(🔥152) | 🔍 LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers(LLM显微镜:揭示标点符号在Transformer上下文记忆中的隐藏作用) [03:08] TOP2(🔥89) | 📚 SurveyX: Academic Survey Automation via Large Language Models(SurveyX 基于大型语言模型的学术调查自动化系统) [05:42] TOP3(🔥65) | 🎥 VideoGrain: Modulating Space-Time Attenti...
2025.02.28 | 自我校正提升数学推理,强化学习优化医疗推理。 28.02.2025 13:50
本期的 19 篇论文如下: [00:23] 🧠 Self-rewarding correction for mathematical reasoning(自我奖励的数学推理校正) [01:03] 🧠 MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning(MedVLM-R1:通过强化学习激励视觉语言模型的医疗推理能力) [01:53] 🧠 R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts(R2-T2:测试时重路由在多模态...
2025.02.27 | Kanana提升韩英双语效率,GHOST 2.0实现高保真头部转移。 27.02.2025 13:12
本期的 18 篇论文如下: [00:23] 🌐 Kanana: Compute-efficient Bilingual Language Models(Kanana:计算高效的双语语言模型) [00:54] 👤 GHOST 2.0: generative high-fidelity one shot transfer of heads(GHOST 2.0:生成高保真一次性头部转移) [01:43] 🎥 TheoremExplainAgent: Towards Multimodal Explanations for LLM Theorem Understanding(定理解释代理:面向大语言模型定理理解的多模态解释) [02:21] 🤖 Agentic Re...
2025.02.26 | OmniAlign-V提升多模态模型对齐,SpargeAttn加速注意力计算 26.02.2025 10:26
本期的 14 篇论文如下: [00:23] 🤖 OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference(OmniAlign-V:迈向多模态大语言模型与人类偏好增强对齐) [01:06] ⚡ SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference(SpargeAttn:准确稀疏注意力加速任意模型推理) [01:53] 🖼 KV-Edit: Training-Free Image Editing for Precise Background Preservation(KV-编辑:无需训练的图像编辑...
2025.02.25 | 长上下文优化创新,视觉扩散高效通用。 25.02.2025 14:37
本期的 20 篇论文如下: [00:27] 📖 Thus Spake Long-Context Large Language Model(长上下文大语言模型如是说) [01:09] 🌈 DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks(用于视觉感知任务的通用扩散模型) [01:48] 🚀 Slamming: Training a Speech Language Model on One GPU in a Day(撞击:在一天内使用单个GPU训练语音语言模型) [02:32] 🎥 VideoGrain: Modulating Space-Time Attention for Mu...
2025.02.24 | 高效学术调查生成,标点符号关键作用 24.02.2025 15:04
本期的 20 篇论文如下: [00:23] 📚 SurveyX: Academic Survey Automation via Large Language Models(基于大型语言模型的学术调查自动化) [01:10] 🔍 LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers(LLM显微镜:揭示标点符号在Transformer上下文记忆中的隐藏作用) [01:50] 🚗 MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction(MaskGWM:结合视...
【周末特辑】2月第3周最火AI论文 | MLGym推动AI代理评估,Qwen2.5-VL提升多模态表现。 22.02.2025 14:53
本期的 5 篇论文如下: [00:42] TOP1(🔥138) | 🧠 MLGym: A New Framework and Benchmark for Advancing AI Research Agents(MLGym:推进AI研究代理的新框架与基准) [03:30] TOP2(🔥131) | 🌐 Qwen2.5-VL Technical Report(Qwen2.5-VL 技术报告) [06:56] TOP3(🔥130) | ⚡ Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention(原生稀疏注意力:硬件对齐与原生可训练的稀疏注意力) [09:20] T...
2025.02.21 | AI代理评估新框架,LLM学科表现差异显著。 21.02.2025 18:02
本期的 20 篇论文如下: [00:26] 🧠 MLGym: A New Framework and Benchmark for Advancing AI Research Agents(MLGym:推进AI研究代理的新框架与基准) [01:18] 📚 SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines(SuperGPQA:扩展LLM评估至285个研究生学科) [02:04] 🌐 SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features(SigLIP...
2025.02.20 | 提升视觉感知,强化自动驾驶安全。 20.02.2025 15:12
本期的 20 篇论文如下: [00:24] 🌐 Qwen2.5-VL Technical Report(Qwen2.5-VL 技术报告) [01:10] 🚗 RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning(RAD:基于大规模3DGS强化学习的端到端驾驶策略训练) [01:50] 🎶 SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation(SongGen:用于文本到歌曲生成的单阶段自回归Transformer) [02:38] 🧠 Mo...
2025.02.19 | 数据高效语音处理,嵌入空间压缩创新。 19.02.2025 14:35
本期的 20 篇论文如下: [00:25] 🎙 Soundwave: Less is More for Speech-Text Alignment in LLMs(声波:减少数据需求,优化语音与文本对齐在LLMs中的应用) [01:05] 🔍 Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity(将1568个Token压缩到一个向量并再次解压:探索嵌入空间容量的极限) [01:48] 🌊 Continuous Diffusion Model for Language Modeling(连续扩散...
2025.02.18 | 稀疏注意力提升效率,机器人起身策略优化。 18.02.2025 21:21
本期的 29 篇论文如下: [00:23] ⚡ Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention(原生稀疏注意力:硬件对齐与原生可训练的稀疏注意力) [01:10] 🤖 Learning Getting-Up Policies for Real-World Humanoid Robots(学习真实世界人形机器人起身策略) [01:55] 🧠 ReLearn: Unlearning via Learning for Large Language Models(ReLearn:通过学习实现大型语言模型的遗忘) [02:35] 💻 SWE...
2025.02.17 | RAS加速扩散变换器,视频生成提升质量 17.02.2025 15:43
本期的 21 篇论文如下: [00:22] 🌐 Region-Adaptive Sampling for Diffusion Transformers(区域自适应采样扩散变换器) [01:05] 🎥 Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model(步进视频生成技术报告:视频基础模型的实践、挑战与未来) [01:48] 🌊 Large Language Diffusion Models(大规模语言扩散模型) [02:31] 🧠 ZeroBench: An Impossible Visual Benchmark for C...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.