duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2025.08.04 | 扩散语言模型变长去噪,高效省资源;PixNerd图像扩散,高效高质量。 05.08.2025

本期的 11 篇论文如下: [00:22] 🔄 Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models(超越固定长度:扩散大语言模型的可变长度去噪) [00:44] 🎨 PixNerd: Pixel Neural Field Diffusion(PixNerd:像素神经场扩散) [01:11] 💡 SWE-Exp: Experience-Driven Software Issue Resolution(SWE-Exp:经验驱动的软件问题解决) [01:38] 🔍 Multimodal Referring Segmentation: A Survey(多模态指代表...

【月末特辑】7月最火AI论文 | GSPO稳训练;序列级裁剪降方差;上下文工程综述,动态拼装信息流 04.08.2025

本期的 10 篇论文如下: [00:30] TOP1(🔥257) | 🚀 Group Sequence Policy Optimization(组序列策略优化) [02:21] TOP2(🔥227) | 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [03:33] TOP3(🔥207) | 🧠 GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning(GLM-4.1V-Thinking:基于可扩展强化学习的通用多模态推理) [05:02] T...

【周末特辑】8月第1周最火AI论文 | ARPO用高熵分叉省预算;混元世界一句话生成可编辑3D场景 03.08.2025

本期的 5 篇论文如下: [00:32] TOP1(🔥114) | 🤖 Agentic Reinforced Policy Optimization(智能体强化策略优化) [02:17] TOP2(🔥94) | 🌍 HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels(混元世界 1.0:从文字或像素生成沉浸式、可探索、可交互的3D世界) [05:04] TOP3(🔥76) | 🏆 Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving(Seed-Prover...

2025.08.01 | Seed-Prover融合LLM解决IMO数学题;Phi-Ground提升GUI感知精度。 01.08.2025

本期的 15 篇论文如下: [00:22] 🏆 Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving(Seed-Prover:自动化定理证明的深度与广度推理) [01:04] 🎯 Phi-Ground Tech Report: Advancing Perception in GUI Grounding(Phi-Ground 技术报告:提升 GUI 接地感知能力) [01:30] 🤔 C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations(C3:探索复杂对话挑战...

2025.07.31 | ScreenCoder自动化UI转代码;Falcon-H1混合架构,提升长序列效率。 01.08.2025

本期的 9 篇论文如下: [00:22] 💻 ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents(ScreenCoder:模块化多模态智能体赋能前端视觉代码生成) [01:02] 🚀 Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance(Falcon-H1:重塑效率与性能的混合架构语言模型系列) [01:33] 💥 BANG: Dividing 3D Assets via Generative Explod...

2025.07.30 | 混元世界从文字像素生成沉浸3D世界;X-Omni用强化学习提升图像生成质量。 31.07.2025

本期的 8 篇论文如下: [00:23] 🌍 HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels(混元世界 1.0:从文字或像素生成沉浸式、可探索、可交互的3D世界) [00:56] ✨ X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again(X-Omni:强化学习让离散自回归图像生成模型再展辉煌) [01:59] 🚀 CUDA-L1: Improving CUDA Optimizat...

2025.07.29 | ARPO提升LLM工具交互性能;ARC-Hunyuan-Video-7B深耕短视频理解。 30.07.2025

本期的 15 篇论文如下: [00:23] 🤖 Agentic Reinforced Policy Optimization(智能体强化策略优化) [00:55] 🧠 ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts(ARC-Hunyuan-Video-7B:真实世界短视频的结构化理解) [01:35] 🚀 Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning(Rep-MTL:释放表示层任务显著性在多任务学习中的力量) [02:03] 🌐 R...

2025.07.28 | GPTQ揭示为Babai算法,保障精度;TTD-DR以扩散模型生成高质量研究报告。 29.07.2025

本期的 5 篇论文如下: [00:25] 💡 The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm(LLM 量化的几何学:GPTQ 作为 Babai 最近平面算法) [00:52] ✨ Deep Researcher with Test-Time Diffusion(基于测试时扩散的深度研究智能体) [01:40] 🔧 Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement(规范自校正:通过测试时细化缓解上下文奖励破解) [...

【周末特辑】7月第4周最火AI论文 | GUI-G2:高斯奖励提升GUI定位;MiroMind-M1:开源数学推理LLM 26.07.2025

本期的 5 篇论文如下: [00:36] TOP1(🔥118) | 🎯 GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding(GUI-G$^2$: 基于高斯奖励模型的GUI定位) [02:14] TOP2(🔥108) | 🧮 MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization(MiroMind-M1:通过上下文感知多阶段策略优化实现数学推理的开源进展) [05:19] TOP3(🔥96) | ♾ Beyond Context Limits: Subco...

2025.07.25 | GSPO解决大模型训练崩溃;MUR提升LLM推理效率。 26.07.2025

本期的 15 篇论文如下: [00:24] 🚀 Group Sequence Policy Optimization(组序列策略优化) [00:53] 🧠 MUR: Momentum Uncertainty guided Reasoning for Large Language Models(MUR:面向大型语言模型的动量不确定性引导推理) [01:30] 🧠 LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization(LAPO:内化推理效率的长度自适应策略优化) [02:09] 🎬 Captain Cinema: Towards Short Movie Gener...

2025.07.24 | MLLMs视觉感知仍不足;Yume模型可生成交互虚拟世界。 25.07.2025

本期的 9 篇论文如下: [00:23] 👁 Pixels, Patterns, but No Poetry: To See The World like Humans(像素、模式,却无诗意:像人类一样感知世界) [00:56] 🌌 Yume: An Interactive World Generation Model(Yume:交互式世界生成模型) [01:29] ✨ DesignLab: Designing Slides Through Iterative Detection and Correction(DesignLab:通过迭代检测与修正进行幻灯片设计) [02:14] 🧠 Can One Domain Help Others? A Data-Cent...

2025.07.23 | TIM模型突破LLM上下文限制;Step-Audio 2提升多模态语音对话。 24.07.2025

本期的 15 篇论文如下: [00:24] ♾ Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning(超越上下文限制:用于长程推理的潜意识线索) [01:05] 🔊 Step-Audio 2 Technical Report(Step-Audio 2 技术报告) [01:41] 🚀 MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning(MegaScience:推动科学推理后训练数据集的前沿) [02:23] ⚡ Upsample What Matters: Region-Adap...

2025.07.22 | MiroMind-M1提升数学推理;GUI-G$^2$高斯奖励助GUI定位。 22.07.2025

本期的 15 篇论文如下: [00:25] 🧮 MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization(MiroMind-M1:通过上下文感知多阶段策略优化实现数学推理的开源进展) [01:00] 🎯 GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding(GUI-G$^2$: 用于GUI定位的高斯奖励建模) [01:42] ⛓ The Invisible Leash: Why RLVR May Not Escape Its Origin(隐形束缚:R...

2025.07.21 | dLLM新型安全漏洞,现有防御不足;俄语语音合成,数据与标注是核心。 22.07.2025

本期的 10 篇论文如下: [00:20] 😈 The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs(隐藏在面具后的恶魔:扩散大语言模型的一种新兴安全漏洞) [01:12] 🎤 A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models(解决俄语语音生成模型中语音与韵律挑战的数据中心框架) [02:07] 🧩 Franca: Nested Matryoshka Clustering for Scalab...

【周末特辑】7月第3周最火AI论文 | 上下文工程提升LLM性能;反射生成模型提高推理效率。 20.07.2025

本期的 5 篇论文如下: [00:39] TOP1(🔥116) | 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [02:35] TOP2(🔥86) | 🧠 Test-Time Scaling with Reflective Generative Model(基于反射生成模型的测试时缩放) [04:31] TOP3(🔥74) | 🤔 Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination(推理还是记忆?数据污染导致强化学习...

2025.07.18 | 优化LLMs上下文;提升视觉语言模型效率 19.07.2025

本期的 15 篇论文如下: [00:27] 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [01:16] 🧠 VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning(VisionThink:基于强化学习的智能高效视觉语言模型) [02:08] 📸 $π^3$: Scalable Permutation-Equivariant Visual Geometry Learning($\pi^3$:可扩展的置换等变视觉几何学习) [02:52] 🤖 The Im...

2025.07.17 | RAG提升LLM推理;PhysX生成物理3D资产 18.07.2025

本期的 13 篇论文如下: [00:26] 🧠 Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs(具身智能RAG与深度推理:LLM中RAG推理系统综述) [01:17] 🧱 PhysX: Physical-Grounded 3D Asset Generation(PhysX:基于物理的3D资产生成) [02:04] 🚗 MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding(MMHU:一个用于人类行为理解的大规模多模态基准) [03:05] 🚀 SWE...

2025.07.16 | VLV自编码器降低训练成本;EXAONE 4.0增强推理能力。 17.07.2025

本期的 8 篇论文如下: [00:28] 💡 Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models(视觉-语言-视觉自编码器:从扩散模型中进行可扩展的知识蒸馏) [01:27] 🤖 EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes(EXAONE 4.0:融合非推理与推理模式的统一大型语言模型) [02:24] ⚖ Scaling Laws for Optimal Data Mixtures(最优数据混合...

2025.07.15 | 数据集支持虚拟人生成;强化学习需防数据污染。 16.07.2025

本期的 12 篇论文如下: [00:24] 🗣 SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation(SpeakerVid-5M:用于视听二元交互式虚拟人生成的大规模高质量数据集) [01:12] 🤔 Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination(推理还是记忆?数据污染导致强化学习结果不可靠) [02:03] 🤖 EmbRACE-3K: Embodied Reasonin...

2025.07.14 | 高效推理路径选择;压缩光场令牌渲染 14.07.2025

本期的 14 篇论文如下: [00:22] 🧠 Test-Time Scaling with Reflective Generative Model(基于反射生成模型的测试时缩放) [00:59] 💡 CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering(CLiFT:用于计算高效和自适应神经渲染的压缩光场令牌) [01:34] 💻 NeuralOS: Towards Simulating Operating Systems via Neural Generative Models(NeuralOS:迈向通过神经生成模型模拟操作系...

【周末特辑】7月第2周最火AI论文 | 长视频推理框架创新;内存操作系统提升AI性能 13.07.2025

本期的 5 篇论文如下: [00:42] TOP1(🔥109) | 🎬 Scaling RL to Long Videos(强化学习驱动视觉语言模型扩展至长视频) [02:54] TOP2(🔥106) | 🧠 MemOS: A Memory OS for AI System(MemOS:面向人工智能系统的内存操作系统) [05:19] TOP3(🔥91) | 🖼 T-LoRA: Single Image Diffusion Model Customization Without Overfitting(T-LoRA:无过拟合的单图像扩散模型定制) [07:51] TOP4(🔥88) | 💡 SingLoRA: Low Rank Adaptation...

2025.07.11 | 长视频推理效率提升;单图像定制模型防过拟合。 11.07.2025

本期的 15 篇论文如下: [00:25] 🎬 Scaling RL to Long Videos(强化学习驱动视觉语言模型扩展至长视频) [01:10] 🖼 T-LoRA: Single Image Diffusion Model Customization Without Overfitting(T-LoRA:无过拟合的单图像扩散模型定制) [01:49] 🖼 Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology(可追踪证据增强的视觉基础推理:评估与方法) [02:28] 🤖 OST-Bench: Evaluating the Capabi...

2025.07.10 | 零样本运动生成突破;4K图像超分辨率提升。 10.07.2025

本期的 14 篇论文如下: [00:22] 🤸 Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data(趋向于零:基于百万级数据的零样本运动生成) [01:03] 🖼 4KAgent: Agentic Any Image to 4K Super-Resolution(4KAgent:将任意图像转化为4K超分辨率的智能体系统) [01:39] 🖼 Perception-Aware Policy Optimization for Multimodal Reasoning(多模态推理的感知感知策略优化) [02:24] 🧪 Rethinking Verification...

2025.07.09 | 潜在推理提升LLM表达能力;SingLoRA优化低秩适应性能。 09.07.2025

本期的 15 篇论文如下: [00:25] 🤔 A Survey on Latent Reasoning(潜在推理研究综述) [00:59] 💡 SingLoRA: Low Rank Adaptation Using a Single Matrix(SingLoRA:使用单矩阵的低秩适应) [01:47] 🧩 OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion(OmniPart:基于语义解耦和结构内聚的部件感知三维生成) [02:36] 🤖 CriticLean: Critic-Guided Reinforcement Learning for Mathema...

2025.07.08 | MemOS提升内存管理效率;MLM与CLM结合优化编码器训练。 08.07.2025

本期的 15 篇论文如下: [00:21] 🧠 MemOS: A Memory OS for AI System(MemOS:面向人工智能系统的内存操作系统) [01:07] 🤔 Should We Still Pretrain Encoders with Masked Language Modeling?(我们是否还应该使用掩码语言模型预训练编码器?) [01:43] 🎥 4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture(4DSloMo:基于异步捕获的高速场景4D重建) [02:22] 🤖 DreamVLA: A Vision-Language-Act...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.