duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2025.08.04 | 扩散语言模型变长去噪,高效省资源;PixNerd图像扩散,高效高质量。 05.08.2025 5:39
本期的 11 篇论文如下: [00:22] 🔄 Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models(超越固定长度:扩散大语言模型的可变长度去噪) [00:44] 🎨 PixNerd: Pixel Neural Field Diffusion(PixNerd:像素神经场扩散) [01:11] 💡 SWE-Exp: Experience-Driven Software Issue Resolution(SWE-Exp:经验驱动的软件问题解决) [01:38] 🔍 Multimodal Referring Segmentation: A Survey(多模态指代表...
【月末特辑】7月最火AI论文 | GSPO稳训练;序列级裁剪降方差;上下文工程综述,动态拼装信息流 04.08.2025 18:10
本期的 10 篇论文如下: [00:30] TOP1(🔥257) | 🚀 Group Sequence Policy Optimization(组序列策略优化) [02:21] TOP2(🔥227) | 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [03:33] TOP3(🔥207) | 🧠 GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning(GLM-4.1V-Thinking:基于可扩展强化学习的通用多模态推理) [05:02] T...
【周末特辑】8月第1周最火AI论文 | ARPO用高熵分叉省预算;混元世界一句话生成可编辑3D场景 03.08.2025 11:27
本期的 5 篇论文如下: [00:32] TOP1(🔥114) | 🤖 Agentic Reinforced Policy Optimization(智能体强化策略优化) [02:17] TOP2(🔥94) | 🌍 HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels(混元世界 1.0:从文字或像素生成沉浸式、可探索、可交互的3D世界) [05:04] TOP3(🔥76) | 🏆 Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving(Seed-Prover...
2025.08.01 | Seed-Prover融合LLM解决IMO数学题;Phi-Ground提升GUI感知精度。 01.08.2025 9:39
本期的 15 篇论文如下: [00:22] 🏆 Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving(Seed-Prover:自动化定理证明的深度与广度推理) [01:04] 🎯 Phi-Ground Tech Report: Advancing Perception in GUI Grounding(Phi-Ground 技术报告:提升 GUI 接地感知能力) [01:30] 🤔 C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations(C3:探索复杂对话挑战...
2025.07.31 | ScreenCoder自动化UI转代码;Falcon-H1混合架构,提升长序列效率。 01.08.2025 6:26
本期的 9 篇论文如下: [00:22] 💻 ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents(ScreenCoder:模块化多模态智能体赋能前端视觉代码生成) [01:02] 🚀 Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance(Falcon-H1:重塑效率与性能的混合架构语言模型系列) [01:33] 💥 BANG: Dividing 3D Assets via Generative Explod...
2025.07.30 | 混元世界从文字像素生成沉浸3D世界;X-Omni用强化学习提升图像生成质量。 31.07.2025 6:27
本期的 8 篇论文如下: [00:23] 🌍 HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels(混元世界 1.0:从文字或像素生成沉浸式、可探索、可交互的3D世界) [00:56] ✨ X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again(X-Omni:强化学习让离散自回归图像生成模型再展辉煌) [01:59] 🚀 CUDA-L1: Improving CUDA Optimizat...
2025.07.29 | ARPO提升LLM工具交互性能;ARC-Hunyuan-Video-7B深耕短视频理解。 30.07.2025 10:31
本期的 15 篇论文如下: [00:23] 🤖 Agentic Reinforced Policy Optimization(智能体强化策略优化) [00:55] 🧠 ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts(ARC-Hunyuan-Video-7B:真实世界短视频的结构化理解) [01:35] 🚀 Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning(Rep-MTL:释放表示层任务显著性在多任务学习中的力量) [02:03] 🌐 R...
2025.07.28 | GPTQ揭示为Babai算法,保障精度;TTD-DR以扩散模型生成高质量研究报告。 29.07.2025 4:05
本期的 5 篇论文如下: [00:25] 💡 The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm(LLM 量化的几何学:GPTQ 作为 Babai 最近平面算法) [00:52] ✨ Deep Researcher with Test-Time Diffusion(基于测试时扩散的深度研究智能体) [01:40] 🔧 Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement(规范自校正:通过测试时细化缓解上下文奖励破解) [...
【周末特辑】7月第4周最火AI论文 | GUI-G2:高斯奖励提升GUI定位;MiroMind-M1:开源数学推理LLM 26.07.2025 15:55
本期的 5 篇论文如下: [00:36] TOP1(🔥118) | 🎯 GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding(GUI-G$^2$: 基于高斯奖励模型的GUI定位) [02:14] TOP2(🔥108) | 🧮 MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization(MiroMind-M1:通过上下文感知多阶段策略优化实现数学推理的开源进展) [05:19] TOP3(🔥96) | ♾ Beyond Context Limits: Subco...
2025.07.25 | GSPO解决大模型训练崩溃;MUR提升LLM推理效率。 26.07.2025 9:47
本期的 15 篇论文如下: [00:24] 🚀 Group Sequence Policy Optimization(组序列策略优化) [00:53] 🧠 MUR: Momentum Uncertainty guided Reasoning for Large Language Models(MUR:面向大型语言模型的动量不确定性引导推理) [01:30] 🧠 LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization(LAPO:内化推理效率的长度自适应策略优化) [02:09] 🎬 Captain Cinema: Towards Short Movie Gener...
2025.07.24 | MLLMs视觉感知仍不足;Yume模型可生成交互虚拟世界。 25.07.2025 6:32
本期的 9 篇论文如下: [00:23] 👁 Pixels, Patterns, but No Poetry: To See The World like Humans(像素、模式,却无诗意:像人类一样感知世界) [00:56] 🌌 Yume: An Interactive World Generation Model(Yume:交互式世界生成模型) [01:29] ✨ DesignLab: Designing Slides Through Iterative Detection and Correction(DesignLab:通过迭代检测与修正进行幻灯片设计) [02:14] 🧠 Can One Domain Help Others? A Data-Cent...
2025.07.23 | TIM模型突破LLM上下文限制;Step-Audio 2提升多模态语音对话。 24.07.2025 11:15
本期的 15 篇论文如下: [00:24] ♾ Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning(超越上下文限制:用于长程推理的潜意识线索) [01:05] 🔊 Step-Audio 2 Technical Report(Step-Audio 2 技术报告) [01:41] 🚀 MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning(MegaScience:推动科学推理后训练数据集的前沿) [02:23] ⚡ Upsample What Matters: Region-Adap...
2025.07.22 | MiroMind-M1提升数学推理;GUI-G$^2$高斯奖励助GUI定位。 22.07.2025 13:22
本期的 15 篇论文如下: [00:25] 🧮 MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization(MiroMind-M1:通过上下文感知多阶段策略优化实现数学推理的开源进展) [01:00] 🎯 GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding(GUI-G$^2$: 用于GUI定位的高斯奖励建模) [01:42] ⛓ The Invisible Leash: Why RLVR May Not Escape Its Origin(隐形束缚:R...
2025.07.21 | dLLM新型安全漏洞,现有防御不足;俄语语音合成,数据与标注是核心。 22.07.2025 8:46
本期的 10 篇论文如下: [00:20] 😈 The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs(隐藏在面具后的恶魔:扩散大语言模型的一种新兴安全漏洞) [01:12] 🎤 A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models(解决俄语语音生成模型中语音与韵律挑战的数据中心框架) [02:07] 🧩 Franca: Nested Matryoshka Clustering for Scalab...
【周末特辑】7月第3周最火AI论文 | 上下文工程提升LLM性能;反射生成模型提高推理效率。 20.07.2025 13:31
本期的 5 篇论文如下: [00:39] TOP1(🔥116) | 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [02:35] TOP2(🔥86) | 🧠 Test-Time Scaling with Reflective Generative Model(基于反射生成模型的测试时缩放) [04:31] TOP3(🔥74) | 🤔 Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination(推理还是记忆?数据污染导致强化学习...
2025.07.18 | 优化LLMs上下文;提升视觉语言模型效率 19.07.2025 13:39
本期的 15 篇论文如下: [00:27] 🧮 A Survey of Context Engineering for Large Language Models(大型语言模型上下文工程综述) [01:16] 🧠 VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning(VisionThink:基于强化学习的智能高效视觉语言模型) [02:08] 📸 $π^3$: Scalable Permutation-Equivariant Visual Geometry Learning($\pi^3$:可扩展的置换等变视觉几何学习) [02:52] 🤖 The Im...
2025.07.17 | RAG提升LLM推理;PhysX生成物理3D资产 18.07.2025 12:16
本期的 13 篇论文如下: [00:26] 🧠 Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs(具身智能RAG与深度推理:LLM中RAG推理系统综述) [01:17] 🧱 PhysX: Physical-Grounded 3D Asset Generation(PhysX:基于物理的3D资产生成) [02:04] 🚗 MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding(MMHU:一个用于人类行为理解的大规模多模态基准) [03:05] 🚀 SWE...
2025.07.16 | VLV自编码器降低训练成本;EXAONE 4.0增强推理能力。 17.07.2025 7:44
本期的 8 篇论文如下: [00:28] 💡 Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models(视觉-语言-视觉自编码器:从扩散模型中进行可扩展的知识蒸馏) [01:27] 🤖 EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes(EXAONE 4.0:融合非推理与推理模式的统一大型语言模型) [02:24] ⚖ Scaling Laws for Optimal Data Mixtures(最优数据混合...
2025.07.15 | 数据集支持虚拟人生成;强化学习需防数据污染。 16.07.2025 11:12
本期的 12 篇论文如下: [00:24] 🗣 SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation(SpeakerVid-5M:用于视听二元交互式虚拟人生成的大规模高质量数据集) [01:12] 🤔 Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination(推理还是记忆?数据污染导致强化学习结果不可靠) [02:03] 🤖 EmbRACE-3K: Embodied Reasonin...
2025.07.14 | 高效推理路径选择;压缩光场令牌渲染 14.07.2025 10:27
本期的 14 篇论文如下: [00:22] 🧠 Test-Time Scaling with Reflective Generative Model(基于反射生成模型的测试时缩放) [00:59] 💡 CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering(CLiFT:用于计算高效和自适应神经渲染的压缩光场令牌) [01:34] 💻 NeuralOS: Towards Simulating Operating Systems via Neural Generative Models(NeuralOS:迈向通过神经生成模型模拟操作系...
【周末特辑】7月第2周最火AI论文 | 长视频推理框架创新;内存操作系统提升AI性能 13.07.2025 12:13
本期的 5 篇论文如下: [00:42] TOP1(🔥109) | 🎬 Scaling RL to Long Videos(强化学习驱动视觉语言模型扩展至长视频) [02:54] TOP2(🔥106) | 🧠 MemOS: A Memory OS for AI System(MemOS:面向人工智能系统的内存操作系统) [05:19] TOP3(🔥91) | 🖼 T-LoRA: Single Image Diffusion Model Customization Without Overfitting(T-LoRA:无过拟合的单图像扩散模型定制) [07:51] TOP4(🔥88) | 💡 SingLoRA: Low Rank Adaptation...
2025.07.11 | 长视频推理效率提升;单图像定制模型防过拟合。 11.07.2025 11:04
本期的 15 篇论文如下: [00:25] 🎬 Scaling RL to Long Videos(强化学习驱动视觉语言模型扩展至长视频) [01:10] 🖼 T-LoRA: Single Image Diffusion Model Customization Without Overfitting(T-LoRA:无过拟合的单图像扩散模型定制) [01:49] 🖼 Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology(可追踪证据增强的视觉基础推理:评估与方法) [02:28] 🤖 OST-Bench: Evaluating the Capabi...
2025.07.10 | 零样本运动生成突破;4K图像超分辨率提升。 10.07.2025 10:35
本期的 14 篇论文如下: [00:22] 🤸 Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data(趋向于零:基于百万级数据的零样本运动生成) [01:03] 🖼 4KAgent: Agentic Any Image to 4K Super-Resolution(4KAgent:将任意图像转化为4K超分辨率的智能体系统) [01:39] 🖼 Perception-Aware Policy Optimization for Multimodal Reasoning(多模态推理的感知感知策略优化) [02:24] 🧪 Rethinking Verification...
2025.07.09 | 潜在推理提升LLM表达能力;SingLoRA优化低秩适应性能。 09.07.2025 11:06
本期的 15 篇论文如下: [00:25] 🤔 A Survey on Latent Reasoning(潜在推理研究综述) [00:59] 💡 SingLoRA: Low Rank Adaptation Using a Single Matrix(SingLoRA:使用单矩阵的低秩适应) [01:47] 🧩 OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion(OmniPart:基于语义解耦和结构内聚的部件感知三维生成) [02:36] 🤖 CriticLean: Critic-Guided Reinforcement Learning for Mathema...
2025.07.08 | MemOS提升内存管理效率;MLM与CLM结合优化编码器训练。 08.07.2025 11:04
本期的 15 篇论文如下: [00:21] 🧠 MemOS: A Memory OS for AI System(MemOS:面向人工智能系统的内存操作系统) [01:07] 🤔 Should We Still Pretrain Encoders with Masked Language Modeling?(我们是否还应该使用掩码语言模型预训练编码器?) [01:43] 🎥 4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture(4DSloMo:基于异步捕获的高速场景4D重建) [02:22] 🤖 DreamVLA: A Vision-Language-Act...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.