duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2025.10.28 | Point Transformer无标对齐长空间;代码递归统一粗细粒度 28.10.2025

本期的 15 篇论文如下: [00:23] 🎼 Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations(Concerto:2D-3D联合自监督学习涌现空间表征) [01:06] 🧩 ReCode: Unify Plan and Action for Universal Granularity Control(ReCode:用递归代码统一规划与行动,实现通用粒度控制) [01:44] 🤖 A Survey of Data Agents: Emerging Paradigm or Overstated Hype?(数据智能体全景透视:新范式还是泡沫?)...

2025.10.27 | DeepAgent一步推理+ToolPO;视频即提示DiT秒控百种语义 27.10.2025

本期的 15 篇论文如下: [00:27] 🧠 DeepAgent: A General Reasoning Agent with Scalable Toolsets(DeepAgent:具备可扩展工具集的通用推理智能体) [01:01] 🎬 Video-As-Prompt: Unified Semantic Control for Video Generation(视频即提示:统一语义控制的视频生成新范式) [01:35] 🔧 From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model(从去噪到精修:视觉-语言扩散模型的纠错式生...

【周末特辑】10月第4周最火AI论文 | 内部概率+投票剪尾,RPC省样本提精度 26.10.2025

本期的 5 篇论文如下: [00:29] TOP1(🔥135) | 🧠 A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning(大模型推理中内部概率与自洽性桥接的理论研究) [03:02] TOP2(🔥104) | 🚀 Efficient Long-context Language Model Training by Core Attention Disaggregation(通过核心注意力拆解实现高效长上下文语言模型训练) [05:29] TOP3(🔥100) | 🧠 LightMem: Lightweight and Efficient...

2025.10.24 | AdaSPEC挑40% token提速两成;AutoPage 10美分生成交互网页 24.10.2025

本期的 15 篇论文如下: [00:23] 🎯 AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders(AdaSPEC:面向高效推测解码的选择性知识蒸馏) [00:57] 🤖 Human-Agent Collaborative Paper-to-Page Crafting for Under $0.1(低成本人机协作论文一键成页:低于0.1美元) [01:35] 🔍 Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence(Open-o3视频:显式时空证据支撑的开放...

2025.10.23 | 线性注意力显存降十倍;动态裁剪PPO稳提分 23.10.2025

本期的 15 篇论文如下: [00:19] 🧠 Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning(每一种注意力都重要:面向长上下文推理的高效混合架构) [00:59] ⚖ BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping(BAPO:通过自适应裁剪的平衡策略优化稳定LLM离策略强化学习) [01:40] 🧠 LoongRL:Reinforcement Learning...

2025.10.22 | LightMem压缩记忆千倍提速12倍;闭环世界模型微调8万数据反超巨兽 22.10.2025

本期的 14 篇论文如下: [00:19] 🧠 LightMem: Lightweight and Efficient Memory-Augmented Generation(LightMem:轻量高效的记忆增强生成框架) [00:55] 🌀 World-in-World: World Models in a Closed-Loop World(世界中的世界:闭环环境下的世界模型) [01:44] 🖼 UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation(UniGenBench++:面向文本到图像生成的统一语义评测基准) [02:29] 🧪 C...

2025.10.21 | 模型不懂光影折射;小模型也能写报告 21.10.2025

本期的 13 篇论文如下: [00:21] 🪞 PICABench: How Far Are We from Physically Realistic Image Editing?(PICABench:我们离物理真实的图像编辑还有多远?) [01:04] 🤖 DeepAnalyze: Agentic Large Language Models for Autonomous Data Science(DeepAnalyze:面向自主数据科学的智能体大模型) [01:50] 🗜 Glyph: Scaling Context Windows via Visual-Text Compression(Glyph:通过视觉-文本压缩扩展上下文窗口长度) [02:23...

2025.10.20 | RPC剪枝提速保准;OmniVinci小数据跨模态称王 20.10.2025

本期的 15 篇论文如下: [00:20] 🧠 A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning(大模型推理中内部概率与自洽性桥接的理论研究) [01:04] 🌐 OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM(OmniVinci:面向全模态理解大模型的架构与数据增强) [01:44] 🎬 Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset(...

【周末特辑】10月第3周最火AI论文 | 量化噪声变探索,单卡跑RL;冻结编码器放语义,DiT生成新纪录 18.10.2025

本期的 5 篇论文如下: [00:40] TOP1(🔥154) | 🚀 QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs(QeRL:超越效率——面向大语言模型的量化增强强化学习) [02:19] TOP2(🔥138) | 🧠 Diffusion Transformers with Representation Autoencoders(基于表示自编码器的扩散Transformer) [04:54] TOP3(🔥134) | 🎯 Spatial Forcing: Implicit Spatial Representation Alignment for Vision-languag...

2025.10.17 | AI眼镜预判式服务;视频生成补想象力 17.10.2025

本期的 11 篇论文如下: [00:25] 👓 AI for Service: Proactive Assistance with AI Glasses(AI服务:AI眼镜的主动式协助) [01:06] 🎬 ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints(ImagerySearch:面向超越语义依赖约束的自适应测试时搜索视频生成) [01:43] 🎯 LaSeR: Reinforcement Learning with Last-Token Self-Rewarding(LaSeR:基于末词元自奖励的强化学习...

2025.10.16 | UniMoE一统语音音乐;注意力图点亮大模型推理 16.10.2025

本期的 15 篇论文如下: [00:21] 🎧 UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE(UniMoE-Audio:基于动态容量MoE的统一语音与音乐生成模型) [00:57] 🔍 Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization(注意力照亮大模型推理:预规划-锚定节奏实现细粒度策略优化) [01:38] ⚡ FlashWorld: High-quality 3D Scene Generation...

2025.10.15 | 像素级自监督ViT刷新生成基准;多智能体评测网文翻译新标尺 15.10.2025

本期的 14 篇论文如下: [00:20] 🖼 Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training(通过自监督预训练推进端到端像素空间生成建模) [00:53] 📚 DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation(DITING:面向网络小说翻译评测的多智能体基准框架) [01:41] 🌐 Scaling Language-Centric Omnimodal Representation Learning(以语言为中心的跨模态...

2025.10.14 | 量化误差变奖励,单卡训32B;面向多模态大模型的音视频评测基准 14.10.2025

本期的 15 篇论文如下: [00:23] 🚀 QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs(QeRL:超越效率——面向大语言模型的量化增强强化学习) [01:22] 🧠 Diffusion Transformers with Representation Autoencoders(基于表示自编码器的扩散Transformer) [02:12] 🎬 OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs(OmniVideoBench:面向全向多模态大模型的音...

2025.10.13 | 桌面交互预训练解锁机器人潜能;统一模型赋予相机空间想象力 13.10.2025

本期的 14 篇论文如下: [00:20] 🖥 D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI(D2E:利用桌面数据规模化视觉-动作预训练以迁移至具身智能) [01:13] 📷 Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation(基于相机的统一多模态理解与生成模型) [01:56] 🎨 TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffus...

【周末特辑】10月第2周最火AI论文 | 递归小模型刷爆推理榜;未来经验点亮零奖励学习 12.10.2025

本期的 5 篇论文如下: [00:33] TOP1(🔥300) | 🧠 Less is More: Recursive Reasoning with Tiny Networks(小而精:用微型网络递归推理) [02:16] TOP2(🔥164) | 🌱 Agent Learning via Early Experience(基于早期经验的主体学习) [04:15] TOP3(🔥105) | 🧠 Apriel-1.5-15b-Thinker(Apriel-1.5-15B-Thinker:以小博大实现前沿多模态推理的15B开源模型) [06:17] TOP4(🔥97) | 🧠 MM-HELIX: Boosting Multimodal Long-Chain Ref...

2025.10.10 | 早期经验的Agent Learning;图文交错反思链跃升至24.9% 10.10.2025

本期的 14 篇论文如下: [00:16] 🌱 Agent Learning via Early Experience(基于早期经验的主体学习) [00:50] 🧠 MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization(MM-HELIX:以整体平台与自适应混合策略优化激发多模态长链反思推理) [01:42] 🧪 From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reaso...

2025.10.09 | Ming-UniVision统一视觉词表;KV-Cache直连让大模型秒聊 09.10.2025

本期的 15 篇论文如下: [00:21] 🔄 Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer(Ming-UniVision:用统一连续视觉词表打通图像理解与生成) [00:59] 🧠 Cache-to-Cache: Direct Semantic Communication Between Large Language Models(缓存到缓存:大模型间的直接语义通信) [01:32] 🌀 Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation a...

2025.10.08 | TaTToo用外挂代码干翻大模型;4B小模型32步逼近闭源巨头 08.10.2025

本期的 15 篇论文如下: [00:24] 📊 TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning(TaTToo:面向表格推理测试时扩展的“工具落地思维”过程奖励模型) [00:57] 🔍 Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs(Fathom-DeepResearch:解锁小模型长程信息检索与综合的钥匙) [01:39] 🚀 Fast-dLLM v2: Efficient Block-Diffusion LLM(Fast-dLLM v...

2025.10.07 | 论文秒变演讲;Video-LMM后训练突破 07.10.2025

本期的 15 篇论文如下: [00:21] 🎬 Paper2Video: Automatic Video Generation from Scientific Papers(论文自动生成学术演讲视频) [00:55] 🎬 Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models(Video-LMM后训练:深入剖析大型多模态模型的视频推理) [01:38] 🎬 VChain: Chain-of-Visual-Thought for Reasoning in Video Generation(VChain:面向视频生成推理的视觉思维链) [02:14]...

2025.10.06 | 15B小模型追平DeepSeek-R1;渐进蒸馏128 token省八成算力 06.10.2025

本期的 15 篇论文如下: [00:28] 🧠 Apriel-1.5-15b-Thinker(Apriel-1.5-15B-Thinker:以小博大实现前沿多模态推理的15B开源模型) [01:04] 🚀 Efficient Multi-modal Large Language Models via Progressive Consistency Distillation(基于渐进一致性蒸馏的高效多模态大模型) [01:42] 🧩 Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition(组合...

【周末特辑】10月第1周最火AI论文 | Transformer长出大脑的壳;LongLive把长视频做成直播 05.10.2025

本期的 5 篇论文如下: [00:43] TOP1(🔥323) | 🐣 The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain(幼龙破壳: Transformer 与大脑模型之间缺失的环节) [02:38] TOP2(🔥167) | 🎬 LongLive: Real-time Interactive Long Video Generation(LongLive:实时交互式长视频生成框架) [05:04] TOP3(🔥150) | 🔥 MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP U...

2025.10.03 | LongCodeZip删得快准;迈向分钟级高质量视频生成 03.10.2025

本期的 15 篇论文如下: [00:22] 🗜 LongCodeZip: Compress Long Context for Code Language Models(LongCodeZip:面向代码大模型的长上下文压缩方法) [00:56] 🎬 Self-Forcing++: Towards Minute-Scale High-Quality Video Generation(自增强++:迈向分钟级高质量视频生成) [01:38] 🧠 ExGRPO: Learning to Reason from Experience(基于经验的群体相对策略优化:让大模型学会从经验中推理) [02:32] 🥷 StealthAttack: Robust...

2025.10.02 | MCTS破局RLVR瓶颈;GEM开源智能体训练场 02.10.2025

本期的 15 篇论文如下: [00:19] 🧠 DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search(DeepSearch:以蒙特卡洛树搜索破解强化学习可验证奖励瓶颈) [01:20] 🤖 GEM: A Gym for Agentic LLMs(GEM:面向智能体大模型的开放训练场) [01:57] 🧠 VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators(VLA-RF...

【月末特辑】9月最火AI论文 | 群体RL共享降本;SAPO让旧机也能训大模型 02.10.2025

本期的 10 篇论文如下: [00:29] TOP1(🔥640) | 🤝 Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing(共享即关爱:基于集体RL经验共享的高效大模型后训练) [02:49] TOP2(🔥341) | 🔒 A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code(A.S.E:一个用于评估AI生成代码安全的仓库级基准) [04:59] TOP3(🔥218) | 🤖 VLA-Adapter: An Effective Paradigm f...

2025.10.01 | 自对弈零标注训练;MCP代理深度评测 01.10.2025

本期的 15 篇论文如下: [00:20] 🎮 Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play(Vision-Zero:基于策略化博弈自对弈的可扩展视觉语言模型自我提升) [00:59] 🔥 MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use(MCPMark:面向真实且全面的MCP应用场景的压力测试基准) [01:36] 🐣 The Dragon Hatchling: The Missing Link between the Transformer and Models...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.