duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

【周末特辑】4月第2周最火AI论文 | SmolVLM优化多模态模型性能;OmniSVG提升SVG生成质量。 12.04.2025

本期的 5 篇论文如下: [00:44] TOP1(🔥149) | 💡 SmolVLM: Redefining small and efficient multimodal models(SmolVLM:重新定义小型高效多模态模型) [03:07] TOP2(🔥125) | 🎨 OmniSVG: A Unified Scalable Vector Graphics Generation Model(OmniSVG:一个统一的可扩展矢量图形生成模型) [05:57] TOP3(🔥90) | 🎬 One-Minute Video Generation with Test-Time Training(基于测试时训练的分钟级视频生成) [08:13] TOP4(🔥...

2025.04.11 | Kimi-VL模型表现优异;VCR-Bench评估推理瓶颈。 11.04.2025

本期的 14 篇论文如下: [00:22] 🧠 Kimi-VL Technical Report(Kimi-VL技术报告) [01:05] 🎬 VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning(VCR-Bench:一个用于视频链式思考推理的综合评估框架) [01:54] 🖼 MM-IFEngine: Towards Multimodal Instruction Following(MM-IFEngine: 面向多模态指令跟随) [02:35] 🖼 VisualCloze: A Universal Image Generation Framework via Visual I...

2025.04.10 | DDT提升图像生成质量;GenDoP优化相机轨迹生成。 10.04.2025

本期的 15 篇论文如下: [00:25] 🎨 DDT: Decoupled Diffusion Transformer(解耦扩散Transformer) [01:05] 🎬 GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography(GenDoP:基于自回归的相机轨迹生成,如同电影摄影师一般) [01:49] 🔍 OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens(OLMoTrace:将语言模型的输出追溯到数万亿的训练文本) [02:28] 🖼 A Un...

2025.04.09 | OmniSVG生成高质量SVG图形;Skywork R1V多模态推理出色。 09.04.2025

本期的 13 篇论文如下: [00:22] 🎨 OmniSVG: A Unified Scalable Vector Graphics Generation Model(OmniSVG:一个统一的可扩展矢量图形生成模型) [01:02] 🧠 Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought(Skywork R1V:以思维链引领多模态推理) [01:42] 🖼 An Empirical Study of GPT-4o Image Generation Capabilities(GPT-4o图像生成能力实证研究) [02:22] 🚀 Hogwild! Inference: Parallel LLM...

2025.04.08 | 分钟级AI视频生成;小型模型超越大型模型 08.04.2025

本期的 15 篇论文如下: [00:21] 🎬 One-Minute Video Generation with Test-Time Training(基于测试时训练的分钟级视频生成) [01:03] 💡 SmolVLM: Redefining small and efficient multimodal models(SmolVLM:重新定义小型高效多模态模型) [01:39] 🖼 URECA: Unique Region Caption Anything(URECA:独特区域描述一切) [02:17] 🧰 T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language...

2025.04.07 | 多语言基准测试揭示LLMs跨语言泛化局限,具身智能新方法提升规划效率与适应性。 07.04.2025

本期的 15 篇论文如下: [00:23] 🛠 Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving(Multi-SWE-bench:一个用于问题解决的多语言基准测试) [01:07] 🧠 Agentic Knowledgeable Self-awareness(具身智能的知识型自我感知) [01:49] 🧮 MegaMath: Pushing the Limits of Open Math Corpora(MegaMath:推动开放数学语料库的极限) [02:32] 🤖 SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge...

【月末特辑】3月最火AI论文 | 稀疏自编码器提升文本检测,动态Tanh优化Transformer 06.04.2025

本期的 10 篇论文如下: [00:42] TOP1(🔥226) | 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [03:07] TOP2(🔥153) | 🧠 Transformers without Normalization(无需归一化的Transformer) [04:59] TOP3(🔥136) | 🎥 DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation(DropletVideo:探...

【周末特辑】4月第1周最火AI论文 | 智能体设计挑战,视觉文本生成创新。 05.04.2025

本期的 5 篇论文如下: [00:40] TOP1(🔥101) | 🧠 Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems(具身智能体的进展与挑战:从脑启发智能到进化、协作与安全系统) [03:17] TOP2(🔥83) | 🖼 TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes(TextCrafter:复杂视觉场景中准确渲染多重文本) [05:27] TOP3(🔥80)...

2025.04.04 | 智能体自主提升,视觉编辑推理重要。 04.04.2025

本期的 15 篇论文如下: [00:19] 🧠 Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems(具身智能体的进展与挑战:从脑启发智能到进化、协作与安全系统) [01:01] 🖼 Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing(超越像素的展望:推理驱动的视觉编辑基准测试) [01:41] 🖼 GPT-ImgEval: A Comprehensive Ben...

2025.04.03 | MergeVQ高效生成高质量图像,类R1-Zero提升视觉空间推理。 03.04.2025

本期的 15 篇论文如下: [00:23] 🎨 MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization(MergeVQ:一种用于视觉生成和表示的统一框架,具有解耦的Token合并和量化) [01:00] 🧠 Improved Visual-Spatial Reasoning via R1-Zero-Like Training(通过类R1-Zero训练改进视觉空间推理) [01:45] 🎮 AnimeGamer: Infinite Anime Life Simulation with Next Gam...

2025.04.02 | 视频生成精度提升,强化学习增强视频理解。 02.04.2025

本期的 15 篇论文如下: [00:21] 🎬 Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation(Any2Caption:将任意条件解析为描述以实现可控视频生成) [01:01] 🎬 Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1(探索强化学习对视频理解的影响:来自SEED-Bench-R1的见解) [01:48] ⚖ JudgeLRM: Large Reasoning Models as a Judge(Judge...

2025.04.01 | 多文本渲染新方法,电影级对话角色合成 01.04.2025

本期的 15 篇论文如下: [00:22] 🖼 TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes(TextCrafter:复杂视觉场景中准确渲染多个文本) [00:59] 🎬 MoCha: Towards Movie-Grade Talking Character Synthesis(MoCha:面向电影级对话角色合成) [01:39] 🔍 What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models(什么、如何、何地以及如何有效?大型语言模型中测试...

2025.03.31 | 减少token使用,提升领域效率。 31.03.2025

本期的 15 篇论文如下: [00:22] 💡 AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation(AdaptiVocab:通过轻量级词汇自适应增强LLM在特定领域的效率) [01:01] 🤖 Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback(探索人类反馈强化学习中的数据缩放趋势与影响) [01:41] 🤔 Think Before Recommend: Unleashing the Latent Reaso...

【周末特辑】3月第4周最火AI论文 | 稀疏自编码器解读LLM推理特征,多模态模型创新。 29.03.2025

本期的 5 篇论文如下: [00:37] TOP1(🔥109) | 🧠 I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders(我已经覆盖了所有基础:通过稀疏自编码器解读大型语言模型中的推理特征) [02:42] TOP2(🔥92) | 🤖 Qwen2.5-Omni Technical Report(Qwen2.5-Omni技术报告) [05:10] TOP3(🔥83) | 🎬 Video-T1: Test-Time Scaling for Video Generation(Video-T1:面向...

2025.03.28 | 视频推理提升,GUI动作预测优化 28.03.2025

本期的 15 篇论文如下: [00:22] 🧠 Video-R1: Reinforcing Video Reasoning in MLLMs(Video-R1:增强多模态大语言模型中的视频推理) [01:02] 📱 UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning(UI-R1:通过强化学习增强GUI代理的动作预测) [01:41] 🤯 Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models(挑战推理的边界:一个面向大型语...

2025.03.27 | Dita跨模态策略优异,Qwen2.5-Omni多模态实时响应。 27.03.2025

本期的 15 篇论文如下: [00:26] 🤖 Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy(Dita:扩展扩散Transformer以实现通用视觉-语言-动作策略) [01:07] 🤖 Qwen2.5-Omni Technical Report(Qwen2.5-Omni技术报告) [01:46] 🧩 LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?(乐高拼图:多模态大型语言模型在多步空间推理方面的表现如何?) [02:35] 🎬 Wan: Open and...

2025.03.26 | 视频预测性能提升,多模态预训练效果显著。 26.03.2025

本期的 15 篇论文如下: [00:22] 🎬 Long-Context Autoregressive Video Modeling with Next-Frame Prediction(基于下一帧预测的长程上下文自回归视频建模) [01:01] 🖼 CoMP: Continual Multimodal Pre-training for Vision Foundation Models(CoMP:面向视觉基础模型的持续多模态预训练) [01:42] 🎬 Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation(探索大...

2025.03.25 | 稀疏自编码器解读LLM中的推理特征,交互视频革新 25.03.2025

本期的 15 篇论文如下: [00:24] 🧠 I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders(我已经覆盖了所有基础:通过稀疏自编码器解读大型语言模型中的推理特征) [01:03] 🎮 Position: Interactive Generative Video as Next-Generation Game Engine(立场:交互式生成视频作为下一代游戏引擎) [01:47] 🎬 Video-T1: Test-Time Scaling for Video Generati...

2025.03.24 | 多智能体协作提升性能,苏格拉底式对话优化提示。 24.03.2025

本期的 15 篇论文如下: [00:22] 🧠 MAPS: A Multi-Agent Framework Based on Big Seven Personality and Socratic Guidance for Multimodal Scientific Problem Solving(MAPS:一个基于大七人格和苏格拉底指导的多智能体框架,用于多模态科学问题求解) [01:09] 🤖 MARS: A Multi-Agent Framework Incorporating Socratic Guidance for Automated Prompt Optimization(MARS:一个融合苏格拉底式指导的多智能体自动提示优化框架...

【周末特辑】3月第3周最火AI论文 | 序列建模创新,视频渲染突破 22.03.2025

本期的 5 篇论文如下: [00:37] TOP1(🔥118) | 🦢 RWKV-7 "Goose" with Expressive Dynamic State Evolution(RWKV-7 "Goose":具有表达性动态状态演化的序列建模) [02:36] TOP2(🔥115) | 🎥 ReCamMaster: Camera-Controlled Generative Rendering from A Single Video(ReCamMaster:基于单视频的相机控制生成式渲染) [05:18] TOP3(🔥89) | 🤖 DAPO: An Open-Source LLM Reinforcement Learning System at Scale(DAPO:一个大...

2025.03.21 | 蒸馏提升超分辨率效率,优化推理减少计算负担。 21.03.2025

本期的 15 篇论文如下: [00:23] 🖼 One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation(基于蒸馏的单步残差转移扩散超分辨率) [01:01] 🤔 Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models(停止过度思考:大型语言模型高效推理综述) [01:38] 🚀 Unleashing Vecset Diffusion Model for Fast Shape Generation(释放Vecset扩散模型以实现快速形状生成) [02:18]...

2025.03.20 | 自适应前瞻采样优化推理;强化学习提升3D网格质量 20.03.2025

本期的 15 篇论文如下: [00:23] 🔍 $φ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation($\phi$-解码:用于平衡推理时探索与利用的自适应前瞻采样) [01:08] 🎨 DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning(DeepMesh:基于强化学习的自回归艺术家网格创建) [01:51] 🌷 TULIP: Towards Unified Language-Image Pretraining(TULIP:迈向统...

2025.03.19 | 动态序列建模优势,视频生成理解挑战 19.03.2025

本期的 15 篇论文如下: [00:21] 🦢 RWKV-7 "Goose" with Expressive Dynamic State Evolution(RWKV-7 "Goose":具有表达性动态状态演化的序列建模) [00:55] 🤯 Impossible Videos(不可能的视频) [01:38] 🎨 Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM(Creation-MMBench:评估多模态大型语言模型中具有上下文感知能力的创造性智能) [02:17] 🤖 DAPO: An Open-Source LLM Reinforcement Learn...

2025.03.18 | 视频生成新方法,人形机器人新框架 18.03.2025

本期的 15 篇论文如下: [00:21] 🎥 DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation(DropletVideo:探索整体时空一致性视频生成的数据集与方法) [01:10] 🤖 Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills(Being-0:一个具有视觉-语言模型和模块化技能的人形机器人代理) [01:49] 🖼 DreamRenderer: Taming Multi-Instance Attrib...

2025.03.17 | 新相机轨迹生成,稀疏性提升图像质量 17.03.2025

本期的 15 篇论文如下: [00:25] 🎥 ReCamMaster: Camera-Controlled Generative Rendering from A Single Video(ReCamMaster:基于单视频的相机控制生成式渲染) [01:11] 💡 PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity(PLADIS:通过利用稀疏性,在扩散模型推理时突破注意力机制的限制) [01:50] 🤖 Adversarial Data Collection: Human-Collaborative Perturbatio...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.