duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
【周末特辑】4月第2周最火AI论文 | SmolVLM优化多模态模型性能;OmniSVG提升SVG生成质量。 12.04.2025 12:45
本期的 5 篇论文如下: [00:44] TOP1(🔥149) | 💡 SmolVLM: Redefining small and efficient multimodal models(SmolVLM:重新定义小型高效多模态模型) [03:07] TOP2(🔥125) | 🎨 OmniSVG: A Unified Scalable Vector Graphics Generation Model(OmniSVG:一个统一的可扩展矢量图形生成模型) [05:57] TOP3(🔥90) | 🎬 One-Minute Video Generation with Test-Time Training(基于测试时训练的分钟级视频生成) [08:13] TOP4(🔥...
2025.04.11 | Kimi-VL模型表现优异;VCR-Bench评估推理瓶颈。 11.04.2025 10:32
本期的 14 篇论文如下: [00:22] 🧠 Kimi-VL Technical Report(Kimi-VL技术报告) [01:05] 🎬 VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning(VCR-Bench:一个用于视频链式思考推理的综合评估框架) [01:54] 🖼 MM-IFEngine: Towards Multimodal Instruction Following(MM-IFEngine: 面向多模态指令跟随) [02:35] 🖼 VisualCloze: A Universal Image Generation Framework via Visual I...
2025.04.10 | DDT提升图像生成质量;GenDoP优化相机轨迹生成。 10.04.2025 10:49
本期的 15 篇论文如下: [00:25] 🎨 DDT: Decoupled Diffusion Transformer(解耦扩散Transformer) [01:05] 🎬 GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography(GenDoP:基于自回归的相机轨迹生成,如同电影摄影师一般) [01:49] 🔍 OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens(OLMoTrace:将语言模型的输出追溯到数万亿的训练文本) [02:28] 🖼 A Un...
2025.04.09 | OmniSVG生成高质量SVG图形;Skywork R1V多模态推理出色。 09.04.2025 9:42
本期的 13 篇论文如下: [00:22] 🎨 OmniSVG: A Unified Scalable Vector Graphics Generation Model(OmniSVG:一个统一的可扩展矢量图形生成模型) [01:02] 🧠 Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought(Skywork R1V:以思维链引领多模态推理) [01:42] 🖼 An Empirical Study of GPT-4o Image Generation Capabilities(GPT-4o图像生成能力实证研究) [02:22] 🚀 Hogwild! Inference: Parallel LLM...
2025.04.08 | 分钟级AI视频生成;小型模型超越大型模型 08.04.2025 10:50
本期的 15 篇论文如下: [00:21] 🎬 One-Minute Video Generation with Test-Time Training(基于测试时训练的分钟级视频生成) [01:03] 💡 SmolVLM: Redefining small and efficient multimodal models(SmolVLM:重新定义小型高效多模态模型) [01:39] 🖼 URECA: Unique Region Caption Anything(URECA:独特区域描述一切) [02:17] 🧰 T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language...
2025.04.07 | 多语言基准测试揭示LLMs跨语言泛化局限,具身智能新方法提升规划效率与适应性。 07.04.2025 11:16
本期的 15 篇论文如下: [00:23] 🛠 Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving(Multi-SWE-bench:一个用于问题解决的多语言基准测试) [01:07] 🧠 Agentic Knowledgeable Self-awareness(具身智能的知识型自我感知) [01:49] 🧮 MegaMath: Pushing the Limits of Open Math Corpora(MegaMath:推动开放数学语料库的极限) [02:32] 🤖 SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge...
【月末特辑】3月最火AI论文 | 稀疏自编码器提升文本检测,动态Tanh优化Transformer 06.04.2025 25:09
本期的 10 篇论文如下: [00:42] TOP1(🔥226) | 🤖 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders(基于稀疏自编码器的人工文本检测特征分析) [03:07] TOP2(🔥153) | 🧠 Transformers without Normalization(无需归一化的Transformer) [04:59] TOP3(🔥136) | 🎥 DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation(DropletVideo:探...
【周末特辑】4月第1周最火AI论文 | 智能体设计挑战,视觉文本生成创新。 05.04.2025 12:58
本期的 5 篇论文如下: [00:40] TOP1(🔥101) | 🧠 Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems(具身智能体的进展与挑战:从脑启发智能到进化、协作与安全系统) [03:17] TOP2(🔥83) | 🖼 TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes(TextCrafter:复杂视觉场景中准确渲染多重文本) [05:27] TOP3(🔥80)...
2025.04.04 | 智能体自主提升,视觉编辑推理重要。 04.04.2025 11:06
本期的 15 篇论文如下: [00:19] 🧠 Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems(具身智能体的进展与挑战:从脑启发智能到进化、协作与安全系统) [01:01] 🖼 Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing(超越像素的展望:推理驱动的视觉编辑基准测试) [01:41] 🖼 GPT-ImgEval: A Comprehensive Ben...
2025.04.03 | MergeVQ高效生成高质量图像,类R1-Zero提升视觉空间推理。 03.04.2025 10:53
本期的 15 篇论文如下: [00:23] 🎨 MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization(MergeVQ:一种用于视觉生成和表示的统一框架,具有解耦的Token合并和量化) [01:00] 🧠 Improved Visual-Spatial Reasoning via R1-Zero-Like Training(通过类R1-Zero训练改进视觉空间推理) [01:45] 🎮 AnimeGamer: Infinite Anime Life Simulation with Next Gam...
2025.04.02 | 视频生成精度提升,强化学习增强视频理解。 02.04.2025 11:28
本期的 15 篇论文如下: [00:21] 🎬 Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation(Any2Caption:将任意条件解析为描述以实现可控视频生成) [01:01] 🎬 Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1(探索强化学习对视频理解的影响:来自SEED-Bench-R1的见解) [01:48] ⚖ JudgeLRM: Large Reasoning Models as a Judge(Judge...
2025.04.01 | 多文本渲染新方法,电影级对话角色合成 01.04.2025 11:40
本期的 15 篇论文如下: [00:22] 🖼 TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes(TextCrafter:复杂视觉场景中准确渲染多个文本) [00:59] 🎬 MoCha: Towards Movie-Grade Talking Character Synthesis(MoCha:面向电影级对话角色合成) [01:39] 🔍 What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models(什么、如何、何地以及如何有效?大型语言模型中测试...
2025.03.31 | 减少token使用,提升领域效率。 31.03.2025 10:50
本期的 15 篇论文如下: [00:22] 💡 AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation(AdaptiVocab:通过轻量级词汇自适应增强LLM在特定领域的效率) [01:01] 🤖 Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback(探索人类反馈强化学习中的数据缩放趋势与影响) [01:41] 🤔 Think Before Recommend: Unleashing the Latent Reaso...
【周末特辑】3月第4周最火AI论文 | 稀疏自编码器解读LLM推理特征,多模态模型创新。 29.03.2025 13:03
本期的 5 篇论文如下: [00:37] TOP1(🔥109) | 🧠 I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders(我已经覆盖了所有基础:通过稀疏自编码器解读大型语言模型中的推理特征) [02:42] TOP2(🔥92) | 🤖 Qwen2.5-Omni Technical Report(Qwen2.5-Omni技术报告) [05:10] TOP3(🔥83) | 🎬 Video-T1: Test-Time Scaling for Video Generation(Video-T1:面向...
2025.03.28 | 视频推理提升,GUI动作预测优化 28.03.2025 10:44
本期的 15 篇论文如下: [00:22] 🧠 Video-R1: Reinforcing Video Reasoning in MLLMs(Video-R1:增强多模态大语言模型中的视频推理) [01:02] 📱 UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning(UI-R1:通过强化学习增强GUI代理的动作预测) [01:41] 🤯 Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models(挑战推理的边界:一个面向大型语...
2025.03.27 | Dita跨模态策略优异,Qwen2.5-Omni多模态实时响应。 27.03.2025 11:01
本期的 15 篇论文如下: [00:26] 🤖 Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy(Dita:扩展扩散Transformer以实现通用视觉-语言-动作策略) [01:07] 🤖 Qwen2.5-Omni Technical Report(Qwen2.5-Omni技术报告) [01:46] 🧩 LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?(乐高拼图:多模态大型语言模型在多步空间推理方面的表现如何?) [02:35] 🎬 Wan: Open and...
2025.03.26 | 视频预测性能提升,多模态预训练效果显著。 26.03.2025 10:56
本期的 15 篇论文如下: [00:22] 🎬 Long-Context Autoregressive Video Modeling with Next-Frame Prediction(基于下一帧预测的长程上下文自回归视频建模) [01:01] 🖼 CoMP: Continual Multimodal Pre-training for Vision Foundation Models(CoMP:面向视觉基础模型的持续多模态预训练) [01:42] 🎬 Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation(探索大...
2025.03.25 | 稀疏自编码器解读LLM中的推理特征,交互视频革新 25.03.2025 11:13
本期的 15 篇论文如下: [00:24] 🧠 I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders(我已经覆盖了所有基础:通过稀疏自编码器解读大型语言模型中的推理特征) [01:03] 🎮 Position: Interactive Generative Video as Next-Generation Game Engine(立场:交互式生成视频作为下一代游戏引擎) [01:47] 🎬 Video-T1: Test-Time Scaling for Video Generati...
2025.03.24 | 多智能体协作提升性能,苏格拉底式对话优化提示。 24.03.2025 11:24
本期的 15 篇论文如下: [00:22] 🧠 MAPS: A Multi-Agent Framework Based on Big Seven Personality and Socratic Guidance for Multimodal Scientific Problem Solving(MAPS:一个基于大七人格和苏格拉底指导的多智能体框架,用于多模态科学问题求解) [01:09] 🤖 MARS: A Multi-Agent Framework Incorporating Socratic Guidance for Automated Prompt Optimization(MARS:一个融合苏格拉底式指导的多智能体自动提示优化框架...
【周末特辑】3月第3周最火AI论文 | 序列建模创新,视频渲染突破 22.03.2025 12:56
本期的 5 篇论文如下: [00:37] TOP1(🔥118) | 🦢 RWKV-7 "Goose" with Expressive Dynamic State Evolution(RWKV-7 "Goose":具有表达性动态状态演化的序列建模) [02:36] TOP2(🔥115) | 🎥 ReCamMaster: Camera-Controlled Generative Rendering from A Single Video(ReCamMaster:基于单视频的相机控制生成式渲染) [05:18] TOP3(🔥89) | 🤖 DAPO: An Open-Source LLM Reinforcement Learning System at Scale(DAPO:一个大...
2025.03.21 | 蒸馏提升超分辨率效率,优化推理减少计算负担。 21.03.2025 10:55
本期的 15 篇论文如下: [00:23] 🖼 One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation(基于蒸馏的单步残差转移扩散超分辨率) [01:01] 🤔 Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models(停止过度思考:大型语言模型高效推理综述) [01:38] 🚀 Unleashing Vecset Diffusion Model for Fast Shape Generation(释放Vecset扩散模型以实现快速形状生成) [02:18]...
2025.03.20 | 自适应前瞻采样优化推理;强化学习提升3D网格质量 20.03.2025 10:59
本期的 15 篇论文如下: [00:23] 🔍 $φ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation($\phi$-解码:用于平衡推理时探索与利用的自适应前瞻采样) [01:08] 🎨 DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning(DeepMesh:基于强化学习的自回归艺术家网格创建) [01:51] 🌷 TULIP: Towards Unified Language-Image Pretraining(TULIP:迈向统...
2025.03.19 | 动态序列建模优势,视频生成理解挑战 19.03.2025 10:54
本期的 15 篇论文如下: [00:21] 🦢 RWKV-7 "Goose" with Expressive Dynamic State Evolution(RWKV-7 "Goose":具有表达性动态状态演化的序列建模) [00:55] 🤯 Impossible Videos(不可能的视频) [01:38] 🎨 Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM(Creation-MMBench:评估多模态大型语言模型中具有上下文感知能力的创造性智能) [02:17] 🤖 DAPO: An Open-Source LLM Reinforcement Learn...
2025.03.18 | 视频生成新方法,人形机器人新框架 18.03.2025 10:54
本期的 15 篇论文如下: [00:21] 🎥 DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation(DropletVideo:探索整体时空一致性视频生成的数据集与方法) [01:10] 🤖 Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills(Being-0:一个具有视觉-语言模型和模块化技能的人形机器人代理) [01:49] 🖼 DreamRenderer: Taming Multi-Instance Attrib...
2025.03.17 | 新相机轨迹生成,稀疏性提升图像质量 17.03.2025 11:08
本期的 15 篇论文如下: [00:25] 🎥 ReCamMaster: Camera-Controlled Generative Rendering from A Single Video(ReCamMaster:基于单视频的相机控制生成式渲染) [01:11] 💡 PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity(PLADIS:通过利用稀疏性,在扩散模型推理时突破注意力机制的限制) [01:50] 🤖 Adversarial Data Collection: Human-Collaborative Perturbatio...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.