duan
HuggingFace 每日AI论文速递
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2025.09.01 | R-4B模型优化思考效率;EO-1提升机器人控制能力 01.09.2025 7:57
本期的 15 篇论文如下: [00:24] 🧠 R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning(R-4B: 通过双模式退火和强化学习激励多模态大语言模型的通用自动思考能力) [00:59] 🤖 EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control(具身一体视觉:交错视觉-文本-动作预训练用于通用机器人控制) [01:29] 🔒 A.S.E: A...
【月末特辑】8月最火AI论文 | 科学AI模型缩小性能差距;图像模型解决文本渲染与编辑 31.08.2025 13:04
本期的 10 篇论文如下: [00:30] TOP1(🔥242) | 🧪 Intern-S1: A Scientific Multimodal Foundation Model(Intern-S1:一个科学多模态基础模型) [01:36] TOP2(🔥239) | 🎨 Qwen-Image Technical Report(Qwen-Image技术报告) [02:46] TOP3(🔥227) | 🤔 Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens(LLM思维链推理是海市蜃楼吗?一个数据分布的视角) [04:14] TOP4(🔥220) | 🚀 DINOv3(DINOv3:...
【周末特辑】8月第5周最火AI论文 | 多模态模型效率提升;自博弈策略提高多样性 30.08.2025 6:47
本期的 5 篇论文如下: [00:36] TOP1(🔥161) | 🚀 InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency(InternVL3.5:提升开源多模态模型在通用性、推理能力和效率上的表现) [01:25] TOP2(🔥114) | 📈 Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR(超越Pass@1:变分问题合成的自博弈策略持续提升RLVR) [02:23] TOP3(🔥108) | 🚀 AgentFly: Fine-...
2025.08.29 | 稳定文本到图像生成;高效数学推理 29.08.2025 8:01
本期的 15 篇论文如下: [00:24] ⚖ Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning(Pref-GRPO:基于成对偏好奖励的GRPO用于稳定的文本到图像强化学习) [00:57] 🧠 rStar2-Agent: Agentic Reasoning Technical Report(rStar2-Agent:智能体推理技术报告) [01:28] 🎨 USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning(USO: 通过解...
2025.08.28 | 推理分解减幻觉;可解释性编码信息 28.08.2025 7:15
本期的 14 篇论文如下: [00:25] 🧠 Self-Rewarding Vision-Language Model via Reasoning Decomposition(通过推理分解的自奖励视觉语言模型) [00:49] 🔍 Beyond Transcription: Mechanistic Interpretability in ASR(超越转录:自动语音识别中的机械可解释性) [01:22] 🤖 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies(离散扩散VLA:将离散扩散引入视觉-语言...
2025.08.27 | 物理模型评估显不足;树算法优化提效降本 27.08.2025 7:29
本期的 15 篇论文如下: [00:23] 🔬 CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics(CMPhysBench:用于评估凝聚态物理中大语言模型的基准测试) [00:57] 🌳 TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling(TreePO: 通过启发式树建模弥合策略优化与效果和推理效率之间的差距) [01:21] 🗣 VibeVoice T...
2025.08.26 | 提升模型推理效率;增强生成语义对齐 26.08.2025 7:23
本期的 15 篇论文如下: [00:24] 🚀 InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency(InternVL3.5:提升开源多模态模型在通用性、推理能力和效率上的表现) [00:52] 🧠 Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation(Visual-CoG:阶段感知强化学习与指导链用于文本到图像生成) [01:19] 🎨 MV-RAG: Retrieval Augment...
2025.08.25 | 无微调智能体高效学习;四足机器人长周期探索 25.08.2025 8:02
本期的 15 篇论文如下: [00:23] 🚀 AgentFly: Fine-tuning LLM Agents without Fine-tuning LLMs(AgentFly:无需微调LLM即可微调LLM智能体) [00:48] 🐕 ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks(ODYSSEY:开放世界四足机器人长周期任务探索与操作) [01:24] 📈 Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR(超越Pass@1:变分问题合成的自博弈策...
【周末特辑】8月第4周最火AI论文 | 视觉模型新突破;科学多模态领先 24.08.2025 8:00
本期的 5 篇论文如下: [00:39] TOP1(🔥172) | 🚀 DINOv3(DINOv3:视觉基础模型新里程碑) [01:39] TOP2(🔥170) | 🧪 Intern-S1: A Scientific Multimodal Foundation Model(Intern-S1:一个科学多模态基础模型) [03:08] TOP3(🔥100) | 🤖 Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL(智能体链:基于多智能体蒸馏与智能体强化学习的端到端智能体基础模型) [04:18] TOP...
2025.08.22 | 科学多模态缩小差距;GUI自动化解决挑战 23.08.2025 7:23
本期的 15 篇论文如下: [00:22] 🧪 Intern-S1: A Scientific Multimodal Foundation Model(Intern-S1:一个科学多模态基础模型) [00:46] 🤖 Mobile-Agent-v3: Foundamental Agents for GUI Automation(Mobile-Agent-v3:GUI自动化基础智能体) [01:10] ✅ Deep Think with Confidence(置信深思) [01:31] 🤔 LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries(LiveMCP-101:在挑战性查...
2025.08.21 | 金融大模型认知诊断;DuPO优化自验证 22.08.2025 8:12
本期的 15 篇论文如下: [00:22] 🧠 From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models(从分数到技能:金融大语言模型认知诊断评估框架) [00:49] ✅ DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization(DuPO:通过双重偏好优化实现大模型可靠自验证) [01:17] 🔮 FutureX: An Advanced Live Benchmark for LLM Agents in Future Predicti...
2025.08.20 | 智能体链提升效率;长视频3D重建优化 21.08.2025 7:13
本期的 15 篇论文如下: [00:23] 🤖 Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL(智能体链:基于多智能体蒸馏与智能体强化学习的端到端智能体基础模型) [00:52] 🎥 LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos(LongSplat:针对随意长视频的鲁棒无姿态3D高斯泼溅) [01:13] 🛠 Prompt Orchestration Markup Language(提示编排标记语言) [0...
2025.08.19 | Ovis2.5提升多模态;ComoRAG优化长叙事推理 20.08.2025 8:19
本期的 15 篇论文如下: [00:20] ✨ Ovis2.5 Technical Report(Ovis2.5 技术报告) [00:51] 🧠 ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning(ComoRAG:一种认知启发式记忆组织RAG,用于有状态长叙事推理) [01:14] 🎥 4DNeX: Feed-Forward 4D Generative Modeling Made Easy(4DNeX:前馈4D生成建模轻松实现) [01:38] ✨ Next Visual Granularity Generation(下一视觉粒度生成...
2025.08.18 | 超越图像思考;自搜索强化 18.08.2025 6:44
本期的 13 篇论文如下: [00:19] 💡 Thyme: Think Beyond Images(Thyme:超越图像的思考) [00:48] 🧠 SSRL: Self-Search Reinforcement Learning(SSRL:自搜索强化学习) [01:16] 🚀 DINOv3(DINOv3:视觉基础模型新里程碑) [01:42] 🔍 PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing(PaperRegister:通过分层寄存器索引提升灵活粒度论文搜索) [02:13] 🚀 XQuant: Breaking the...
【周末特辑】8月第3周最火AI论文 | GLM-4.5统一智能体推理编程;We-Math提升视觉数学推理 17.08.2025 6:30
本期的 5 篇论文如下: [00:32] TOP1(🔥139) | 🚀 GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models(GLM-4.5:智能体、推理与编程(ARC)基础模型) [01:44] TOP2(🔥121) | 📚 We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning(We-Math 2.0:一个激励视觉数学推理的多功能数学手册系统) [02:46] TOP3(🔥110) | 🧠 ReasonRank: Empowering Passage Ranking with Str...
2025.08.15 | 数学推理手册提升模型能力;连续令牌生成图像模型 16.08.2025 6:29
本期的 12 篇论文如下: [00:23] 📚 We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning(We-Math 2.0:一个激励视觉数学推理的多功能数学手册系统) [00:50] 🚀 NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale(NextStep-1:迈向大规模连续令牌自回归图像生成) [01:17] 🎨 ToonComposer: Streamlining Cartoon Production with Generative Post-...
2025.08.14 | 分子推理框架提升性能;视频身份控制轻量高效 14.08.2025 7:05
本期的 15 篇论文如下: [00:17] 🧪 Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery(Mol-R1:迈向分子发现中的显式长链思维推理) [00:38] ✨ Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation(Stand-In:视频生成中轻量级即插即用的身份控制) [01:06] 🎥 Story2Board: A Training-Free Approach for Expressive Storyboard Generation(Story2Board:一种富有表现力的...
2025.08.13 | 多模态AI突破;3D世界生成 13.08.2025 6:53
本期的 15 篇论文如下: [00:22] 🤖 WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent(WebWatcher:突破视觉-语言深度研究智能体的新前沿) [00:45] 🌎 Matrix-3D: Omnidirectional Explorable 3D World Generation(Matrix-3D:全向可探索三维世界生成) [01:17] 🚀 Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL(超越十回合:通过大规模异步强化学习...
2025.08.12 | ReasonRank提升段落排序推理;WideSearch评估智能体广域搜寻 13.08.2025 7:15
本期的 15 篇论文如下: [00:18] 🧠 ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability(ReasonRank:赋予段落排序强大推理能力) [00:41] 🔍 WideSearch: Benchmarking Agentic Broad Info-Seeking(WideSearch:智能体广域信息搜寻基准测试) [01:01] ✨ Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation(Omni-Effects:统一且空间可控的视觉效果生成) [01:26] 🧠 Klear-Rea...
2025.08.11 | GLM-4.5统一智能体推理编程;Voost高保真虚拟试穿试脱 12.08.2025 5:09
本期的 11 篇论文如下: [00:20] 🚀 GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models(GLM-4.5:智能体、推理与编程(ARC)基础模型) [00:47] 👕 Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off(Voost:一种统一且可扩展的双向虚拟试穿与试脱扩散Transformer) [01:11] 🎯 InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimi...
【周末特辑】8月第2周最火AI论文 | CoT推理是幻象;Qwen-Image渲染领先 10.08.2025 6:28
本期的 5 篇论文如下: [00:33] TOP1(🔥174) | 🤔 Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens(LLM思维链推理是海市蜃楼吗?一个数据分布的视角) [01:46] TOP2(🔥158) | 🎨 Qwen-Image Technical Report(Qwen-Image技术报告) [02:59] TOP3(🔥127) | 🤖 VeriGUI: Verifiable Long-Chain GUI Dataset(VeriGUI:可验证的长链GUI数据集) [04:09] TOP4(🔥100) | 🚀 Seed Diffusion: A Large-Scale...
2025.08.08 | 动态微调优推理;零数据自演进强推理 09.08.2025 6:35
本期的 15 篇论文如下: [00:16] ✨ On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification(关于SFT泛化性的研究:一个基于奖励修正的强化学习视角) [00:41] 🌱 R-Zero: Self-Evolving Reasoning LLM from Zero Data(R-Zero:零数据自演进推理大语言模型) [01:00] 🤖 Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation(Genie Envisioner:一个用于...
2025.08.07 | VeriGUI提升代理能力;CoT推理实为模式匹配 07.08.2025 5:47
本期的 13 篇论文如下: [00:20] 🤖 VeriGUI: Verifiable Long-Chain GUI Dataset(VeriGUI:可验证的长链GUI数据集) [00:40] 🤔 Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens(LLM思维链推理是海市蜃楼吗?一个数据分布的视角) [00:59] 💰 Efficient Agents: Building Effective Agents While Reducing Cost(高效智能体:在降低成本的同时构建有效智能体) [01:21] 🌱 SEAgent: Self-Evolving C...
2025.08.06 | 高速推理扩散模型;紧凑视觉生成模型 07.08.2025 6:20
本期的 13 篇论文如下: [00:17] 🚀 Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference(种子扩散:一种具有高速推理能力的大规模扩散语言模型) [00:39] 🎨 Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation(Skywork UniPic:用于视觉理解与生成的统一自回归建模) [01:05] 🎥 LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation...
2025.08.05 | 图像文本渲染编辑创新;上下文检索提升故事理解 06.08.2025 7:28
本期的 15 篇论文如下: [00:18] 🎨 Qwen-Image Technical Report(Qwen-Image技术报告) [00:39] 🔍 SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension(SitEmb-v1.5:改进的上下文感知密集检索用于语义关联与长故事理解) [01:08] 🧬 CellForge: Agentic Design of Virtual Cell Models(CellForge: 虚拟细胞模型的智能体设计) [01:36] 🧠 Beyond the Trade-off: Se...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.