duan

HuggingFace 每日AI论文速递

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】

Author

duan

Category

Technology

Podcast website

www.xiaoyuzhoufm.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

2024.10.24 每日AI论文 | 多图像任务优化,视频生成模型评估 24.10.2024

本期的 10 篇论文如下: [00:25] 🖼 MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models(多图像增强直接偏好优化:大型视觉语言模型) [01:09] 🌍 WorldSimBench: Towards Video Generation Models as World Simulators(世界模拟器:迈向视频生成模型作为世界模拟器) [01:47] 🌊 Scaling Diffusion Language Models via Adaptation from Autoregressive Models(通过自回归模型适...

2024.10.23 每日AI论文 | 视觉冗余减少提升效率,动态三维重建优化镜面场景。 23.10.2024

本期的 8 篇论文如下: [00:27] 🔍 PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction(金字塔式视觉冗余减少:通过金字塔视觉冗余减少加速大型视觉-语言模型) [01:09] 🌟 SpectroMotion: Dynamic 3D Reconstruction of Specular Scenes(光谱运动:镜面场景的动态三维重建) [01:48] 🤖 Aligning Large Language Models via Self-Steering Optimization(通过自引导优化对...

2024.10.22 每日AI论文 | 指南针评判者加速模型评估,SAM2Long提升长视频分割精度。 22.10.2024

本期的 21 篇论文如下: [00:24] 🤖 CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution(指南针评判者-1:一体化评判模型助力模型评估与进化) [01:11] 🌲 SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree(SAM2长:通过无需训练的记忆树增强SAM 2以实现长视频分割) [01:55] 🌐 PUMA: Empowering Unified MLLM with Multi-granular Visual Generation(PU...

2024.10.21 每日AI论文 | 提升网页导航成功率,增强图像生成精细度。 21.10.2024

本期的 12 篇论文如下: [00:27] 🌐 Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation(拥有世界模型的网络代理:学习和利用环境动态进行网页导航) [01:11] 👗 MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models(魔法裁缝:文本到图像扩散模型中的组件可控个性化) [01:48] 💼 UCFE: A User-Centric Financial Expertise Benchmark for La...

【周末特辑】10月第3周最火AI论文 | 多模态大语言模型创新,评估标准统一化。 19.10.2024

本期的 5 篇论文如下: [00:45] TOP1(🔥80) | 🌐 Baichuan-Omni Technical Report(百川-Omni 技术报告) [02:20] TOP2(🔥58) | 📊 MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures(MixEval-X:从现实世界数据混合中进行任意到任意评估) [04:20] TOP3(🔥58) | 🎥 Movie Gen: A Cast of Media Foundation Models(电影生成:媒体基础模型集合) [06:27] TOP4(🔥53) | 🤖 LOKI: A Comprehensive Synthetic Data...

2024.10.18 每日AI论文 | AI评估标准化,电影生成模型领先。 18.10.2024

本期的 31 篇论文如下: [00:23] 📊 MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures(MixEval-X:从现实世界数据混合中进行任意到任意评估) [01:02] 🎥 Movie Gen: A Cast of Media Foundation Models(电影生成:媒体基础模型集合) [01:35] 📱 MobA: A Two-Level Agent System for Efficient Mobile Task Automation(MobA:一种高效移动任务自动化的两级代理系统) [02:18] 🌐 Harnessing Webpage UIs for...

2024.10.17 每日AI论文 | 视觉推理能力待提升,自中心视频理解需改进 17.10.2024

本期的 19 篇论文如下: [00:28] 🧠 HumanEval-V: Evaluating Visual Understanding and Reasoning Abilities of Large Multimodal Models Through Coding Tasks(HumanEval-V:通过编码任务评估大型多模态模型的视觉理解和推理能力) [01:15] 🎥 VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI(VidEgoThink:评估具身AI的自中心视频理解能力) [01:50] 🧠 The Curse of Multi-Modalities:...

2024.10.16 每日AI论文 | 多模态模型幻觉问题,工具使用基准评估。 16.10.2024

本期的 14 篇论文如下: [00:26] 🤖 MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation(多模态大语言模型能看见吗?动态校正解码以减轻幻觉) [01:07] 🛠 MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models(MTU-Bench:大型语言模型的多粒度工具使用基准) [01:47] 📚 LLM$\times$MapReduce: Simplified Long-Sequence Processing using Large Language Models(LLM×MapReduc...

2024.10.15 每日AI论文 | MMIE推动LVLMs发展,LOKI评估合成数据检测。 15.10.2024

本期的 15 篇论文如下: [00:24] 🌐 MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models(大规模多模态交错理解基准测试) [01:06] 🤖 LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models(LOKI:基于大型多模态模型的综合合成数据检测基准) [02:01] 🔍 Toward General Instruction-Following Alignment for Retrieval-Augmented Generation...

2024.10.14 每日AI论文 | 多模态模型Baichuan-Omni开源,Meissonic提升文生图效率 14.10.2024

本期的 16 篇论文如下: [00:25] 🌐 Baichuan-Omni Technical Report(百川-Omni 技术报告) [00:59] 🖼 Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis(Meissonic:高效高分辨率文本到图像生成的掩码生成Transformer复兴) [01:41] 🔧 From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning(从通才到...

【周末特辑】10月第2周最火AI论文 | 差分Transformer提升文本处理,L-Mul算法降低能耗。 12.10.2024

本期的 5 篇论文如下: [00:37] TOP1(🔥128) | 🔍 Differential Transformer(差分Transformer) [02:38] TOP2(🔥125) | ⚡ Addition is All You Need for Energy-efficient Language Models(加法即所需:高效能语言模型) [04:13] TOP3(🔥84) | 🌐 Aria: An Open Multimodal Native Mixture-of-Experts Model(Aria:一个开放的多模态原生混合专家模型) [06:18] TOP4(🔥73) | 🤖 GLEE: A Unified Framework and Benchmark for L...

2024.10.11 每日AI论文 | 数学代码提升推理,前缀量化加速模型 11.10.2024

本期的 21 篇论文如下: [00:25] 🧮 MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code(MathCoder2:通过模型翻译的数学代码进行持续预训练以提升数学推理能力) [01:09] 🚀 PrefixQuant: Static Quantization Beats Dynamic through Prefixed Outliers in LLMs(前缀量化:静态量化通过LLMs中的前缀异常值超越动态量化) [01:59] 🤖 MLLM as Retriever: Interactively Learn...

【月末特辑】9月最火AI论文 | 强化学习提升语言模型,代码智能模型表现优异。 11.10.2024

本期的 10 篇论文如下: [00:40] TOP1(🔥129) | 🤖 Training Language Models to Self-Correct via Reinforcement Learning(通过强化学习训练语言模型进行自我修正) [02:41] TOP2(🔥121) | 🚀 Qwen2.5-Coder Technical Report(Qwen2.5-Coder技术报告) [04:44] TOP3(🔥96) | 🌐 Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models(Molmo 和 PixMo:用于最先进多模态模型的开放权重和开放数...

2024.10.10 每日AI论文 | LLMs经济游戏表现各异,个性化视觉指令提升AI互动。 10.10.2024

本期的 43 篇论文如下: [00:23] 🤖 GLEE: A Unified Framework and Benchmark for Language-based Economic Environments(GLEE:基于语言的经济环境统一框架与基准) [01:09] 👤 Personalized Visual Instruction Tuning(个性化视觉指令微调) [01:48] 🌍 Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation(迈向世界模拟器:基于物理常识的视频生成基准) [02:35] 🖼 IterComp: It...

2024.10.09 每日AI论文 | 长上下文生成能力评估,指令多样性影响泛化 09.10.2024

本期的 9 篇论文如下: [00:28] 📚 LongGenBench: Long-context Generation Benchmark(长上下文生成基准:LongGenBench) [01:11] 🌐 $\textbf{Only-IF}$:Revealing the Decisive Effect of Instruction Diversity on Generalization(仅限IF:揭示指令多样性对泛化的决定性影响) [01:50] 📊 RevisEval: Improving LLM-as-a-Judge via Response-Adapted References(RevisEval:通过响应自适应参考改进LLM作为评判者) [02:35]...

2024.10.08 每日AI论文 | 差分Transformer优化注意力,LLM幻觉研究揭示错误模式。 08.10.2024

本期的 21 篇论文如下: [00:26] 🔍 Differential Transformer(差分Transformer) [01:04] 🧠 LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations(大语言模型知多于表:关于LLM幻觉的内在表征) [01:50] 📹 VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide(视频指南:通过教师指导提升视频扩散模型无需训练) [02:28] 📈 FAN: Fourier Analysis...

2024.10.07 每日AI论文 | 高效能语言模型节能新算法,视觉语言模型推理能力待提升。 07.10.2024

本期的 12 篇论文如下: [00:25] ⚡ Addition is All You Need for Energy-efficient Language Models(加法即所需:高效能语言模型) [01:03] 🧠 NL-Eye: Abductive NLI for Images(NL-Eye:图像的溯因自然语言推理) [01:40] 🔍 Selective Attention Improves Transformer(选择性注意力提升Transformer) [02:17] ⚡ Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding(...

【周末特辑】10月第1周最火AI论文 | Emu3模型表现卓越,最弱环节定律制约LLMs。 05.10.2024

本期的 5 篇论文如下: [00:47] TOP1(🔥73) | 🧠 Emu3: Next-Token Prediction is All You Need(Emu3:下一个词预测是所有你需要的) [02:42] TOP2(🔥48) | 🔗 Law of the Weakest Link: Cross Capabilities of Large Language Models(最弱环节定律:大型语言模型的跨能力) [04:26] TOP3(🔥45) | 🌐 MIO: A Foundation Model on Multimodal Tokens(MIO:基于多模态标记的基础模型) [06:26] TOP4(🔥44) | 🌐 Revisit Large-Sca...

2024.10.04 每日AI论文 | 字幕类型影响模型表现,长视频生成技术突破。 04.10.2024

本期的 19 篇论文如下: [00:24] 🔄 Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models(重新审视大规模图像-文本数据在多模态基础模型预训练中的作用) [01:04] 🎥 Loong: Generating Minute-level Long Videos with Autoregressive Language Models(使用自回归语言模型生成分钟级长视频) [01:39] 🎥 Video Instruction Tuning With Synthetic Data(使用合成数据进行视频指令调优) [02:1...

2024.10.03 每日AI论文 | 分层调试提升代码准确性,多模态模型优化图像任务。 03.10.2024

本期的 20 篇论文如下: [00:23] 🐞 From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging(从代码到正确性:通过分层调试解决代码生成的最后一步) [01:08] 📄 LEOPARD : A Vision Language Model For Text-Rich Multi-Image Tasks(LEOPARD:用于文本丰富的多图像任务的视觉语言模型) [01:48] 📊 Is Preference Alignment Always the Best Option to Enhance LLM-Based Translatio...

2024.10.02 每日AI论文 | 跨能力任务表现受限,边缘设备高效部署模型 02.10.2024

本期的 13 篇论文如下: [00:26] 🔗 Law of the Weakest Link: Cross Capabilities of Large Language Models(最弱环节法则:大型语言模型的跨能力) [01:05] 🌐 TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices(TPI-LLM:在低资源边缘设备上高效服务70B规模的大型语言模型) [01:46] 🌍 Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect(Atlas-Chat:为低资...

2024.10.01 每日AI论文 | 多模态模型提升图像理解,长度控制方法增强生成精确性。 01.10.2024

本期的 11 篇论文如下: [00:26] 🌐 MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning(MM1.5:多模态大语言模型微调的方法、分析与见解) [01:04] 📏 Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models(Ruler:一种用于控制大型语言模型生成长度的模型无关方法) [01:41] 🗣 DiaSynth -- Synthetic Dialogue Generation Framework(DiaSynth -- 合成对话生成框架) [0...

2024.09.30 每日AI论文 | Emu3简化多模态设计,MIO提升视频理解表现。 30.09.2024

本期的 9 篇论文如下: [00:24] 🧠 Emu3: Next-Token Prediction is All You Need(Emu3:下一个词预测是您所需要的全部) [00:53] 🌐 MIO: A Foundation Model on Multimodal Tokens(多模态标记的基础模型:MIO) [01:26] 🔍 VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models(VPTQ:大语言模型的极端低比特向量后训练量化) [02:21] 🎥 PhysGen: Rigid-Body Physics-Grounded Image-to-Vide...

【周末特辑】9月第5周最火AI论文 | 开放权重多模态模型,无调参个性化图像生成。 29.09.2024

本期的 5 篇论文如下: [00:42] TOP1(🔥70) | 🌐 Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models(Molmo 和 PixMo:用于最先进多模态模型的开放权重和开放数据) [02:56] TOP2(🔥64) | 🖼 Imagine yourself: Tuning-Free Personalized Image Generation(想象你自己:无调参个性化图像生成) [05:08] TOP3(🔥48) | 🤖 Programming Every Example: Lifting Pre-training Data Quality like Ex...

2024.09.27 每日AI论文 | 3D感知能力提升,计算开销减少。 27.09.2024

本期的 12 篇论文如下: [00:27] 🌐 LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness(LLaVA-3D:一种简单而有效的路径,赋予多模态模型3D感知能力) [01:10] 🧩 MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models(MaskLLM:大型语言模型的可学习半结构化稀疏性) [01:49] 🎭 EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions(EMOVA:赋予...

Listen to the HuggingFace 每日AI论文速递 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.