Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 11.09.2025 23:21
🤗 Upvotes: 66 | cs. CL Authors: Tong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang, Xinyu Yang, Runpeng Dai, Rui Liu, Huiwen Bao, Chengsong Huang, Heng Huang, Dong Yu Title: Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Arxiv: http://arxiv.org/abs/2509.07980v1 Abstract: Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large langua...
Visual Representation Alignment for Multimodal Large Language Models 11.09.2025 26:13
🤗 Upvotes: 54 | cs. CV Authors: Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim Title: Visual Representation Alignment for Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2509.07979v1 Abstract: Multimodal large language models (MLLMs) trained with visual instr...
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search 11.09.2025 21:52
🤗 Upvotes: 45 | cs. CV, cs. AI, cs. CL Authors: Xin Lai, Junyi Li, Wei Li, Tao Liu, Tianjian Li, Hengshuang Zhao Title: Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Arxiv: http://arxiv.org/abs/2509.07969v1 Abstract: Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing...
Reconstruction Alignment Improves Unified Multimodal Models 11.09.2025 24:13
🤗 Upvotes: 31 | cs. CV, cs. AI, cs. LG Authors: Ji Xie, Trevor Darrell, Luke Zettlemoyer, XuDong Wang Title: Reconstruction Alignment Improves Unified Multimodal Models Arxiv: http://arxiv.org/abs/2509.07295v1 Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training relies on image-text pairs (or sequences) w...
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward 11.09.2025 23:02
🤗 Upvotes: 24 | cs. CV, cs. LG Authors: Yufeng Cheng, Wenxu Wu, Shaojin Wu, Mengqi Huang, Fei Ding, Qian He Title: UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward Arxiv: http://arxiv.org/abs/2509.06818v1 Abstract: Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since w...
Reverse-Engineered Reasoning for Open-Ended Generation 10.09.2025 12:24
🤗 Upvotes: 107 | cs. AI, cs. CL Authors: Haozhe Wang, Haoran Que, Qixin Xu, Minghao Liu, Wangchunshu Zhou, Jiazhan Feng, Wanjun Zhong, Wei Ye, Tong Yang, Wenhao Huang, Ge Zhang, Fangzhen Lin Title: Reverse-Engineered Reasoning for Open-Ended Generation Arxiv: http://arxiv.org/abs/2509.06160v1 Abstract: While the ``deep reasoning'' paradigm has spurred significant advances in verifiable domains li...
Does DINOv3 Set a New Medical Vision Standard? 10.09.2025 10:46
🤗 Upvotes: 28 | cs. CV Authors: Che Liu, Yinda Chen, Haoyuan Shi, Jinpeng Lu, Bailiang Jian, Jiazhen Pan, Linghan Cai, Jiayi Wang, Yundi Zhang, Jun Li, Cosmin I. Bercea, Cheng Ouyang, Chen Chen, Zhiwei Xiong, Benedikt Wiestler, Christian Wachinger, Daniel Rueckert, Wenjia Bai, Rossella Arcucci Title: Does DINOv3 Set a New Medical Vision Standard? Arxiv: http://arxiv.org/abs/2509.06467v1 Abstract:...
Symbolic Graphics Programming with Large Language Models 09.09.2025 13:57
🤗 Upvotes: 31 | cs. CV, cs. LG Authors: Yamei Chen, Haoquan Zhang, Yangyi Huang, Zeju Qiu, Kaipeng Zhang, Yandong Wen, Weiyang Liu Title: Symbolic Graphics Programming with Large Language Models Arxiv: http://arxiv.org/abs/2509.05208v1 Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise visual c...
Set Block Decoding is a Language Model Inference Accelerator 09.09.2025 16:21
🤗 Upvotes: 31 | cs. LG Authors: Itai Gat, Heli Ben-Hamu, Marton Havasi, Daniel Haziza, Jeremy Reizenstein, Gabriel Synnaeve, David Lopez-Paz, Brian Karrer, Yaron Lipman Title: Set Block Decoding is a Language Model Inference Accelerator Arxiv: http://arxiv.org/abs/2509.04185v1 Abstract: Autoregressive next token prediction language models offer powerful capabilities but face significant challenge...
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth 06.09.2025 22:57
🤗 Upvotes: 100 | cs. CL Authors: Yang Wang, Chenghao Xiao, Chia-Yi Hsiao, Zi Yan Chang, Chi-Li Chen, Tyler Loakman, Chenghua Lin Title: Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth Arxiv: http://arxiv.org/abs/2509.03867v1 Abstract: We introduce Drivelology, a unique linguistic phenomenon characterised as "nonsense with depth", utterances that are syntactically coherent yet...
From Editor to Dense Geometry Estimator 06.09.2025 18:45
🤗 Upvotes: 63 | cs. CV, cs. AI Authors: JiYuan Wang, Chunyu Lin, Lei Sun, Rongying Liu, Lang Nie, Mingxing Li, Kang Liao, Xiangxiang Chu, Yao Zhao Title: From Editor to Dense Geometry Estimator Arxiv: http://arxiv.org/abs/2509.04338v1 Abstract: Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inhere...
Towards a Unified View of Large Language Model Post-Training 06.09.2025 23:07
🤗 Upvotes: 42 | cs. LG, cs. AI, cs. CL Authors: Xingtai Lv, Yuxin Zuo, Youbang Sun, Hongyi Liu, Yuntian Wei, Zhekai Chen, Lixuan He, Xuekai Zhu, Kaiyan Zhang, Bingning Wang, Ning Ding, Bowen Zhou Title: Towards a Unified View of Large Language Model Post-Training Arxiv: http://arxiv.org/abs/2509.04419v1 Abstract: Two major sources of training data exist for post-training modern language models: o...
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks 06.09.2025 20:11
🤗 Upvotes: 41 | cs. AI Authors: Haiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu, Jiaxuan Lu, Di Yu, Jianbao Cao, Ben Gao, Jiaqing Xie, Aoran Wang, Wenlong Zhang, Philip Torr, Dongzhan Zhou Title: DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks Arxiv: http://arxiv.org/abs/2509.01396v1 Abstract: Deep research agents have attracted growing attention for their po...
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? 06.09.2025 23:03
🤗 Upvotes: 41 | cs. CL Authors: Qinyan Zhang, Xinping Lei, Ruijie Miao, Yu Fu, Haojie Fan, Le Chang, Jiafan Hou, Dingling Zhang, Zhongfei Hou, Ziqiang Yang, Changxin Pu, Fei Hu, Jingkai Liu, Mengyun Liu, Yang Liu, Xiang Gao, Jiaheng Liu, Tong Yang, Zaiyuan Wang, Ge Zhang, Wenhao Huang Title: Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? Arxiv: http://...
Open Data Synthesis For Deep Research 05.09.2025 23:03
🤗 Upvotes: 37 | cs. CL, cs. AI Authors: Ziyi Xia, Kun Luo, Hongjin Qian, Zheng Liu Title: Open Data Synthesis For Deep Research Arxiv: http://arxiv.org/abs/2509.00375v1 Abstract: Large language models (LLMs) are increasingly expected to go beyond simple factual queries toward Deep Research-tasks that require decomposing questions into sub-problems, coordinating multi-step reasoning, and synthesiz...
Robix: A Unified Model for Robot Interaction, Reasoning and Planning 05.09.2025 21:57
🤗 Upvotes: 33 | cs. AI, cs. CV, cs. RO Authors: Huang Fang, Mengxi Zhang, Heng Dong, Wei Li, Zixuan Wang, Qifeng Zhang, Xueyun Tian, Yucheng Hu, Hang Li Title: Robix: A Unified Model for Robot Interaction, Reasoning and Planning Arxiv: http://arxiv.org/abs/2509.01106v1 Abstract: We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction wi...
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning 04.09.2025 24:16
🤗 Upvotes: 71 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Haoming Wang, Haoyang Zou, Huatong Song, Jiazhan Feng, Junjie Fang, Junting Lu, Longxiang Liu, Qinyu Luo, Shihao Liang, Shijue Huang, Wanjun Zhong, Yining Ye, Yujia Qin, Yuwen Xiong, Yuxin Song, Zhiyong Wu, Bo Li, Chen Dun, Chong Liu, Fuxing Leng, Hanbin Wang, Hao Yu, Haobin Chen, Hongyi Guo, Jing Su, Jingjia Huang, Kai Shen, Kaiyu Shi, Lin...
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model 04.09.2025 23:48
🤗 Upvotes: 63 | cs. CV, cs. LG Authors: Xiyao Wang, Chunyuan Li, Jianwei Yang, Kai Zhang, Bo Liu, Tianyi Xiong, Furong Huang Title: LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Arxiv: http://arxiv.org/abs/2509.00676v1 Abstract: In vision-language modeling, critic models are typically trained to evaluate outputs -- assigning scalar scores or pairwise preferences -- rather t...
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding 04.09.2025 22:32
🤗 Upvotes: 50 | cs. CV, cs. AI Authors: Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu Title: ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding Arxiv: http://arxiv.org/abs/2508.21496v2 Abstract: Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. Howeve...
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion 04.09.2025 20:16
🤗 Upvotes: 39 | cs. CV Authors: Yuan Liu, Zhongyin Zhao, Le Tian, Haicheng Wang, Xubing Ye, Yangxiu You, Zilin Yu, Chuhan Wu, Xiao Zhou, Yang Yu, Jie Zhou Title: POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion Arxiv: http://arxiv.org/abs/2509.01215v1 Abstract: High-quality labeled data is essential for training accurate document conversion models, par...
Baichuan-M2: Scaling Medical Capability with Large Verifier System 04.09.2025 23:34
🤗 Upvotes: 28 | cs. LG, cs. AI Authors: Baichuan-M2 Team, :, Chengfeng Dou, Chong Liu, Fan Yang, Fei Li, Jiyuan Jia, Mingyang Chen, Qiang Ju, Shuai Wang, Shunya Dang, Tianpeng Li, Xiangrong Zeng, Yijie Zhou, Chenzheng Zhu, Da Pan, Fei Deng, Guangwei Ai, Guosheng Dong, Hongda Zhang, Jinyang Tai, Jixiang Hong, Kai Lu, Linzhuang Sun, Peidong Guo, Qian Ma, Rihui Xin, Shihui Yang, Shusen Zhang, Yichua...
Kwai Keye-VL 1.5 Technical Report 04.09.2025 18:19
🤗 Upvotes: 26 | cs. CV Authors: Biao Yang, Bin Wen, Boyang Ding, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Guowang Zhang, Han Shen, Hao Peng, Haojie Ding, Hao Wang, Hengrui Ju, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Muhao Wei, Qiang Wang, Ruitao Wang, Sen Na, Shengn...
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic 04.09.2025 24:18
🤗 Upvotes: 25 | cs. CL Authors: Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Bernard Ghanem Title: Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic Arxiv: http://arxiv.org/abs/2509.01363v1 Abstract: Large language models often require costly optimization, such as reinforcement learning, to master complex reasoning tasks. This work demonstrates that reasoning abili...
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning 03.09.2025 21:59
🤗 Upvotes: 21 | cs. LG, cs. AI Authors: Wenfeng Feng, Penghong Zhao, Guochao Jiang, Chuzhan Hao, Yuewei Zhang, Hao Wang Title: PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning Arxiv: http://arxiv.org/abs/2508.21104v1 Abstract: Critic-free reinforcement learning methods, particularly group policies, have attracted considerable attention for their efficiency in complex task...
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning 02.09.2025 19:58
🤗 Upvotes: 84 | cs. CV, cs. AI, cs. LG Authors: Jie Jiang, Qi Yang, Bolin Ni, Shiming Xiang, Han Hu, Houwen Peng Title: R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Arxiv: http://arxiv.org/abs/2508.21113v1 Abstract: Multimodal Large Language Models (MLLMs) equipped with step-by-step thinking capabilities have demonstrated remar...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.