Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles 04.06.2025 25:00
🤗 Upvotes: 24 | cs. CV, cs. AI, cs. CL Authors: Zifu Wang, Junyi Zhu, Bo Tang, Zhiyu Li, Feiyu Xiong, Jiaqian Yu, Matthew B. Blaschko Title: Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles Arxiv: http://arxiv.org/abs/2505.23590v2 Abstract: The application of rule-based reinforcement learning (RL) to multimodal large language models (MLLMs) introduces unique chal...
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding 04.06.2025 22:06
🤗 Upvotes: 23 | cs. CV Authors: Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu Title: ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding Arxiv: http://arxiv.org/abs/2506.01853v1 Abstract: Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabi...
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning 04.06.2025 20:39
🤗 Upvotes: 21 | cs. CL Authors: Zhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang, Dongfei Cui, Qinjian Zhao, Hui Shen, Jing Xiong, Yi Xin, Yifan Jiang, Yangfan He, Mi Zhang, Shen Yan Title: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning Arxiv: http://arxiv.org/abs/2506.01713v1 Abstract: Multimodal large language models (MLLMs) have shown promising capabilities in...
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 03.06.2025 21:28
🤗 Upvotes: 83 | cs. CL, cs. AI Authors: Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, Yi Dong Title: ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Arxiv: http://arxiv.org/abs/2505.24864v1 Abstract: Recent advances in reasoning-centric language models have highlighted reinforcement learning (RL) as a promising method...
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time 03.06.2025 20:57
🤗 Upvotes: 63 | cs. CL Authors: Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, Huan Zhang Title: AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time Arxiv: http://arxiv.org/abs/2505.24863v1 Abstract: This paper presents AlphaOne ($\alpha$1), a universal framework for modulating reasoning progress in large r...
Time Blindness: Why Video-Language Models Can't See What Humans Can? 03.06.2025 22:28
🤗 Upvotes: 59 | cs. CV, cs. AI Authors: Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed Elhoseiny Title: Time Blindness: Why Video-Language Models Can't See What Humans Can? Arxiv: http://arxiv.org/abs/2505.24867v1 Abstract: Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. However, when spatial informatio...
HardTests: Synthesizing High-Quality Test Cases for LLM Coding 03.06.2025 21:40
🤗 Upvotes: 37 | cs. CL Authors: Zhongmou He, Yee Man Choi, Kexun Zhang, Jiabao Ji, Junting Zhou, Dejia Xu, Ivan Bercovich, Aidan Zhang, Lei Li Title: HardTests: Synthesizing High-Quality Test Cases for LLM Coding Arxiv: http://arxiv.org/abs/2505.24098v1 Abstract: Verifiers play a crucial role in large language model (LLM) reasoning, needed by post-training techniques such as reinforcement learnin...
Large Language Models for Data Synthesis 03.06.2025 22:32
🤗 Upvotes: 36 | cs. LG Authors: Yihong Tang, Menglin Kong, Lijun Sun Title: Large Language Models for Data Synthesis Arxiv: http://arxiv.org/abs/2505.14752v1 Abstract: Generating synthetic data that faithfully captures the statistical structure of real-world distributions is a fundamental challenge in data modeling. Classical approaches often depend on strong parametric assumptions or manual stru...
Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation 03.06.2025 22:03
🤗 Upvotes: 29 | cs. CL, cs. CV Authors: Jiwan Chung, Junhyeok Kim, Siyeol Kim, Jaeyoung Lee, Min Soo Kim, Youngjae Yu Title: Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation Arxiv: http://arxiv.org/abs/2505.18842v1 Abstract: We present v1, a lightweight extension to Multimodal Large Language Models (MLLMs) that enables selective visual revisitation...
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization 03.06.2025 20:55
🤗 Upvotes: 27 | cs. CV Authors: Cailin Zhuang, Ailin Huang, Wei Cheng, Jingwei Wu, Yaoqi Hu, Jiaqi Liao, Zhewei Huang, Hongyuan Wang, Xinyao Liao, Weiwei Cai, Hengyuan Xu, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Chi Zhang Title: ViStoryBench: Comprehensive Benchmark Suite for Story Visualization Arxiv: http://arxiv.org/abs/2505.24862v1 Abstract: Story visualization, which aims to generate a seque...
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models 03.06.2025 22:45
🤗 Upvotes: 21 | cs. CV, cs. AI Authors: Chenbin Pan, Wenbin He, Zhengzhong Tu, Liu Ren Title: DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Arxiv: http://arxiv.org/abs/2505.24025v1 Abstract: The recent explosive interest in the reasoning capabilities of large language models, such as DeepSeek-R1, has demonstrated remarkable success through reinforcement learning-based fi...
Table-R1: Inference-Time Scaling for Table Reasoning 31.05.2025 21:28
🤗 Upvotes: 66 | cs. CL Authors: Zheyuan Yang, Lyuhao Chen, Arman Cohan, Yilun Zhao Title: Table-R1: Inference-Time Scaling for Table Reasoning Arxiv: http://arxiv.org/abs/2505.23621v1 Abstract: In this work, we present the first study to explore inference-time scaling on table reasoning tasks. We develop and evaluate two post-training strategies to enable inference-time scaling: distillation from...
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence 31.05.2025 19:54
🤗 Upvotes: 54 | cs. CV, cs. AI, cs. LG, I.2.6; I.2 Authors: Diankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi Duan Title: Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Arxiv: http://arxiv.org/abs/2505.23747v1 Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spati...
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 31.05.2025 25:44
🤗 Upvotes: 51 | cs. CV, cs. AI, cs. CL Authors: Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao Title: VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Arxiv: http://arxiv.org/abs/2505.23693v1 Abstract: MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-ge...
The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason 31.05.2025 22:04
🤗 Upvotes: 45 | cs. CL Authors: Ang Lv, Ruobing Xie, Xingwu Sun, Zhanhui Kang, Rui Yan Title: The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Arxiv: http://arxiv.org/abs/2505.22653v1 Abstract: Recent studies on post-training large language models (LLMs) for reasoning through reinforcement learning (RL) typically focus on tasks that can be accurately veri...
ZeroGUI: Automating Online GUI Learning at Zero Human Cost 31.05.2025 19:00
🤗 Upvotes: 39 | cs. AI, cs. CL, cs. CV Authors: Chenyu Yang, Shiqian Su, Shi Liu, Xuan Dong, Yue Yu, Weijie Su, Xuehui Wang, Zhaoyang Liu, Jinguo Zhu, Hao Li, Wenhai Wang, Yu Qiao, Xizhou Zhu, Jifeng Dai Title: ZeroGUI: Automating Online GUI Learning at Zero Human Cost Arxiv: http://arxiv.org/abs/2505.23762v1 Abstract: The rapid advancement of large Vision-Language Models (VLMs) has propelled the...
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? 31.05.2025 21:32
🤗 Upvotes: 28 | cs. CV Authors: Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, Lin Sui, Xinhao Li, Yan Zhong, Y. Charles, Xinyu Zhou, Xu Sun Title: VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? Arxiv: http://arxiv.org/abs/2505.23359v1 Abstract: Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large langua...
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering 31.05.2025 21:32
🤗 Upvotes: 21 | cs. CL, cs. AI, cs. SE Authors: Guangtao Zeng, Maohao Shen, Delin Chen, Zhenting Qi, Subhro Das, Dan Gutfreund, David Cox, Gregory Wornell, Wei Lu, Zhang-Wei Hong, Chuang Gan Title: Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering Arxiv: http://arxiv.org/abs/2505.23604v1 Abstract: Language models (LMs) perform well on standardized coding benchma...
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 30.05.2025 22:08
🤗 Upvotes: 84 | cs. LG, cs. AI, cs. CL Authors: Ganqu Cui, Yuchen Zhang, Jiacheng Chen, Lifan Yuan, Zhi Wang, Yuxin Zuo, Haozhan Li, Yuchen Fan, Huayu Chen, Weize Chen, Zhiyuan Liu, Hao Peng, Lei Bai, Wanli Ouyang, Yu Cheng, Bowen Zhou, Ning Ding Title: The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models Arxiv: http://arxiv.org/abs/2505.22617v1 Abstract: This paper aims...
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents 30.05.2025 21:03
🤗 Upvotes: 63 | cs. SE, cs. CL Authors: Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, Simon Karasik, Andrei Andriushchenko, Maria Trofimova, Daria Litvintseva, Boris Yangel Title: SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents Arxiv: http://arxiv.org/abs/2505.20411v1 Abstract: LLM-based agents have...
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing 30.05.2025 23:01
🤗 Upvotes: 59 | cs. CL, cs. AI, cs. LG, cs. PF, I.2.7 Authors: Tianyu Fu, Yi Ge, Yichen You, Enshu Liu, Zhihang Yuan, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang Title: R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing Arxiv: http://arxiv.org/abs/2505.21600v1 Abstract: Large Language Models (LLMs) achieve impressive reasoning capabilities at the cost o...
Skywork Open Reasoner 1 Technical Report 30.05.2025 22:21
🤗 Upvotes: 45 | cs. LG, cs. AI, cs. CL Authors: Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, Siyuan Li, Liang Zeng, Tianwen Wei, Cheng Cheng, Bo An, Yang Liu, Yahui Zhou Title: Skywork Open Reasoner 1 Technical Report Arxiv: http://arxiv.org/abs/2505.22312v2 Abstract: The success of DeepSeek-R1 underscores the signif...
Sherlock: Self-Correcting Reasoning in Vision-Language Models 30.05.2025 21:23
🤗 Upvotes: 44 | cs. CV, cs. CL, cs. LG Authors: Yi Ding, Ruqi Zhang Title: Sherlock: Self-Correcting Reasoning in Vision-Language Models Arxiv: http://arxiv.org/abs/2505.22651v1 Abstract: Reasoning Vision-Language Models (VLMs) have shown promising performance on complex multimodal tasks. However, they still face significant challenges: they are highly sensitive to reasoning errors, require large...
Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO 30.05.2025 22:53
🤗 Upvotes: 37 | cs. CL, cs. AI, cs. CV, cs. LG Authors: Lai Wei, Yuting Li, Chen Wang, Yue Wang, Linghe Kong, Weiran Huang, Lichao Sun Title: Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO Arxiv: http://arxiv.org/abs/2505.22453v1 Abstract: Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforce...
SageAttention2++: A More Efficient Implementation of SageAttention2 30.05.2025 19:42
🤗 Upvotes: 33 | cs. LG, cs. AI, cs. AR, cs. CV Authors: Jintao Zhang, Xiaoming Xu, Jia Wei, Haofeng Huang, Pengle Zhang, Chendong Xiang, Jun Zhu, Jianfei Chen Title: SageAttention2++: A More Efficient Implementation of SageAttention2 Arxiv: http://arxiv.org/abs/2505.21136v2 Abstract: The efficiency of attention is critical because its time complexity grows quadratically with sequence length. Sage...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.