Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Quartet: Native FP4 Training Can Be Optimal for Large Language Models 27.05.2025 22:46
🤗 Upvotes: 55 | cs. LG Authors: Roberto L. Castro, Andrei Panferov, Soroush Tabesh, Oliver Sieberling, Jiale Chen, Mahdi Nikdan, Saleh Ashkboos, Dan Alistarh Title: Quartet: Native FP4 Training Can Be Optimal for Large Language Models Arxiv: http://arxiv.org/abs/2505.14669v1 Abstract: The rapid advancement of large language models (LLMs) has been paralleled by unprecedented increases in computati...
Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models 27.05.2025 20:36
🤗 Upvotes: 51 | cs. AI Authors: Doohyuk Jang, Yoonjeon Kim, Chanjae Park, Hyun Ryu, Eunho Yang Title: Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models Arxiv: http://arxiv.org/abs/2505.17225v1 Abstract: Large language models have demonstrated remarkable proficiency in long and complex reasoning tasks. However, they frequently exhibit a problematic reliance on fami...
One RL to See Them All: Visual Triple Unified Reinforcement Learning 27.05.2025 20:04
🤗 Upvotes: 51 | cs. CV, cs. CL Authors: Yan Ma, Linge Du, Xuyang Shen, Shaoxiang Chen, Pengfei Li, Qibing Ren, Lizhuang Ma, Yuchao Dai, Pengfei Liu, Junjie Yan Title: One RL to See Them All: Visual Triple Unified Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.18129v1 Abstract: Reinforcement learning (RL) has significantly advanced the reasoning capabilities of vision-language models (VLM...
Distilling LLM Agent into Small Models with Retrieval and Code Tools 27.05.2025 21:30
🤗 Upvotes: 49 | cs. CL, cs. AI Authors: Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho, Sung Ju Hwang Title: Distilling LLM Agent into Small Models with Retrieval and Code Tools Arxiv: http://arxiv.org/abs/2505.17612v1 Abstract: Large language models (LLMs) excel at complex reasoning tasks but remain computationally expensive, limiting their practical deployment. To address this, recent work...
QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization 27.05.2025 22:56
🤗 Upvotes: 39 | cs. CL Authors: Weizhou Shen, Chenliang Li, Fanqi Wan, Shengyi Liao, Shaopeng Lai, Bo Zhang, Yingcheng Shi, Yuning Wu, Gang Fu, Zhansheng Li, Bin Yang, Ji Zhang, Fei Huang, Jingren Zhou, Ming Yan Title: QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization Arxiv: http://arxiv.org/abs/2505.18092v1 Abstract: This technical report presents QwenLong-CPRS, a context co...
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? 27.05.2025 22:03
🤗 Upvotes: 38 | cs. AI Authors: Hui Shen, Taiqiang Wu, Qi Han, Yunta Hsieh, Jizhou Wang, Yuyue Zhang, Yuxin Cheng, Zijian Hao, Yuansheng Ni, Xin Wang, Zhongwei Wan, Kai Zhang, Wendong Xu, Jing Xiong, Ping Luo, Wenhu Chen, Chaofan Tao, Zhuoqing Mao, Ngai Wong Title: PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Arxiv: http://arxiv.org/abs/2505.15929v1 Abstract: Existing benchmarks...
Scaling Image and Video Generation via Test-Time Evolutionary Search 27.05.2025 24:37
🤗 Upvotes: 33 | cs. CV, cs. AI, cs. LG Authors: Haoran He, Jiajun Liang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Ling Pan Title: Scaling Image and Video Generation via Test-Time Evolutionary Search Arxiv: http://arxiv.org/abs/2505.17618v1 Abstract: As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scalin...
MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback 27.05.2025 19:36
🤗 Upvotes: 25 | cs. CL, cs. AI, cs. CE Authors: Wanhao Liu, Zonglin Yang, Jue Wang, Lidong Bing, Di Zhang, Dongzhan Zhou, Yuqiang Li, Houqiang Li, Erik Cambria, Wanli Ouyang Title: MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback Arxiv: http://arxiv.org/abs/2505.17873v1 Abstract: Hypothesis ranking is a crucial component of automated scientific discover...
NovelSeek: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification 24.05.2025 20:49
🤗 Upvotes: 86 | cs. AI, cs. CL, cs. CV Authors: NovelSeek Team, Bo Zhang, Shiyang Feng, Xiangchao Yan, Jiakang Yuan, Zhiyin Yu, Xiaohan He, Songtao Huang, Shaowei Hou, Zheng Nie, Zhilong Wang, Jinyao Liu, Runmin Ma, Tianshuo Peng, Peng Ye, Dongzhan Zhou, Shufei Zhang, Xiaosong Wang, Yilan Zhang, Meng Li, Zhongying Tu, Xiangyu Yue, Wangli Ouyang, Bowen Zhou, Lei Bai Title: NovelSeek: When Agent Be...
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models 24.05.2025 23:44
🤗 Upvotes: 49 | cs. CL, cs. AI Authors: Tingchen Fu, Jiawei Gu, Yafu Li, Xiaoye Qu, Yu Cheng Title: Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Arxiv: http://arxiv.org/abs/2505.14810v1 Abstract: Instruction-following is essential for aligning large language models (LLMs) with user intent. While recent reasoning-oriented models exhibit impressive p...
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning 24.05.2025 21:53
🤗 Upvotes: 43 | cs. CL, cs. AI, cs. LG Authors: Guanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Yutao Zhu, Hangyu Mao, Guorui Zhou, Zhicheng Dou, Ji-Rong Wen Title: Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.16410v1 Abstract: Recently, large language models (LLMs) have shown remarkable reasoning capabilities vi...
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning 24.05.2025 17:51
🤗 Upvotes: 37 | cs. CV, cs. AI, cs. CL Authors: Alex Su, Haozhe Wang, Weimin Ren, Fangzhen Lin, Wenhu Chen Title: Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.15966v1 Abstract: Chain-of-thought reasoning has significantly improved the performance of Large Language Models (LLMs) across various domains. However, th...
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models 24.05.2025 20:54
🤗 Upvotes: 36 | cs. CV Authors: Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Gang Yu, Wenbo Zhu, Bernt Schiele, Ming-Hsuan Yang, Xu Yang Title: KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models Arxiv: http://arxiv.org/abs/2505.16707v1 Abstract: Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing...
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 24.05.2025 19:40
🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Benjamin Schneider, Dongfu Jiang, Chao Du, Tianyu Pang, Wenhu Chen Title: QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design Arxiv: http://arxiv.org/abs/2505.16175v1 Abstract: Long-video understanding has emerged as a crucial capability in real-world applications such as video surveillance, meeting summarization, educational lect...
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning 24.05.2025 25:23
🤗 Upvotes: 23 | cs. CV, cs. AI, cs. CL, cs. LG, cs. MM Authors: Chengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang, Linjiang Huang, Xingyu Zeng, Hongsheng Li, Xihui Liu Title: GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.17022v1 Abstract: Visual generation models have made remarkable progress in creating realisti...
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning 24.05.2025 20:15
🤗 Upvotes: 22 | cs. LG, cs. CL, cs. CV Authors: Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, Chongxuan Li Title: LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Arxiv: http://arxiv.org/abs/2505.16933v1 Abstract: In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instructi...
Scaling Diffusion Transformers Efficiently via $μ$P 24.05.2025 22:53
🤗 Upvotes: 21 | cs. LG, cs. AI, cs. CV Authors: Chenyu Zheng, Xinyu Zhang, Rongzhen Wang, Wei Huang, Zhi Tian, Weilin Huang, Jun Zhu, Chongxuan Li Title: Scaling Diffusion Transformers Efficiently via $μ$P Arxiv: http://arxiv.org/abs/2505.15270v1 Abstract: Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyper...
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents 23.05.2025 22:49
🤗 Upvotes: 80 | cs. CL Authors: Hyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Minseok Kang, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Minju Kim, Beong-woo Kwak, Dongjin Kang, Jinyoung Yeo Title: Web-Shepherd: Advancing PRMs for Reinforcing Web Agents Arxiv: http://arxi...
MMaDA: Multimodal Large Diffusion Language Models 23.05.2025 21:00
🤗 Upvotes: 56 | cs. CV Authors: Ling Yang, Ye Tian, Bowen Li, Xinchen Zhang, Ke Shen, Yunhai Tong, Mengdi Wang Title: MMaDA: Multimodal Large Diffusion Language Models Arxiv: http://arxiv.org/abs/2505.15809v1 Abstract: We introduce MMaDA, a novel class of multimodal diffusion foundation models designed to achieve superior performance across diverse domains such as textual reasoning, multimodal un...
Scaling Law for Quantization-Aware Training 23.05.2025 19:44
🤗 Upvotes: 56 | cs. LG, cs. CL Authors: Mengzhao Chen, Chaoyi Zhang, Jing Liu, Yutao Zeng, Zeyue Xue, Zhiheng Liu, Yunshui Li, Jin Ma, Jie Huang, Xun Zhou, Ping Luo Title: Scaling Law for Quantization-Aware Training Arxiv: http://arxiv.org/abs/2505.14302v1 Abstract: Large language models (LLMs) demand substantial computational and memory resources, creating deployment challenges. Quantization-awa...
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning 23.05.2025 18:15
🤗 Upvotes: 43 | cs. CV Authors: Sule Bai, Mingxing Li, Yong Liu, Jing Tang, Haoji Zhang, Lei Sun, Xiangxiang Chu, Yansong Tang Title: UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.14231v1 Abstract: Traditional visual grounding methods primarily focus on single-image scenarios with simple textual references. However, extending th...
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 23.05.2025 20:47
🤗 Upvotes: 41 | cs. CL Authors: Siyue Zhang, Yilun Zhao, Liyuan Geng, Arman Cohan, Anh Tuan Luu, Chen Zhao Title: Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective Arxiv: http://arxiv.org/abs/2505.15045v1 Abstract: Large language model (LLM)-based embedding models, benefiting from large scale pre-training and post-training, have begun to surpass BERT and T5-based models o...
Efficient Agent Training for Computer Use 23.05.2025 23:07
🤗 Upvotes: 32 | cs. AI, cs. CL, cs. LG Authors: Yanheng He, Jiahe Jin, Pengfei Liu Title: Efficient Agent Training for Computer Use Arxiv: http://arxiv.org/abs/2505.13909v1 Abstract: Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces rel...
This Time is Different: An Observability Perspective on Time Series Foundation Models 23.05.2025 22:05
🤗 Upvotes: 28 | cs. LG, cs. AI Authors: Ben Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi, Chris Lettieri, Charles Masson, Hugo Miccinilli, Elise Ramé, Qiqi Ren, Afshin Rostamizadeh, Jean Ogier du Terrail, Anna-Monica Toon, Kan Wang, Stephan Xie, David Asker, Ameet Talwalkar, Othmane Abou-Amal Title: This Time is Different: An Observability Perspective on Time Series Foundation Models...
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping 23.05.2025 17:46
🤗 Upvotes: 24 | cs. CL, cs. AI, cs. LG Authors: Wei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang, Junteng Liu, Yuntian Deng, Yizhe Zhang, Junxian He Title: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Arxiv: http://arxiv.org/abs/2505.15612v1 Abstract: Large Reasoning Models (LRMs) have shown remarkable capabilities in solving complex problems through reinforcement learning...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.