Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Model Merging Scaling Laws in Large Language Models 13.05.2026 21:44
🤗 Upvotes: 26 | cs. AI Authors: Yuanyi Wang, Yanggan Gu, Yiming Zhang, Qi Zhou, Zhaoyi Yan, Congkai Xie, Xinyao Wang, Jianbo Yuan, Hongxia Yang Title: Model Merging Scaling Laws in Large Language Models Arxiv: http://arxiv.org/abs/2509.24244v4 Abstract: We study empirical scaling laws for language model merging measured by cross-entropy. Despite its wide practical use, merging lacks a quantitativ...
SEIF: Self-Evolving Reinforcement Learning for Instruction Following 13.05.2026 21:27
🤗 Upvotes: 25 | cs. CL Authors: Qingyu Ren, Qianyu He, Jiajie Zhu, Xingzhou Chen, Jingwen Chang, Zeye Sun, Han Xia, Fei Yu, Jiaqing Liang, Yanghua Xiao Title: SEIF: Self-Evolving Reinforcement Learning for Instruction Following Arxiv: http://arxiv.org/abs/2605.07465v1 Abstract: Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capab...
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors 13.05.2026 22:26
🤗 Upvotes: 24 | cs. CV Authors: Keming Wu, Yijing Cui, Wenhan Xue, Qijie Wang, Xuan Luo, Zhiyuan Feng, Zuhao Yang, Sudong Wang, Sicong Jiang, Haowei Zhu, Zihan Wang, Ping Nie, Wenhu Chen, Bin Wang Title: WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors Arxiv: http://arxiv.org/abs/2605.10434v1 Abstract: Commercial video generation systems such as...
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models 13.05.2026 22:29
🤗 Upvotes: 21 | cs. CL, cs. AI, cs. LG Authors: Victor Conchello Vendrell, Arnau Padres Masdemont, Niccolò Grillo, Jordi Ros-Giralt, Arash Behboodi, Fabio Valerio Massoli Title: Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Arxiv: http://arxiv.org/abs/2605.07721v1 Abstract: Recurrent LLM architectures have emerged as a promising approach for improvi...
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers 12.05.2026 21:53
🤗 Upvotes: 106 | cs. LG, cs. CV Authors: Pengqi Lu Title: Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Arxiv: http://arxiv.org/abs/2605.06169v1 Abstract: Scaling Diffusion Transformers (DiTs) to hundreds of layers introduces a structural vulnerability: networks can enter a silent, mean-dominated collapse state that homogenizes token representations and...
Flow-OPD: On-Policy Distillation for Flow Matching Models 12.05.2026 26:15
🤗 Upvotes: 79 | cs. CV, cs. AI Authors: Zhen Fang, Wenxuan Huang, Yu Zeng, Yiming Zhao, Shuang Chen, Kaituo Feng, Yunlong Lin, Lin Chen, Zehui Chen, Shaosheng Cao, Feng Zhao Title: Flow-OPD: On-Policy Distillation for Flow Matching Models Arxiv: http://arxiv.org/abs/2605.08063v1 Abstract: Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignm...
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents 12.05.2026 25:21
🤗 Upvotes: 57 | cs. LG, cs. AI Authors: Guankai Li, Jiabin Chen, Yi Xu, Xichen Zhang, Yuan Lu Title: HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents Arxiv: http://arxiv.org/abs/2605.07177v1 Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction r...
Anisotropic Modality Align 12.05.2026 22:54
🤗 Upvotes: 23 | cs. MM, cs. CV Authors: Xiaomin Yu, Yijiang Li, Yuhui Zhang, Hanzhen Zhao, Yue Yang, Hao Tang, Yue Song, Xiaobin Hu, Chengwei Qin, Shuicheng Yan, Hui Xiong Title: Anisotropic Modality Align Arxiv: http://arxiv.org/abs/2605.07825v1 Abstract: Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show th...
Beyond Retrieval: A Multitask Benchmark and Model for Code Search 12.05.2026 21:14
🤗 Upvotes: 22 | cs. SE, cs. AI Authors: Siqiao Xue, Zihan Liao, Jin Qin, Ziyin Zhang, Yixiang Mu, Fan Zhou, Hang Yu Title: Beyond Retrieval: A Multitask Benchmark and Model for Code Search Arxiv: http://arxiv.org/abs/2605.04615v2 Abstract: Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and developer-style qu...
MiA-Signature: Approximating Global Activation for Long-Context Understanding 09.05.2026 11:50
🤗 Upvotes: 41 | cs. CL Authors: Yuqing Li, Jiangnan Li, Mo Yu, Zheng Lin, Weiping Wang, Jie Zhou Title: MiA-Signature: Approximating Global Activation for Long-Context Understanding Arxiv: http://arxiv.org/abs/2605.06416v1 Abstract: A growing body of work in cognitive science suggests that reportable conscious access is associated with \emph{global ignition} over distributed memory systems, while...
When to Trust Imagination: Adaptive Action Execution for World Action Models 09.05.2026 12:11
🤗 Upvotes: 34 | cs. RO, cs. AI Authors: Rui Wang, Yue Zhang, Jiehong Lin, Kuncheng Luo, Jianan Wang, Zhongrui Wang, Xiaojuan Qi Title: When to Trust Imagination: Adaptive Action Execution for World Action Models Arxiv: http://arxiv.org/abs/2605.06222v1 Abstract: World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual ob...
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation 09.05.2026 15:16
🤗 Upvotes: 24 | cs. CV, cs. AI Authors: Tao Liu, Hao Yan, Mengting Chen, Taihang Hu, Zhengrong Yue, Zihao Pan, Jinsong Lan, Xiaoyong Zhu, Ming-Ming Cheng, Bo Zheng, Yaxing Wang Title: Continuous-Time Distribution Matching for Few-Step Diffusion Distillation Arxiv: http://arxiv.org/abs/2605.06376v1 Abstract: Step distillation has become a leading technique for accelerating diffusion models, among...
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation 08.05.2026 22:23
🤗 Upvotes: 109 | cs. CV Authors: Bin Wu, Mengqi Huang, Shaojin Wu, Weinan Jia, Yuxin Wang, Zhendong Mao, Yongdong Zhang Title: Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation Arxiv: http://arxiv.org/abs/2605.03849v1 Abstract: Distillation-based acceleration has become foundational for making autoregressive streaming video diffusion models practical, with...
Stream-T1: Test-Time Scaling for Streaming Video Generation 08.05.2026 21:59
🤗 Upvotes: 94 | cs. CV Authors: Yijing Tu, Shaojin Wu, Mengqi Huang, Wenchuan Wang, Yuxin Wang, Chunxiao Liu, Zhendong Mao Title: Stream-T1: Test-Time Scaling for Streaming Video Generation Arxiv: http://arxiv.org/abs/2605.04461v1 Abstract: While Test-Time Scaling (TTS) offers a promising direction to enhance video generation without the surging costs of training, current test-time video generati...
RLDX-1 Technical Report 08.05.2026 23:03
🤗 Upvotes: 92 | cs. RO, cs. AI, cs. LG Authors: Dongyoung Kim, Huiwon Jang, Myungkyu Koo, Suhyeok Jang, Taeyoung Kim, Beomjun Kim, Byungjun Yoon, Changsung Jang, Daewon Choi, Dongsu Han, Donguk Lee, Heeseung Kwon, Hojin Jeon, Jaehyun Kang, Jaekyoung Bae, Jihyuk Lee, Jimin Lee, John Won, Joonwoo Ahn, Junhyeong Park, Junyoung Sung, Kyungmin Lee, Minseong Han, Minsung Yoon, Sejune Joo, Seonil Son, S...
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents 08.05.2026 25:07
🤗 Upvotes: 84 | cs. CV Authors: Shuang Chen, Kaituo Feng, Hangting Chen, Wenxuan Huang, Dasen Dai, Quanxin Shou, Yunlong Lin, Xiangyu Yue, Shenghua Gao, Tianyu Pang Title: OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Arxiv: http://arxiv.org/abs/2605.05185v1 Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex...
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation 08.05.2026 23:18
🤗 Upvotes: 68 | cs. CV Authors: Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai Title: HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Arxiv: http://arxiv.org/abs/2604.28196v1 Abstract: Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existi...
PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World 08.05.2026 21:38
🤗 Upvotes: 30 | cs. CV Authors: Yunhan Yang, Chunshi Wang, Junliang Ye, Yang Li, Zanxin Chen, Zehuan Huang, Yao Mu, Zhuo Chen, Chunchao Guo, Xihui Liu Title: PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World Arxiv: http://arxiv.org/abs/2605.05163v1 Abstract: Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI...
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration 07.05.2026 24:16
🤗 Upvotes: 75 | cs. SE, cs. AI Authors: Ruofeng Yang, Yongcan Li, Shuai Li Title: ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Arxiv: http://arxiv.org/abs/2605.03042v1 Abstract: This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The p...
OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories 07.05.2026 21:54
🤗 Upvotes: 44 | cs. AI, cs. CL Authors: Yuwen Du, Rui Ye, Shuo Tang, Keduan Huang, Xinyu Zhu, Yuzhu Cai, Siheng Chen Title: OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories Arxiv: http://arxiv.org/abs/2605.04036v1 Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their...
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL 07.05.2026 21:27
🤗 Upvotes: 37 | cs. CV, cs. AI, cs. CL Authors: Sudong Wang, Weiquan Huang, Xiaomin Yu, Zuhao Yang, Hehai Lin, Keming Wu, Chaojun Xiao, Chen Chen, Wenxuan Wang, Beier Zhu, Yunjian Zhang, Chengwei Qin Title: Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL Arxiv: http://arxiv.org/abs/2604.28123v2 Abstract: The standard post-training recipe for large multimodal...
MolmoAct2: Action Reasoning Models for Real-world Deployment 06.05.2026 25:04
🤗 Upvotes: 170 | cs. RO Authors: Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai, Shirui Chen, Yi Ru Wang, Shanli Xing, Jaemin Cho, Jae Sung Park, Ainaz Eftekhar, Peter Sushko, Karen Farley, Angad Wadhwa, Cole Harrison, Winson Han, Ying-Chun Lee, Eli VanderBilt, Rose Hendrix, Suveen Ellawela, Lucas Ngoo, Joyce Chai, Zhongzheng Ren, Ali Farhadi,...
From Context to Skills: Can Language Models Learn from Context Skillfully? 06.05.2026 19:07
🤗 Upvotes: 123 | cs. AI Authors: Shuzheng Si, Haozhe Zhao, Yu Lei, Qingyi Wang, Dingwei Chen, Zhitong Wang, Zhenhailong Wang, Kangyang Luo, Zheng Wang, Gang Chen, Fanchao Qi, Minjia Zhang, Maosong Sun Title: From Context to Skills: Can Language Models Learn from Context Skillfully? Arxiv: http://arxiv.org/abs/2604.27660v2 Abstract: Many real-world tasks require language models (LMs) to reason ove...
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors 05.05.2026 22:14
🤗 Upvotes: 70 | cs. CV Authors: Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo, Chongjie Ye, Lvmin Zhang, Hao Zhao, Anyi Rao Title: UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors Arxiv: http://arxiv.org/abs/2605.00658v1 Abstract: Recent progress has shown that video diffusion models (VDMs) can be repurposed for...
Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction 05.05.2026 23:47
🤗 Upvotes: 27 | cs. AI Authors: Yuxuan Huang, Yihang Chen, Zhiyuan He, Yuxiang Chen, Ka Yiu Lee, Huichi Zhou, Weilin Luo, Meng Fang, Jun Wang Title: Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction Arxiv: http://arxiv.org/abs/2604.27221v1 Abstract: Agentic web search increasingly faces two distinct demands: deep reasoning over a single target, a...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.