Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
AURA: Always-On Understanding and Real-Time Assistance via Video Streams 08.04.2026 23:35
🤗 Upvotes: 39 | cs. CV Authors: Xudong Lu, Yang Bo, Jinpeng Chen, Shuhan Li, Xintong Guo, Huankang Guan, Fang Liu, Dunyuan Xu, Peiwen Sun, Heyang Sun, Rui Liu, Hongsheng Li Title: AURA: Always-On Understanding and Real-Time Assistance via Video Streams Arxiv: http://arxiv.org/abs/2604.04184v1 Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on many video understa...
ClawArena: Benchmarking AI Agents in Evolving Information Environments 08.04.2026 21:28
🤗 Upvotes: 27 | cs. LG, cs. AI, cs. CL Authors: Haonian Ji, Kaiwen Xiong, Siwei Han, Peng Xia, Shi Qiu, Yiyang Zhou, Jiaqi Liu, Jinlong Li, Bingzhou Li, Zeyu Zheng, Cihang Xie, Huaxiu Yao Title: ClawArena: Benchmarking AI Agents in Evolving Information Environments Arxiv: http://arxiv.org/abs/2604.04202v1 Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their...
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing 08.04.2026 22:34
🤗 Upvotes: 26 | cs. CV Authors: Yicheng Xiao, Wenhu Zhang, Lin Song, Yukang Chen, Wenbo Li, Nan Jiang, Tianhe Ren, Haokun Lin, Wei Huang, Haoyang Huang, Xiu Li, Nan Duan, Xiaojuan Qi Title: SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing Arxiv: http://arxiv.org/abs/2604.04911v1 Abstract: Image spatial editing performs geometry-driven transformations, allowing precise control over obj...
LightThinker++: From Reasoning Compression to Memory Management 08.04.2026 20:05
🤗 Upvotes: 25 | cs. CL, cs. AI, cs. IR, cs. LG, cs. MM Authors: Yuqi Zhu, Jintian Zhang, Zhenjie Wan, Yujie Luo, Shuofei Qiao, Zhengke Gui, Da Zheng, Lei Liang, Huajun Chen, Ningyu Zhang Title: LightThinker++: From Reasoning Compression to Memory Management Arxiv: http://arxiv.org/abs/2604.03679v1 Abstract: Large language models (LLMs) excel at complex reasoning, yet their efficiency is limited b...
Self-Distilled RLVR 07.04.2026 21:55
🤗 Upvotes: 89 | cs. LG, cs. CL Authors: Chenxu Yang, Chuanyu Qin, Qingyi Si, Minghui Chen, Naibin Gu, Dingyu Yao, Zheng Lin, Weiping Wang, Jiaqi Wang, Nan Duan Title: Self-Distilled RLVR Arxiv: http://arxiv.org/abs/2604.03128v1 Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, f...
A Simple Baseline for Streaming Video Understanding 07.04.2026 21:50
🤗 Upvotes: 58 | cs. CV Authors: Yujiao Shen, Shulin Tian, Jingkang Yang, Ziwei Liu Title: A Simple Baseline for Streaming Video Understanding Arxiv: http://arxiv.org/abs/2604.02317v1 Abstract: Recent streaming video understanding methods increasingly rely on complex memory mechanisms to handle long video streams. We challenge this trend with a simple finding: a sliding-window baseline that feeds...
Token Warping Helps MLLMs Look from Nearby Viewpoints 07.04.2026 20:05
🤗 Upvotes: 23 | cs. CV Authors: Phillip Y. Lee, Chanho Park, Mingue Park, Seungwoo Yoo, Juil Koo, Minhyuk Sung Title: Token Warping Helps MLLMs Look from Nearby Viewpoints Arxiv: http://arxiv.org/abs/2604.02870v1 Abstract: Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual...
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? 07.04.2026 22:43
🤗 Upvotes: 23 | cs. AI Authors: Qianshan Wei, Yishan Yang, Siyi Wang, Jinglin Chen, Binyu Wang, Jiaming Wang, Shuang Chen, Zechen Li, Yang Shi, Yuqi Tang, Weining Wang, Yi Yu, Chaoyou Fu, Qi Li, Yi-Fan Zhang Title: Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? Arxiv: http://arxiv.org/abs/2604.03016v1 Abstract: Multimodal Large Language Models (MLLMs) are evolving...
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models 04.04.2026 27:27
🤗 Upvotes: 144 | cs. LG, cs. CL Authors: Hao Liang, Zhengyang Zhao, Meiyi Qiang, Mingrui Chen, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu, Conghui He, Bin Cui, Zhiyu Li, Weinan E, Wentao Zhang Title: DataFlex: A Unified Framework for Data-Ce...
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook 04.04.2026 22:54
🤗 Upvotes: 102 | cs. AI Authors: Xinlei Yu, Zhangquan Chen, Yongbo He, Tianyu Fu, Cheng Yang, Chengming Xu, Yue Ma, Xiaobin Hu, Zhe Cao, Jie Xu, Guibin Zhang, Jiale Tao, Jiayi Zhang, Siyuan Ma, Kaituo Feng, Haojie Huang, Youxing Li, Ronghao Chen, Huacan Wang, Chenglin Wu, Zikun Su, Xiaogang Xu, Kelu Yao, Kun Wang, Chen Gao, Yue Liao, Ruqi Huang, Tao Jin, Cheng Tan, Jiangning Zhang, Wenqi Ren, Yan...
Generative World Renderer 04.04.2026 22:41
🤗 Upvotes: 76 | cs. CV Authors: Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan, Ruihan Yu, Yidan Zhang, Bo Zheng, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang Title: Generative World Renderer Arxiv: http://arxiv.org/abs/2604.02329v1 Abstract: Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datas...
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization 04.04.2026 19:28
🤗 Upvotes: 72 | cs. LG Authors: Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Chengcheng Han, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen Title: SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization Arxiv: http://arxiv.org/abs/2604.02268v1 Abstract: Agent skills, structured packages of procedural knowledge and executable resources that agents dynamically...
Steerable Visual Representations 04.04.2026 21:42
🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, Yuki M. Asano Title: Steerable Visual Representations Arxiv: http://arxiv.org/abs/2604.02327v1 Abstract: Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation...
EgoSim: Egocentric World Simulator for Embodied Interaction Generation 04.04.2026 24:47
🤗 Upvotes: 30 | cs. CV, cs. AI Authors: Jinkun Hao, Mingda Jia, Ruiyan Wang, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, Xudong Xu Title: EgoSim: Egocentric World Simulator for Embodied Interaction Generation Arxiv: http://arxiv.org/abs/2604.01001v1 Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently u...
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery 04.04.2026 24:57
🤗 Upvotes: 22 | cs. AI Authors: Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu, Xuan Jiang, Sirui Li, Cathy Wu, Bryan Kian Hsiang Low, Jinhua Zhao, Paul Pu Liang Title: CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery Arxiv: http://arxiv.org/abs/2604.01658v1 Abstract: Large language...
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers 03.04.2026 25:10
🤗 Upvotes: 167 | cs. CR, cs. AI Authors: Songyang Liu, Chaozhuo Li, Chenxu Wang, Jinyu Hou, Zejian Chen, Litian Zhang, Zheng Liu, Qiwei Ye, Yiming Hei, Xi Zhang, Zhongyuan Wang Title: ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers Arxiv: http://arxiv.org/abs/2603.24414v1 Abstract: OpenClaw has rapidly established itself as a leading open-sour...
Terminal Agents Suffice for Enterprise Automation 03.04.2026 24:09
🤗 Upvotes: 71 | cs. SE, cs. AI, cs. CL Authors: Patrice Bechard, Orlando Marquez Ayala, Emily Chen, Jordan Skelton, Sagar Davasam, Srinivas Sunkara, Vikas Yadav, Sai Rajeswar Title: Terminal Agents Suffice for Enterprise Automation Arxiv: http://arxiv.org/abs/2604.00073v1 Abstract: There has been growing interest in building agents that can interact with digital platforms to execute meaningful en...
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome 03.04.2026 24:48
🤗 Upvotes: 53 | cs. AI, cs. CL Authors: Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao, Yibo Wang, Lei Wang, Zhen Zhang, Lu Wang, Yue Deng, Bin Wang, Yifan Zhang, Liangcai Su, Xinyu Wang, He Zhao, Chen Wei, Qiang Ren, Bryan Hooi, An Bo, Shuicheng Yan, Lidong Bing Title: MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome Arxiv: http://arxiv.org/abs/2603....
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners? 03.04.2026 24:35
🤗 Upvotes: 36 | cs. CV, cs. AI Authors: Haonan Han, Jiancheng Huang, Xiaopeng Sun, Junyan He, Rui Yang, Jie Hu, Xiaojiang Peng, Lin Ma, Xiaoming Wei, Xiu Li Title: ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners? Arxiv: http://arxiv.org/abs/2603.25823v1 Abstract: Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fai...
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification 03.04.2026 23:27
🤗 Upvotes: 33 | cs. SE, cs. AI Authors: Zehai He, Wenyi Hong, Zhen Yang, Ziyang Pan, Mingdao Liu, Xiaotao Gu, Jie Tang Title: Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Arxiv: http://arxiv.org/abs/2603.26648v2 Abstract: Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, en...
QuitoBench: A High-Quality Open Time Series Forecasting Benchmark 03.04.2026 24:46
🤗 Upvotes: 25 | cs. LG Authors: Siqiao Xue, Zhaoyang Zhu, Wei Zhang, Rongyao Cai, Rui Wang, Yixiang Mu, Fan Zhou, Jianguo Li, Peng Di, Hang Yu Title: QuitoBench: A High-Quality Open Time Series Forecasting Benchmark Arxiv: http://arxiv.org/abs/2603.26017v1 Abstract: Time series forecasting is critical across finance, healthcare, and cloud computing, yet progress is constrained by a fundamental bo...
Reasoning Shift: How Context Silently Shortens LLM Reasoning 03.04.2026 23:27
🤗 Upvotes: 22 | cs. LG Authors: Gleb Rodionov Title: Reasoning Shift: How Context Silently Shortens LLM Reasoning Arxiv: http://arxiv.org/abs/2604.01161v1 Abstract: Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of th...
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization 02.04.2026 23:55
🤗 Upvotes: 293 | cs. LG Authors: Chiyu Ma, Shuo Yang, Kexin Huang, Jinda Lu, Haoming Meng, Shangshang Wang, Bolin Ding, Soroush Vosoughi, Guoyin Wang, Jingren Zhou Title: FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization Arxiv: http://arxiv.org/abs/2603.19835v3 Abstract: We present Future-KL Influenced Policy Optimization (FIPO), a reinforcement learning algorithm desig...
CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence 02.04.2026 28:20
🤗 Upvotes: 230 | cs. RO, cs. AI, cs. CV, cs. HC Authors: Tianle Zeng, Hanxuan Chen, Yanci Wen, Hong Zhang Title: CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence Arxiv: http://arxiv.org/abs/2603.28032v1 Abstract: The convergence of low-altitude economies, embodied intelligence, and air-ground cooperative systems creates growing demand for...
LongCat-Next: Lexicalizing Modalities as Discrete Tokens 02.04.2026 23:25
🤗 Upvotes: 118 | cs. CV, cs. CL Authors: Meituan LongCat Team, Bin Xiao, Chao Wang, Chengjiang Li, Chi Zhang, Chong Peng, Hang Yu, Hao Yang, Haonan Yan, Haoze Sun, Haozhe Zhao, Hong Liu, Hui Su, Jiaqi Zhang, Jiawei Wang, Jing Li, Kefeng Zhang, Manyuan Zhang, Minhao Jing, Peng Pei, Quan Chen, Taofeng Xue, Tongxin Pan, Xiaotong Li, Xiaoyang Li, Xiaoyu Zhao, Xing Hu, Xinyang Lin, Xunliang Cai, Yan B...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.