Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Heterogeneous Scientific Foundation Model Collaboration 02.05.2026 25:02
🤗 Upvotes: 180 | cs. AI, cs. CL, cs. LG Authors: Zihao Li, Jiaru Zou, Feihao Fang, Xuying Ning, Mengting Ai, Tianxin Wei, Sirui Chen, Xiyuan Yang, Jingrui He Title: Heterogeneous Scientific Foundation Model Collaboration Arxiv: http://arxiv.org/abs/2604.27351v1 Abstract: Agentic large language model systems have demonstrated strong capabilities. However, their reliance on language as the universa...
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling 02.05.2026 20:21
🤗 Upvotes: 70 | cs. CV Authors: Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, Haowei Zhu, Sicong Leng, Zhongyu Yang, Qijie Wang, Sudong Wang, Ziting Wang, Zili Wang, Hui Zhang, Haonan Wang, Hang Zhou, Yifan Pu, Xingxuan Li, Fangneng Zhan, Bo Li, Lidong Bing, Yuxin Song, Ziwei Liu, Wenhu Chen, Jingdong Wang, Xinchao Wang, Xiaojuan Qi, Shijian Lu, Bin Wang Title: Visual Generation in the New E...
Co-Evolving Policy Distillation 02.05.2026 22:49
🤗 Upvotes: 34 | cs. LG Authors: Naibin Gu, Chenxu Yang, Qingyi Si, Chuanyu Qin, Dingyu Yao, Peng Fu, Zheng Lin, Weiping Wang, Nan Duan, Jiaqi Wang Title: Co-Evolving Policy Distillation Arxiv: http://arxiv.org/abs/2604.27083v1 Abstract: RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities...
ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control 02.05.2026 27:23
🤗 Upvotes: 32 | cs. RO Authors: Yanghao Zhou, Jingyu Ma, Yibo Peng, Zhenguo Sun, Yu Bai, Börje F. Karlsson Title: ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control Arxiv: http://arxiv.org/abs/2604.27711v1 Abstract: Humanoid control systems have made significant progress in recent years, yet modeling fluent interaction-rich behavior between a robot, its surroundin...
Efficient Training on Multiple Consumer GPUs with RoundPipe 02.05.2026 23:18
🤗 Upvotes: 24 | cs. DC, cs. AI, cs. LG Authors: Yibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, Jiwu Shu Title: Efficient Training on Multiple Consumer GPUs with RoundPipe Arxiv: http://arxiv.org/abs/2604.27085v1 Abstract: Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe interconnects. Pipeline parallel...
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents 01.05.2026 25:34
🤗 Upvotes: 74 | cs. CV Authors: GLM-V Team, :, Wenyi Hong, Xiaotao Gu, Ziyang Pan, Zhen Yang, Yuting Wang, Yue Wang, Yuanchang Yue, Yu Wang, Yanling Wang, Yan Wang, Xijun Liu, Wenmeng Yu, Weihan Wang, Wei Li, Shuaiqi Duan, Sheng Yang, Ruiliang Lv, Mingdao Liu, Lihang Pan, Ke Ning, Junhui Ji, Jinjiang Wang, Jing Chen, Jiazheng Xu, Jiale Zhu, Jiale Cheng, Ji Qi, Guobing Gan, Guo Wang, Cong Yao, Zij...
Large Language Models Explore by Latent Distilling 01.05.2026 22:37
🤗 Upvotes: 56 | cs. CL, cs. AI, cs. LG Authors: Yuanhao Zeng, Ao Lu, Lufei Li, Zheng Zhang, Yexin Li, Kan Ren Title: Large Language Models Explore by Latent Distilling Arxiv: http://arxiv.org/abs/2604.24927v1 Abstract: Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yields surface-level lexical variation, limit...
RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments 01.05.2026 24:14
🤗 Upvotes: 49 | cs. CV Authors: Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov, Jaafar Mahmoud, Sergey Kolyubin Title: RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments Arxiv: http://arxiv.org/abs/2604.26067v1 Abstract: We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that...
ClawGym: A Scalable Framework for Building Effective Claw Agents 01.05.2026 26:01
🤗 Upvotes: 38 | cs. CL, cs. AI, cs. LG Authors: Fei Bai, Huatong Song, Shuang Sun, Daixuan Cheng, Yike Yang, Chuan Hao, Renyuan Li, Feng Chang, Yuan Wei, Ran Tao, Bryan Dai, Jian Yang, Wayne Xin Zhao Title: ClawGym: A Scalable Framework for Building Effective Claw Agents Arxiv: http://arxiv.org/abs/2604.26904v1 Abstract: Claw-style environments support multi-step workflows over local files, tools...
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models 01.05.2026 22:25
🤗 Upvotes: 37 | cs. CL, cs. AI, cs. LG Authors: Gongbo Zhang, Wen Wang, Ye Tian, Li Yuan Title: Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models Arxiv: http://arxiv.org/abs/2604.26951v1 Abstract: Diffusion large language models (dLLMs) offer parallel decoding and bidirectional context, but state-of-the-art dLLMs require billions of parameters for competitive p...
Recursive Multi-Agent Systems 30.04.2026 25:02
🤗 Upvotes: 131 | cs. AI, cs. CL, cs. LG Authors: Xiyuan Yang, Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus J. Buehler, Jingrui He, James Zou Title: Recursive Multi-Agent Systems Arxiv: http://arxiv.org/abs/2604.25917v1 Abstract: Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the sam...
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora 30.04.2026 24:15
🤗 Upvotes: 75 | cs. SE, cs. AI Authors: Chenkai Pan, Xinglong Xu, Yuhang Xu, Yujun Wu, Siyuan Li, Jintao Chen, Conghui He, Jingxuan Wei, Cheng Tan Title: Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora Arxiv: http://arxiv.org/abs/2604.24819v1 Abstract: Reliably transferring specialized human knowledge from text into large language models remains a fund...
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios 30.04.2026 23:08
🤗 Upvotes: 37 | cs. CL Authors: Jinxiang Meng, Shaoping Huang, Fangyu Lei, Jingyu Guo, Haoxiang Liu, Jiahao Su, Sihan Wang, Yao Wang, Enrui Wang, Ye Yang, Hongze Chai, Jinming Lv, Anbang Yu, Huangjing Zhang, Yitong Zhang, Yiming Huang, Zeyao Ma, Shizhu He, Jun Zhao, Kang Liu Title: DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Arxiv: http://arxiv.org/abs/2604.25914v1 Ab...
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery 30.04.2026 21:40
🤗 Upvotes: 26 | cs. AI Authors: Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang, Jin-Ge Yao, Zheng Liu, Jingying Shao, Jianlyu Chen, Hongjin Qian, Xi Yang, Qian Yu, Hao Li, Chen Yue, Xiaan Du, Yuyang Wang, Yesheng Liu, Haiyu Xu, Zhicheng Dou Title: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery Arxiv: http://arxiv.org/abs/2604.25256v1 Abstract: Autonomous scientifi...
Meta-CoT: Enhancing Granularity and Generalization in Image Editing 30.04.2026 22:01
🤗 Upvotes: 24 | cs. CV, cs. AI, cs. LG, cs. MM Authors: Shiyi Zhang, Yiji Cheng, Tiankai Hang, Zijin Yin, Runze He, Yu Xu, Wenxun Dai, Yunlong Lin, Chunyu Wang, Qinglin Lu, Yansong Tang Title: Meta-CoT: Enhancing Granularity and Generalization in Image Editing Arxiv: http://arxiv.org/abs/2604.24625v1 Abstract: Unified multi-modal understanding/generative models have shown improved image editing p...
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models 30.04.2026 23:20
🤗 Upvotes: 22 | cs. CV Authors: Jiayi Guo, Linqing Wang, Jiangshan Wang, Yang Yue, Zeyu Liu, Zhiyuan Zhao, Qinglin Lu, Gao Huang, Chunyu Wang Title: Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models Arxiv: http://arxiv.org/abs/2604.25636v1 Abstract: Unified multimodal models (UMMs) integrate visual understanding and generation within a...
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 29.04.2026 21:09
🤗 Upvotes: 102 | cs. CV Authors: Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang, Zeyu Zhang, Yefei He, Yanbo Ding, Xirui Hu, Donny Y. Chen, Zhiyuan He, Yuqing Yang, Bohan Zhuang Title: World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Arxiv: http://arxiv.org/abs/2604.24764v1 Abstract: Recent video foundation models demonstrate impressive visual synthesis but frequently suffer fr...
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company 29.04.2026 25:57
🤗 Upvotes: 100 | cs. AI Authors: Zhengxu Yu, Yu Fu, Zhiyuan He, Yuxuan Huang, Lee Ka Yiu, Meng Fang, Weilin Luo, Jun Wang Title: From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company Arxiv: http://arxiv.org/abs/2604.22446v1 Abstract: Individual agent capabilities have advanced rapidly through modular skills and tool integrations, yet multi-agent systems remain constrained...
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning 29.04.2026 22:53
🤗 Upvotes: 57 | cs. CV Authors: Yiming Zhang, Jiacheng Chen, Jiaqi Tan, Yongsen Mao, Wenhu Chen, Angel X. Chang Title: ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning Arxiv: http://arxiv.org/abs/2604.24300v1 Abstract: Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. Fir...
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation 29.04.2026 22:29
🤗 Upvotes: 47 | cs. CV Authors: Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, Tianhong Li, Mengzhao Chen, Yatai Ji, Sen He, Jonas Schult, Belinda Zeng, Tao Xiang, Wenhu Chen, Ping Luo, Luke Zettlemoyer, Yuren Cong Title: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Arxiv: http://arxiv.org/abs/2604.24763v1 Abstract: Unified multimodal models typi...
Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms 29.04.2026 29:53
🤗 Upvotes: 42 | cs. RO Authors: Qi Li, Bo Yin, Weiqi Huang, Ruhao Liu, Bojun Zou, Runpeng Yu, Jingwen Ye, Weihao Yu, Xinchao Wang Title: Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms Arxiv: http://arxiv.org/abs/2604.23775v1 Abstract: Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of...
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents 29.04.2026 22:23
🤗 Upvotes: 27 | cs. CV, cs. SE Authors: Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, Jiawei Gu, Pengfei Zhou, Rui Huang, Ziqi Zhao, Shengyuan Ding, Ailing Yu, Bo Peng, Bowei Xia, Hao Sun, Haotian Liang, Ji Xie, Jiajun Chen, Jiajun Song, Liu Yang, Ming Xu, Qionglin Qiu, Runhao Fu, Shengfang Zhai, Shijian Wang, Tengfei Ma, Tianyi Wu, W...
SketchVLM: Vision language models can annotate images to explain thoughts and guide users 29.04.2026 21:34
🤗 Upvotes: 22 | cs. CV, cs. AI Authors: Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen Title: SketchVLM: Vision language models can annotate images to explain thoughts and guide users Arxiv: http://arxiv.org/abs/2604.22875v2 Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In co...
Video Analysis and Generation via a Semantic Progress Function 28.04.2026 20:59
🤗 Upvotes: 42 | cs. CV Authors: Gal Metzer, Sagi Polaczek, Ali Mahdavi-Amiri, Raja Giryes, Daniel Cohen-Or Title: Video Analysis and Generation via a Semantic Progress Function Arxiv: http://arxiv.org/abs/2604.22554v1 Abstract: Transformations produced by image and video generation models often evolve in a highly non-linear manner: long stretches where the content barely changes are followed by s...
DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction 28.04.2026 25:02
🤗 Upvotes: 27 | eess. IV, cs. CV Authors: Shiyan Su, Ruyi Zha, Danli Shi, Hongdong Li, Xuelian Cheng Title: DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction Arxiv: http://arxiv.org/abs/2604.21518v1 Abstract: Neural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT)...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.