Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

One-Minute Video Generation with Test-Time Training 09.04.2025

🤗 Upvotes: 61 | cs. CV Authors: Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, Xiaolong Wang Title: One-Minute Video Generation with Test-Time Training Arxiv: http://arxiv.org/abs/2504.05298v1 Abstract: Transformers today still struggle to generate one-minu...

Rethinking Reflection in Pre-Training 09.04.2025

🤗 Upvotes: 52 | cs. CL, cs. AI Authors: Essential AI, :, Darsh J Shah, Peter Rushton, Somanshu Singla, Mohit Parmar, Kurt Smith, Yash Vanjani, Ashish Vaswani, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Anthony Polloreno, Ashish Tanwer, Burhan Drak Sibai, Divya S Mansingka, Divya Shivaprasad, Ishaan Shah, Karl Stratos, Khoi Nguyen, Michael Callahan, Michael Pust, Mrinal Iyer, Philip...

URECA: Unique Region Caption Anything 09.04.2025

🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Sangbeom Lim, Junwan Kim, Heeji Yoon, Jaewoo Jung, Seungryong Kim Title: URECA: Unique Region Caption Anything Arxiv: http://arxiv.org/abs/2504.05305v1 Abstract: Region-level captioning aims to generate natural language descriptions for specific image regions while highlighting their distinguishing features. However, existing methods struggle to produce uni...

T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models 09.04.2025

🤗 Upvotes: 29 | cs. CL, cs. AI Authors: Minki Kang, Jongwon Jeong, Jaewoong Cho Title: T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models Arxiv: http://arxiv.org/abs/2504.04718v1 Abstract: Recent studies have demonstrated that test-time compute scaling effectively improves the performance of small language models (sLMs). However, prior research has mainly...

Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving 08.04.2025

🤗 Upvotes: 30 | cs. SE, cs. AI, cs. CL Authors: Daoguang Zan, Zhirong Huang, Wei Liu, Hanwu Chen, Linhao Zhang, Shulin Xin, Lu Chen, Qi Liu, Xiaojian Zhong, Aoyan Li, Siyao Liu, Yongsheng Xiao, Liangqiang Chen, Yuyu Zhang, Jing Su, Tianyu Liu, Rui Long, Kai Shen, Liang Xiang Title: Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving Arxiv: http://arxiv.org/abs/2504.02605v1 Abstract: The...

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems 05.04.2025

🤗 Upvotes: 98 | cs. AI Authors: Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, Yuheng Cheng, Suyuchen Wang, Xiaoqiang Wang, Yuyu Luo, Haibo Jin, Peiyan Zhang, Ollie Liu, Jiaqi Chen, Huan Zhang, Zhaoyang Yu, Haochen Shi, Boyan Li, Dekun Wu, Fengwei Teng, Xiaojun Jia, Jiawei Xu, Jinyu Xiang, Yizhang Lin, Tianming Liu, To...

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing 05.04.2025

🤗 Upvotes: 55 | cs. CV Authors: Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Hao Li, Zicheng Zhang, Guangtao Zhai, Junchi Yan, Hua Yang, Xue Yang, Haodong Duan Title: Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing Arxiv: http://arxiv.org/abs/2504.02826v1 Abstract: Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation,...

ZClip: Adaptive Spike Mitigation for LLM Pre-Training 05.04.2025

🤗 Upvotes: 47 | cs. LG, cs. CL Authors: Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury, Fabian Güra Title: ZClip: Adaptive Spike Mitigation for LLM Pre-Training Arxiv: http://arxiv.org/abs/2504.02507v1 Abstract: Training large language models (LLMs) presents numerous challenges, including gradient instability and loss spikes. These phenomena can lead to catastrophic divergence, requiring costly...

GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation 05.04.2025

🤗 Upvotes: 34 | cs. CV Authors: Zhiyuan Yan, Junyan Ye, Weijia Li, Zilong Huang, Shenghai Yuan, Xiangyang He, Kaiqing Lin, Jun He, Conghui He, Li Yuan Title: GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation Arxiv: http://arxiv.org/abs/2504.02782v1 Abstract: The recent breakthroughs in OpenAI's GPT4o model have demonstrated surprisingly good capabilities in image gen...

Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme 05.04.2025

🤗 Upvotes: 24 | cs. LG, cs. CL, cs. CV Authors: Yan Ma, Steffi Chern, Xuyang Shen, Yiran Zhong, Pengfei Liu Title: Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme Arxiv: http://arxiv.org/abs/2504.02587v1 Abstract: Reinforcement learning (RL) has recently shown strong potential in improving the reasoning capabilities of la...

WikiVideo: Article Generation from Multiple Videos 05.04.2025

🤗 Upvotes: 24 | cs. CV, cs. CL Authors: Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme Title: WikiVideo: Article Generation from Multiple Videos Arxiv: http://arxiv.org/abs/2504.00939v1 Abstract: We present the challenging task of automatically creating a high-level Wikipedia-style article that aggregates informati...

MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization 04.04.2025

🤗 Upvotes: 57 | cs. CV, cs. AI Authors: Siyuan Li, Luyuan Zhang, Zedong Wang, Juanxi Tian, Cheng Tan, Zicheng Liu, Chang Yu, Qingsong Xie, Haonan Lu, Haoqian Wang, Zhen Lei Title: MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization Arxiv: http://arxiv.org/abs/2504.00999v1 Abstract: Masked Image Modeling (MIM) with Vector Quantizat...

AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 04.04.2025

🤗 Upvotes: 30 | cs. CV Authors: Junhao Cheng, Yuying Ge, Yixiao Ge, Jing Liao, Ying Shan Title: AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction Arxiv: http://arxiv.org/abs/2504.01014v1 Abstract: Recent advancements in image and video synthesis have opened up new promise in generative games. One particularly intriguing application is transforming characters from anime fi...

Understanding R1-Zero-Like Training: A Critical Perspective 04.04.2025

🤗 Upvotes: 25 | cs. LG, cs. AI, cs. CL Authors: Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin Title: Understanding R1-Zero-Like Training: A Critical Perspective Arxiv: http://arxiv.org/abs/2503.20783v1 Abstract: DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervis...

Towards Physically Plausible Video Generation via VLM Planning 04.04.2025

🤗 Upvotes: 25 | cs. CV, cs. AI Authors: Xindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, Xu Jia Title: Towards Physically Plausible Video Generation via VLM Planning Arxiv: http://arxiv.org/abs/2503.23368v2 Abstract: Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highl...

DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance 04.04.2025

🤗 Upvotes: 24 | cs. CV, cs. AI Authors: Yuxuan Luo, Zhengkun Rong, Lizhen Wang, Longhao Zhang, Tianshu Hu, Yongming Zhu Title: DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance Arxiv: http://arxiv.org/abs/2504.01724v2 Abstract: While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-g...

VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step 04.04.2025

🤗 Upvotes: 22 | cs. CV Authors: Hanyang Wang, Fangfu Liu, Jiawei Chi, Yueqi Duan Title: VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step Arxiv: http://arxiv.org/abs/2504.01956v2 Abstract: Recovering 3D scenes from sparse views is a challenging task due to its inherent ill-posed problem. Conventional methods have developed specialized solutions (e.g., geometry regular...

START: Self-taught Reasoner with Tools 08.03.2025

🤗 Upvotes: 49 | cs. CL Authors: Chengpeng Li, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang, Beichen Zhang, Xiang Wang, Bowen Yu, Binyuan Hui, Junyang Lin, Dayiheng Liu Title: START: Self-taught Reasoner with Tools Arxiv: http://arxiv.org/abs/2503.04625v1 Abstract: Large reasoning models (LRMs) like OpenAI-o1 and DeepSeek-R1 have demonstrated remarkable capabilities in complex reasoning tasks through th...

Token-Efficient Long Video Understanding for Multimodal LLMs 08.03.2025

🤗 Upvotes: 41 | cs. CV Authors: Jindong Jiang, Xiuyu Li, Zhijian Liu, Muyang Li, Guo Chen, Zhiqi Li, De-An Huang, Guilin Liu, Zhiding Yu, Kurt Keutzer, Sungjin Ahn, Jan Kautz, Hongxu Yin, Yao Lu, Song Han, Wonmin Byeon Title: Token-Efficient Long Video Understanding for Multimodal LLMs Arxiv: http://arxiv.org/abs/2503.04130v1 Abstract: Recent advances in video-based multimodal large language mode...

LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM 08.03.2025

🤗 Upvotes: 33 | cs. CL Authors: Sambal Shikhar, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jean Lahoud, Fahad Khan, Rao Muhammad Anwer, Salman Khan, Hisham Cholakkal Title: LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM Arxiv: http://arxiv.org/abs/2503.04724v1 Abstract: Recent advancements in speech-to-speech dialogue systems leverage LLMs for multimodal interactions, yet...

EgoLife: Towards Egocentric Life Assistant 08.03.2025

🤗 Upvotes: 21 | cs. CV Authors: Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Joerg Widmer, Francesco Gringoli, Lei Yang, Bo Li, Ziwei Liu Title: EgoLife: Towards Egocentric Life Assistant Arxiv: http://arxiv...

Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers 07.03.2025

🤗 Upvotes: 42 | cs. CL, cs. AI Authors: Yiran Zhao, Chaoqun Liu, Yue Deng, Jiahao Ying, Mahani Aljunied, Zhaodonghui Li, Lidong Bing, Hou Pong Chan, Yu Rong, Deli Zhao, Wenxuan Zhang Title: Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers Arxiv: http://arxiv.org/abs/2503.00865v1 Abstract: Large language models (LLMs) have revolutionized natural language processin...

HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs 07.03.2025

🤗 Upvotes: 27 | cs. CL, cs. HC Authors: Tin Nguyen, Logan Bolton, Mohammad Reza Taesiri, Anh Totti Nguyen Title: HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs Arxiv: http://arxiv.org/abs/2503.02003v2 Abstract: An Achilles heel of Large Language Models (LLMs) is their tendency to hallucinate non-factual statements. A response mixed of factual and non-factual statem...

Process-based Self-Rewarding Language Models 07.03.2025

🤗 Upvotes: 27 | cs. CL, cs. AI Authors: Shimao Zhang, Xiao Liu, Xin Zhang, Junxiao Liu, Zheheng Luo, Shujian Huang, Yeyun Gong Title: Process-based Self-Rewarding Language Models Arxiv: http://arxiv.org/abs/2503.03746v1 Abstract: Large Language Models have demonstrated outstanding performance across various downstream tasks and have been widely applied in multiple scenarios. Human-annotated prefe...

Visual-RFT: Visual Reinforcement Fine-Tuning 05.03.2025

🤗 Upvotes: 44 | cs. CV Authors: Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang Title: Visual-RFT: Visual Reinforcement Fine-Tuning Arxiv: http://arxiv.org/abs/2503.01785v1 Abstract: Reinforcement Fine-Tuning (RFT) in Large Reasoning Models like OpenAI o1 learns from feedback on its answers, which is especially useful in applications when fine-tuning...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.