Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass 24.01.2025

🤗 Upvotes: 3 | cs. CV, cs. AI, cs. GR, cs. RO Authors: Jianing Yang, Alexander Sax, Kevin J. Liang, Mikael Henaff, Hao Tang, Ang Cao, Joyce Chai, Franziska Meier, Matt Feiszli Title: Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass Arxiv: http://arxiv.org/abs/2501.13928v1 Abstract: Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in appli...

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training 23.01.2025

🤗 Upvotes: 61 | cs. AI Authors: Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, Jiecao Chen Title: Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Arxiv: http://arxiv.org/abs/2501.11425v1 Abstract: Large Language Models (LLMs) agents are increasingly pivotal for addressing complex tasks in interactive environments. Existing work mainly focuses on enhancin...

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 23.01.2025

🤗 Upvotes: 59 | cs. CV, cs. AI, cs. CL Authors: Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, Arman Cohan Title: MMVU: Measuring Expert-Level Multi-Discipline Video Understanding Arxiv: http://arxiv.org/abs/2501.1238...

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models 23.01.2025

🤗 Upvotes: 51 | cs. LG, cs. CL Authors: Zihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin Title: Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Arxiv: http://arxiv.org/abs/2501.11873v1 Abstract: This paper revisits the implementation of $\textbf{L}$oad-$\textbf{b}$alanc...

TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space 23.01.2025

🤗 Upvotes: 32 | cs. CV Authors: Daniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov, Shiran Zada, Ariel Ephrat, Tomer Michaeli, Inbar Mosseri, Tali Dekel Title: TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space Arxiv: http://arxiv.org/abs/2501.12224v1 Abstract: We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image dif...

UI-TARS: Pioneering Automated GUI Interaction with Native Agents 23.01.2025

🤗 Upvotes: 31 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian...

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 23.01.2025

🤗 Upvotes: 26 | cs. CV, cs. CL Authors: Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang Title: InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Arxiv: http://arxiv.org/abs/2501.12368v1 Abstract: Despite the promising performance of Large Vision Language Models (L...

Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks 23.01.2025

🤗 Upvotes: 20 | cs. CL, cs. CV Authors: Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, Heng Ji Title: Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks Arxiv: http://arxiv.org/abs/2501.11733v1 Abstract: Smartphones have become indispensable in modern life, yet navigating complex tasks on mobile devices often remains frustrating. Recent advancem...

Reasoning Language Models: A Blueprint 23.01.2025

🤗 Upvotes: 18 | cs. AI, cs. CL Authors: Maciej Besta, Julia Barth, Eric Schreiber, Ales Kubicek, Afonso Catarino, Robert Gerstenberger, Piotr Nyczyk, Patrick Iff, Yueling Li, Sam Houliston, Tomasz Sternal, Marcin Copik, Grzegorz Kwaśniewski, Jürgen Müller, Łukasz Flis, Hannes Eberhard, Hubert Niewiadomski, Torsten Hoefler Title: Reasoning Language Models: A Blueprint Arxiv: http://arxiv.org/abs/2...

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation 23.01.2025

🤗 Upvotes: 16 | cs. CV Authors: Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, Huiwen Shi, Sicong Liu, Junta Wu, Yihang Lian, Fan Yang, Ruining Tang, Zebin He, Xinzhou Wang, Jian Liu, Xuhui Zuo, Zhuo Chen, Biwen Lei, Haohan Weng, Jing Xu, Yiling Zhu, Xinhai Liu, Lixin Xu, Changrong Hu, Tianyu Huang, Lifu Wang, Jih...

Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments 23.01.2025

🤗 Upvotes: 15 | cs. LG, cs. AI Authors: Hongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, Sercan Ö. Arık Title: Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments Arxiv: http://arxiv.org/abs/2501.10893v1 Abstract: Autonomous agents powered by large language models (LLMs) have the potential to enhance human capabilities, assisting with digital...

GameFactory: Creating New Games with Generative Interactive Videos 22.01.2025

🤗 Upvotes: 48 | cs. CV Authors: Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, Xihui Liu Title: GameFactory: Creating New Games with Generative Interactive Videos Arxiv: http://arxiv.org/abs/2501.08325v1 Abstract: Generative game engines have the potential to revolutionize game development by autonomously creating new content and reducing manual workload. However, existing video-based g...

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos 22.01.2025

🤗 Upvotes: 8 | cs. CV Authors: Zhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao, Bingyi Kang, Jiashi Feng, Xiaojie Jin Title: VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Arxiv: http://arxiv.org/abs/2501.09781v1 Abstract: This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models li...

SEAL: Entangled White-box Watermarks on Low-Rank Adaptation 22.01.2025

🤗 Upvotes: 2 | cs. AI, cs. CR Authors: Giyeong Oh, Saejin Kim, Woohyun Cho, Sangkyu Lee, Jiwan Chung, Dokyung Song, Youngjae Yu Title: SEAL: Entangled White-box Watermarks on Low-Rank Adaptation Arxiv: http://arxiv.org/abs/2501.09284v2 Abstract: Recently, LoRA and its variants have become the de facto strategy for training and sharing task-specific versions of large pretrained models, thanks to t...

The Lessons of Developing Process Reward Models in Mathematical Reasoning 15.01.2025

🤗 Upvotes: 53 | cs. CL, cs. AI, cs. LG Authors: Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin Title: The Lessons of Developing Process Reward Models in Mathematical Reasoning Arxiv: http://arxiv.org/abs/2501.07301v1 Abstract: Process Reward Models (PRMs) emerge as a promising approach for process supervision in mathematical re...

Tensor Product Attention Is All You Need 15.01.2025

🤗 Upvotes: 38 | cs. CL, cs. AI, cs. LG Authors: Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Zhen Qin, Yang Yuan, Quanquan Gu, Andrew Chi-Chih Yao Title: Tensor Product Attention Is All You Need Arxiv: http://arxiv.org/abs/2501.06425v1 Abstract: Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during in...

$\text{Transformer}^2$: Self-adaptive LLMs 15.01.2025

🤗 Upvotes: 25 | cs. LG, cs. AI, cs. CL Authors: Qi Sun, Edoardo Cetin, Yujin Tang Title: $\text{Transformer}^2$: Self-adaptive LLMs Arxiv: http://arxiv.org/abs/2501.06252v2 Abstract: Self-adaptive large language models (LLMs) aim to solve the challenges posed by traditional fine-tuning methods, which are often computationally intensive and static in their ability to handle diverse tasks. We intro...

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction 15.01.2025

🤗 Upvotes: 21 | cs. CL, cs. AI, cs. HC, cs. SD, eess. AS Authors: Qian Chen, Yafeng Chen, Yanni Chen, Mengzhe Chen, Yingda Chen, Chong Deng, Zhihao Du, Ruize Gao, Changfeng Gao, Zhifu Gao, Yabin Li, Xiang Lv, Jiaqing Liu, Haoneng Luo, Bin Ma, Chongjia Ni, Xian Shi, Jialong Tang, Hui Wang, Hao Wang, Wen Wang, Yuxuan Wang, Yunlan Xu, Fan Yu, Zhijie Yan, Yexin Yang, Baosong Yang, Xian Yang, Guanrou...

VideoAuteur: Towards Long Narrative Video Generation 15.01.2025

🤗 Upvotes: 21 | cs. CV Authors: Junfei Xiao, Feng Cheng, Lu Qi, Liangke Gui, Jiepeng Cen, Zhibei Ma, Alan Yuille, Lu Jiang Title: VideoAuteur: Towards Long Narrative Video Generation Arxiv: http://arxiv.org/abs/2501.06173v1 Abstract: Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in g...

O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning 15.01.2025

🤗 Upvotes: 18 | cs. CL Authors: Zhongzhen Huang, Gui Geng, Shengyi Hua, Zhen Huang, Haoyang Zou, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang Title: O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning Arxiv: http://arxiv.org/abs/2501.06458v1 Abstract: Building upon our previous investigations of O1 replication (Part 1: Journey Learning [Qin et al., 2024] and Part 2: Disti...

WebWalker: Benchmarking LLMs in Web Traversal 15.01.2025

🤗 Upvotes: 16 | cs. CL, cs. AI Authors: Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, Deyu Zhou, Pengjun Xie, Fei Huang Title: WebWalker: Benchmarking LLMs in Web Traversal Arxiv: http://arxiv.org/abs/2501.07572v2 Abstract: Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. H...

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training 15.01.2025

🤗 Upvotes: 12 | cs. LG, cs. AI, cs. CL Authors: Tianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu, Zhangyang Wang, Shiwei Liu Title: SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training Arxiv: http://arxiv.org/abs/2501.06842v1 Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks, yet their training remains highly resource-intensive and su...

UnCommon Objects in 3D 15.01.2025

🤗 Upvotes: 8 | cs. CV, cs. AI, cs. GR Authors: Xingchen Liu, Piyush Tayal, Jianyuan Wang, Jesus Zarzar, Tom Monnier, Konstantinos Tertikas, Jiali Duan, Antoine Toisoul, Jason Y. Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, David Novotny Title: UnCommon Objects in 3D Arxiv: http://arxiv.org/abs/2501.07574v1 Abstract: We introduce Uncommon Objects in 3D (uCO3D), a new object-centric d...

VideoRAG: Retrieval-Augmented Generation over Video Corpus 14.01.2025

🤗 Upvotes: 43 | cs. CV, cs. AI, cs. CL, cs. IR, cs. LG Authors: Soyeong Jeong, Kangsan Kim, Jinheon Baek, Sung Ju Hwang Title: VideoRAG: Retrieval-Augmented Generation over Video Corpus Arxiv: http://arxiv.org/abs/2501.05874v1 Abstract: Retrieval-Augmented Generation (RAG) is a powerful strategy to address the issue of generating factually incorrect outputs in foundation models by retrieving exte...

OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 14.01.2025

🤗 Upvotes: 29 | cs. CV, cs. AI Authors: Yifei Li, Junbo Niu, Ziyang Miao, Chunjiang Ge, Yuanhang Zhou, Qihao He, Xiaoyi Dong, Haodong Duan, Shuangrui Ding, Rui Qian, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang Title: OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? Arxiv: http://arxiv.org/abs/2501.05510v1 Abstract: Temporal Awareness, the ability to...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.