Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Spotlight on Token Perception for Multimodal Reinforcement Learning 15.10.2025 23:52
🤗 Upvotes: 31 | cs. CV Authors: Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng Title: Spotlight on Token Perception for Multimodal Reinforcement Learning Arxiv: http://arxiv.org/abs/2510.09285v1 Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in...
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment 15.10.2025 24:01
🤗 Upvotes: 31 | cs. LG, cs. AI, cs. CL Authors: Jinghao Zhang, Naishan Zheng, Ruilin Li, Dongzhou Cheng, Zheming Liang, Feng Zhao, Jiaqi Wang Title: RLFR: Extending Reinforcement Learning for LLMs with Flow Environment Arxiv: http://arxiv.org/abs/2510.10201v1 Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising framework for improving reasoning abili...
DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training 15.10.2025 22:13
🤗 Upvotes: 26 | cs. CV Authors: Haoran Feng, Dizhe Zhang, Xiangtai Li, Bo Du, Lu Qi Title: DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training Arxiv: http://arxiv.org/abs/2510.11712v1 Abstract: In this work, we propose DiT360, a DiT-based framework that performs hybrid training on perspective and panoramic data for panoramic image generation. For the issues of maintaining geometr...
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration 15.10.2025 24:24
🤗 Upvotes: 26 | cs. CV Authors: Xinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua, Linli Yao, Yang Shi, Bozhou Li, Yuanxing Zhang, Qiang Liu, Pengfei Wan, Liang Wang, Tieniu Tan Title: AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration Arxiv: http://arxiv.org/abs/2510.10395v1 Abstract: Audiovisual video captioning aims to generate semantically rich descriptions with temporal...
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models 15.10.2025 25:21
🤗 Upvotes: 25 | cs. CV Authors: Haomin Wang, Jinhui Yin, Qi Wei, Wenguang Zeng, Lixin Gu, Shenglong Ye, Zhangwei Gao, Yaohui Wang, Yanting Zhang, Yuanqi Li, Yanwen Guo, Wenhai Wang, Kai Chen, Yu Qiao, Hongjie Zhang Title: InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2510.11341v1 Abstract: General SVG modeling remains challenging due to fra...
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions 15.10.2025 21:38
🤗 Upvotes: 25 | cs. CL, cs. AI Authors: Tao Yu, Zhengbo Zhang, Zhiheng Lyu, Junhao Gong, Hongzhu Yi, Xinming Wang, Yuxuan Zhou, Jiabing Yang, Ping Nie, Yan Huang, Wenhu Chen Title: BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions Arxiv: http://arxiv.org/abs/2510.10666v2 Abstract: Efficiently solving real-world problems with LLMs increasingly hinges on their ability to in...
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI 14.10.2025 23:49
🤗 Upvotes: 104 | cs. AI, cs. CV, cs. RO Authors: Suwhan Choi, Jaeyoon Jung, Haebin Seong, Minchan Kim, Minyeong Kim, Yongjun Cho, Yoonshik Kim, Yubeen Park, Youngjae Yu, Yunsung Lee Title: D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI Arxiv: http://arxiv.org/abs/2510.05684v1 Abstract: Large language models leverage internet-scale text data, yet embodied AI rem...
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation 14.10.2025 23:18
🤗 Upvotes: 86 | cs. CV Authors: Kang Liao, Size Wu, Zhonghua Wu, Linyi Jin, Chao Wang, Yikai Wang, Fei Wang, Wei Li, Chen Change Loy Title: Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation Arxiv: http://arxiv.org/abs/2510.08673v1 Abstract: Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typicall...
TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling 14.10.2025 21:39
🤗 Upvotes: 38 | cs. CV Authors: Hyunmin Cho, Donghoon Ahn, Susung Hong, Jee Eun Kim, Seungryong Kim, Kyong Hwan Jin Title: TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling Arxiv: http://arxiv.org/abs/2510.04533v1 Abstract: Recent diffusion models achieve the state-of-the-art performance in image generation, but often suffer from semantic inconsistencies or halluci...
AutoPR: Let's Automate Your Academic Promotion! 14.10.2025 22:35
🤗 Upvotes: 38 | cs. CL Authors: Qiguang Chen, Zheng Yan, Mingda Yang, Libo Qin, Yixin Yuan, Hanjing Li, Jinhao Liu, Yiyan Ji, Dengyun Peng, Jiannan Guan, Mengkang Hu, Yantao Du, Wanxiang Che Title: AutoPR: Let's Automate Your Academic Promotion! Arxiv: http://arxiv.org/abs/2510.09558v1 Abstract: As the volume of peer-reviewed research surges, scholars increasingly rely on social platforms for dis...
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs 14.10.2025 25:37
🤗 Upvotes: 37 | cs. LG, cs. AI, cs. CL Authors: Yumin Choi, Dongki Kim, Jinheon Baek, Sung Ju Hwang Title: Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs Arxiv: http://arxiv.org/abs/2510.09201v1 Abstract: Large Language Models (LLMs) have shown remarkable success, and their multimodal expansions (MLLMs) further unlock capabilities spanning images, videos, and other...
BEAR: Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied Capabilities 14.10.2025 26:40
🤗 Upvotes: 30 | cs. CV, cs. RO Authors: Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma, Ziyan Chen, Yaokun Han, Renrui Zhang, Zitiantao Lin, Shiji Xin, Yijian Huang, Kai Cheng, Peiheng Wang, Jiazheng Liu, Jiayi Zhang, Yizhe Zhu, Wenqing Wang, Yiran Qin, Xupeng Zhu, Haojie Huang, Lawson L. S. Wong Title: BEAR: Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied Capabilities Arxiv: ht...
StreamingVLM: Real-Time Understanding for Infinite Video Streams 14.10.2025 21:22
🤗 Upvotes: 26 | cs. CV, cs. AI, cs. CL Authors: Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Kelly Peng, Yao Lu, Song Han Title: StreamingVLM: Real-Time Understanding for Infinite Video Streams Arxiv: http://arxiv.org/abs/2510.09608v1 Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite...
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels 14.10.2025 24:41
🤗 Upvotes: 22 | cs. CL, cs. AI Authors: Zhepeng Cen, Haolin Chen, Shiyu Wang, Zuxin Liu, Zhiwei Liu, Ding Zhao, Silvio Savarese, Caiming Xiong, Huan Wang, Weiran Yao Title: Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels Arxiv: http://arxiv.org/abs/2510.06499v1 Abstract: Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast...
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution 14.10.2025 22:59
🤗 Upvotes: 22 | cs. SE, cs. AI, cs. CL Authors: Terry Yue Zhuo, Xiaolong Jin, Hange Liu, Juyong Jiang, Tianyang Liu, Chen Gong, Bhupesh Bishnoi, Vaisakhi Mishra, Marek Suppa, Noah Ziems, Saiteja Utpala, Ming Xu, Guangyu Song, Kaixin Li, Yuhan Cao, Bo Liu, Zheng Liu, Sabina Abdurakhmanova, Wenhao Yu, Mengzhao Jia, Jihan Yao, Kenneth Hamilton, Kumar Shridhar, Minh Chien Vu, Dingmin Wang, Jiawei Liu...
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 14.10.2025 24:42
🤗 Upvotes: 22 | cs. AI, cs. CL Authors: Yi Lu, Jianing Wang, Linsen Guo, Wei He, Hongyin Tang, Tao Gui, Xuanjing Huang, Xuezhi Cao, Wei Wang, Xunliang Cai Title: R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? Arxiv: http://arxiv.org/abs/2510.08189v1 Abstract: Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek-R1) have led to remar...
Agent Learning via Early Experience 11.10.2025 22:45
🤗 Upvotes: 124 | cs. AI, cs. CL, cs. IR, cs. LG Authors: Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, Jian Xie, Yuxuan Sun, Boyu Gou, Qi Qi, Zihang Meng, Jianwei Yang, Ning Zhang, Xian Li, Ashish Shah, Dat Huynh, Hengduo Li, Zi Yang, Sara Cao, Lawrence Jang, Shuyan Zhou, Jiacheng Zhu, Huan Sun, Jason Weston, Yu Su, Yifan...
MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization 11.10.2025 21:36
🤗 Upvotes: 92 | cs. CV Authors: Xiangyu Zhao, Junming Lin, Tianhao Liang, Yifan Zhou, Wenhao Chai, Yuzhe Gu, Weiyun Wang, Kai Chen, Gen Luo, Wenwei Zhang, Junchi Yan, Hua Yang, Haodong Duan, Xue Yang Title: MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization Arxiv: http://arxiv.org/abs/2510.08540v1 Abstract: While current Mu...
MemMamba: Rethinking Memory Patterns in State Space Model 11.10.2025 24:07
🤗 Upvotes: 56 | cs. LG, cs. AI, cs. CL Authors: Youjin Wang, Yangjingyi Chen, Jiahao Yan, Jiaxuan Lu, Xiao Sun Title: MemMamba: Rethinking Memory Patterns in State Space Model Arxiv: http://arxiv.org/abs/2510.03279v1 Abstract: With the explosive growth of data, long-sequence modeling has become increasingly important in tasks such as natural language processing and bioinformatics. However, existi...
UniVideo: Unified Understanding, Generation, and Editing for Videos 11.10.2025 26:06
🤗 Upvotes: 47 | cs. CV Authors: Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhu Chen Title: UniVideo: Unified Understanding, Generation, and Editing for Videos Arxiv: http://arxiv.org/abs/2510.08377v1 Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. I...
From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning 11.10.2025 26:23
🤗 Upvotes: 42 | cs. AI, cs. CL Authors: Cheng Yang, Jiaxuan Lu, Haiyuan Wan, Junchi Yu, Feiwei Qin Title: From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning Arxiv: http://arxiv.org/abs/2509.23768v1 Abstract: The chemical reaction recommendation is to select proper reaction condition parameters for chemical reactions, which is pivotal to accelerating ch...
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs 11.10.2025 21:23
🤗 Upvotes: 38 | cs. CL, cs. AI, cs. LG Authors: Soyeong Jeong, Taehee Jung, Sung Ju Hwang, Joo-Kyung Kim, Dongyeop Kang Title: When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs Arxiv: http://arxiv.org/abs/2510.07499v1 Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge-intens...
Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning 11.10.2025 24:32
🤗 Upvotes: 38 | cs. LG, cs. AI Authors: Yoonjeon Kim, Doohyuk Jang, Eunho Yang Title: Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning Arxiv: http://arxiv.org/abs/2510.03259v1 Abstract: Recent studies on reasoning models explore the meta-awareness of language models, the ability to know how to think by itself. We argue that large reasoning models lack this meta-awar...
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning 11.10.2025 24:37
🤗 Upvotes: 36 | cs. CV Authors: Minghong Cai, Qiulin Wang, Zongli Ye, Wenze Liu, Quande Liu, Weicai Ye, Xintao Wang, Pengfei Wan, Kun Gai, Xiangyu Yue Title: VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning Arxiv: http://arxiv.org/abs/2510.08555v1 Abstract: We introduce the task of arbitrary spatio-temporal video completion, where a video is...
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety 11.10.2025 24:20
🤗 Upvotes: 33 | cs. CL Authors: Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, Benjamin Van Durme, Daniel Khashabi, Jason Weston, Hongyuan Zhan Title: The Alignment Waltz: Jointly Training Agents to Collaborate for Safety Arxiv: http://arxiv.org/abs/2510.08240v1 Abstract: Harnessing the power of LLMs requires a delicate dance between being helpful and harm...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.