Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective 04.07.2025 23:55
🤗 Upvotes: 24 | cs. RO Authors: Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, Yaodong Yang Title: A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Arxiv: http://arxiv.org/abs/2507.01925v1 Abstract: The remarkable advancements of vision and l...
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning 03.07.2025 24:47
🤗 Upvotes: 141 | cs. CV, cs. AI, cs. LG Authors: GLM-V Team, :, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiali Chen, Jing Che...
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning 03.07.2025 21:50
🤗 Upvotes: 35 | cs. AI, cs. CL Authors: Maggie Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu, Seungone Kim, Minxin Du, Radha Poovendran, Graham Neubig, Xiang Yue Title: Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Arxiv: http://arxiv.org/abs/2507.00432v1 Abstract: Math reasoning has become the poster child of progress in large language models (LLM...
SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks 03.07.2025 18:57
🤗 Upvotes: 33 | cs. CL, cs. AI Authors: Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu, Ronan Le Bras, Taira Anderson, Jonathan Bragg, Joseph Chee Chang, Jesse Dodge, Matt Latzke, Yixin Liu, Charles McGrady, Xiangru Tang, Zihang Wang, Chen Zhao, Hannaneh Hajishirzi, Doug Downey, Arman Cohan Title: SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks Arxiv:...
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings 03.07.2025 20:58
🤗 Upvotes: 30 | cs. CV, cs. AI, cs. CL Authors: Haonan Chen, Hong Liu, Yuping Luo, Liang Wang, Nan Yang, Furu Wei, Zhicheng Dou Title: MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Arxiv: http://arxiv.org/abs/2506.23115v1 Abstract: Multimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks. Howev...
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation 03.07.2025 20:53
🤗 Upvotes: 29 | cs. CV, cs. AI, cs. LG Authors: Xingyang Li, Muyang Li, Tianle Cai, Haocheng Xi, Shuo Yang, Yujun Lin, Lvmin Zhang, Songlin Yang, Jinbo Hu, Kelly Peng, Maneesh Agrawala, Ion Stoica, Kurt Keutzer, Song Han Title: Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation Arxiv: http://arxiv.org/abs/2506.19852v1 Abstract: Recent advances in diffusion...
Ovis-U1 Technical Report 02.07.2025 22:15
🤗 Upvotes: 51 | cs. CV, cs. AI Authors: Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang, Liangfu Cao, Pengxin Zhan, Lunhao Duan, Shiyin Lu, Minghao Fu, Xiaohao Chen, Jianshan Zhao, Yang Li, Qing-Guo Chen Title: Ovis-U1 Technical Report Arxiv: http://arxiv.org/abs/2506.23044v2 Abstract: In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, te...
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning 02.07.2025 20:38
🤗 Upvotes: 27 | cs. AI, cs. CL, cs. LG Authors: Bo Liu, Leon Guertler, Simon Yu, Zichen Liu, Penghui Qi, Daniel Balcells, Mickel Liu, Cheston Tan, Weiyan Shi, Min Lin, Wee Sun Lee, Natasha Jaques Title: SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning Arxiv: http://arxiv.org/abs/2506.24119v2 Abstract: Recent advances in reinforcement lea...
VMoBA: Mixture-of-Block Attention for Video Diffusion Models 02.07.2025 17:07
🤗 Upvotes: 26 | cs. CV Authors: Jianzong Wu, Liang Hou, Haotian Yang, Xin Tao, Ye Tian, Pengfei Wan, Di Zhang, Yunhai Tong Title: VMoBA: Mixture-of-Block Attention for Video Diffusion Models Arxiv: http://arxiv.org/abs/2506.23858v1 Abstract: The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high...
Calligrapher: Freestyle Text Image Customization 02.07.2025 22:33
🤗 Upvotes: 24 | cs. CV Authors: Yue Ma, Qingyan Bai, Hao Ouyang, Ka Leong Cheng, Qiuyu Wang, Hongyu Liu, Zichen Liu, Haofan Wang, Jingye Chen, Yujun Shen, Qifeng Chen Title: Calligrapher: Freestyle Text Image Customization Arxiv: http://arxiv.org/abs/2506.24123v1 Abstract: We introduce Calligrapher, a novel diffusion-based framework that innovatively integrates advanced text customization with ar...
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing 01.07.2025 22:20
🤗 Upvotes: 46 | cs. GR, cs. CV Authors: Jiacheng Chen, Ramin Mehran, Xuhui Jia, Saining Xie, Sanghyun Woo Title: BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing Arxiv: http://arxiv.org/abs/2506.17450v2 Abstract: We present BlenderFusion, a generative visual compositing framework that synthesizes new scenes by recomposing objects, camera, and background. It follows a layering-...
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs 01.07.2025 20:54
🤗 Upvotes: 30 | cs. CV, cs. AI, cs. HC, cs. MM Authors: Boyuan Sun, Jiaxing Zhao, Xihan Wei, Qibin Hou Title: LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Arxiv: http://arxiv.org/abs/2506.21862v1 Abstract: In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods m...
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation 01.07.2025 23:43
🤗 Upvotes: 25 | cs. CV Authors: Bowen Chen, Mengyi Zhao, Haomiao Sun, Li Chen, Xu Wang, Kang Du, Xinglong Wu Title: XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation Arxiv: http://arxiv.org/abs/2506.21416v1 Abstract: Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, partic...
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback 17.06.2025 25:01
🤗 Upvotes: 39 | cs. CL Authors: Dongwei Jiang, Alvin Zhang, Andrew Wang, Nicholas Andrews, Daniel Khashabi Title: Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback Arxiv: http://arxiv.org/abs/2506.11930v1 Abstract: Recent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoro...
Effective Red-Teaming of Policy-Adherent Agents 17.06.2025 19:35
🤗 Upvotes: 33 | cs. MA, cs. AI, cs. CL, cs. CR Authors: Itay Nakash, George Kour, Koren Lazar, Matan Vetzler, Guy Uziel, Ateret Anaby-Tavor Title: Effective Red-Teaming of Policy-Adherent Agents Arxiv: http://arxiv.org/abs/2506.09600v1 Abstract: Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lie...
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation 17.06.2025 21:27
🤗 Upvotes: 29 | cs. CV Authors: Min-Seop Kwak, Junho Kim, Sangdoo Yun, Dongyoon Han, Taekyoung Kim, Seungryong Kim, Jin-Hwa Kim Title: Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation Arxiv: http://arxiv.org/abs/2506.11924v1 Abstract: We introduce a diffusion-based framework that performs aligned novel view image and geometry generation via a warping-and-inpa...
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 14.06.2025 21:54
🤗 Upvotes: 63 | cs. CL, cs. AI, cs. MA Authors: Yu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang, Chenghao Xiao, Long Li, Yu Rong, Wenbing Huang, Qifeng Bai, Tingyang Xu Title: ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning Arxiv: http://arxiv.org/abs/2506.09513v1 Abstract: Though reasoning-based large language models (LLMs) have excelled in mathematics and programming,...
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks 14.06.2025 22:04
🤗 Upvotes: 40 | cs. SE, cs. AI Authors: Lianghong Guo, Yanlin Wang, Caihua Li, Pengyu Yang, Jiachi Chen, Wei Tao, Yingtian Zou, Duyu Tang, Zibin Zheng Title: SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks Arxiv: http://arxiv.org/abs/2506.10954v1 Abstract: Constructing large-scale datasets for the GitHub issue resolution task is crucial for both tr...
Text-Aware Image Restoration with Diffusion Models 14.06.2025 23:44
🤗 Upvotes: 34 | cs. CV, cs. AI, cs. LG Authors: Jaewon Min, Jin Hyeon Kim, Paul Hyunbin Cho, Jaeeun Lee, Jihye Park, Minkyu Park, Sangpil Kim, Hyunhee Park, Seungryong Kim Title: Text-Aware Image Restoration with Diffusion Models Arxiv: http://arxiv.org/abs/2506.09993v1 Abstract: Image restoration aims to recover degraded images. However, existing diffusion-based restoration methods, despite grea...
AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation 14.06.2025 20:09
🤗 Upvotes: 30 | cs. MA, cs. CV Authors: Haoyuan Shi, Yunxin Li, Xinyu Chen, Longyue Wang, Baotian Hu, Min Zhang Title: AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation Arxiv: http://arxiv.org/abs/2506.10540v1 Abstract: Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remain...
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos 14.06.2025 22:05
🤗 Upvotes: 29 | cs. CV, cs. AI, cs. MM Authors: Jiashuo Yu, Yue Wu, Meng Chu, Zhifei Ren, Zizheng Huang, Pei Chu, Ruijie Zhang, Yinan He, Qirui Li, Songze Li, Zhenxiang Li, Zhongying Tu, Conghui He, Yu Qiao, Yali Wang, Yi Wang, Limin Wang Title: VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Arxiv: http://arxiv.org/abs/2506.10857v1 Abstract: We present VRBench, the first l...
Discrete Audio Tokens: More Than a Survey! 14.06.2025 24:50
🤗 Upvotes: 24 | cs. SD, cs. AI, cs. CL, eess. AS Authors: Pooneh Mousavi, Gallil Maimon, Adel Moumen, Darius Petermann, Jiatong Shi, Haibin Wu, Haici Yang, Anastasia Kuznetsova, Artem Ploujnikov, Ricard Marxer, Bhuvana Ramabhadran, Benjamin Elizalde, Loren Lugosch, Jinyu Li, Cem Subakan, Phil Woodland, Minje Kim, Hung-yi Lee, Shinji Watanabe, Yossi Adi, Mirco Ravanelli Title: Discrete Audio Token...
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models 13.06.2025 20:39
🤗 Upvotes: 76 | cs. CL, cs. LG Authors: Pengyi Li, Matvey Skripkin, Alexander Zubrey, Andrey Kuznetsov, Ivan Oseledets Title: Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Arxiv: http://arxiv.org/abs/2506.06395v3 Abstract: Large language models (LLMs) excel at reasoning, yet post-training remains critical for aligning their behavior with task goals. Existing reinforcement...
Seedance 1.0: Exploring the Boundaries of Video Generation Models 13.06.2025 20:44
🤗 Upvotes: 49 | cs. CV Authors: Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xiaojie Li, Xunsong Li, Yifu Li, Shanchuan Lin, Zhijie Lin, Jiawei Liu, Shu Liu, Xiaonan Nie, Zhiwu Qing, Yuxi Ren, Li Sun, Zhi Tian, Rui Wang, Sen Wang, Guoqiang Wei, Guohong Wu, Jie Wu, Ruiqi Xia, Fei Xiao, Xuefeng Xiao, Jiangqiao Yan, Ceyuan Yang, Jianchao Ya...
Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation 13.06.2025 21:16
🤗 Upvotes: 37 | cs. LG Authors: Xinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen, Beidi Chen Title: Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation Arxiv: http://arxiv.org/abs/2506.09991v1 Abstract: Autoregressive Large Language Models (AR-LLMs) frequently exhibit implicit parallelism in sequential generation. Inspired by this, we introduce Multiverse, a new...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.