Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
MAXS: Meta-Adaptive Exploration with LLM Agents 16.01.2026 21:42
🤗 Upvotes: 82 | cs. AI Authors: Jian Zhang, Zhiyuan Wang, Zhangqi Wang, Yu He, Haoran Luo, li yuan, Lingling Zhang, Rui Mao, Qika Lin, Jun Liu Title: MAXS: Meta-Adaptive Exploration with LLM Agents Arxiv: http://arxiv.org/abs/2601.09259v1 Abstract: Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, e...
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning 16.01.2026 20:35
🤗 Upvotes: 47 | cs. LG, cs. CL Authors: Shaotian Yan, Kaiyuan Liu, Chen Shen, Bing Wang, Sinan Fan, Jun Zhang, Yue Wu, Zheng Wang, Jieping Ye Title: Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning Arxiv: http://arxiv.org/abs/2601.09088v1 Abstract: In this report, we introduce DASD-4B-Thinking, a lightweight yet highly capable, fully open-source reasoning model. It achie...
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning 16.01.2026 24:42
🤗 Upvotes: 41 | cs. CV, cs. AI, cs. LG, cs. RO Authors: Chi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen, Jan Kautz, Yu-Chiang Frank Wang, Fu-En Yang Title: Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning Arxiv: http://arxiv.org/abs/2601.09708v1 Abstract: Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing ada...
SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL 16.01.2026 26:56
🤗 Upvotes: 36 | cs. CV, cs. AI Authors: Lijun Liu, Linwei Chen, Zhishou Zhang, Meng Tian, Hengfu Cui, Ruiyang Li, Zhaocheng Liu, Qiang Ju, Qianxi Li, Hong-Yu Zhou Title: SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL Arxiv: http://arxiv.org/abs/2601.09136v1 Abstract: General-purpose Large Vision-Language Models (LVLMs), des...
OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG 16.01.2026 25:28
🤗 Upvotes: 26 | cs. CL, cs. AI, cs. IR Authors: Fengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang, Jia Ao Sun, Zheyuan Liu, Chao Zhang, Tetsuya Sakai, Jian-Yun Nie Title: OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG Arxiv: http://arxiv.org/abs/2601.09028v1 Abstract: The development of large language models (LLMs) has achieved superior performance in a range...
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding 16.01.2026 23:20
🤗 Upvotes: 22 | cs. CV Authors: Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun Title: OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding Arxiv: http://arxiv.org/abs/2601.09575v1 Abstract: We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Give...
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences 15.01.2026 23:25
🤗 Upvotes: 62 | cs. SE, cs. AI Authors: Qihao Wang, Ziming Cheng, Shuo Zhang, Fan Liu, Rui Xu, Heng Lian, Kunyi Wang, Xiaoming Yu, Jianghao Yin, Sen Hu, Yue Hu, Shaolei Zhang, Yanbing Liu, Ronghao Chen, Huacan Wang Title: MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences Arxiv: http://arxiv.org/abs/2601.06789v2 Abstract: While autonomous software engineering (SWE)...
Solar Open Technical Report 15.01.2026 21:06
🤗 Upvotes: 53 | cs. CL Authors: Sungrae Park, Sanghoon Kim, Jungho Cho, Gyoungjin Gim, Dawoon Jung, Mikyoung Cha, Eunhae Choo, Taekgyu Hong, Minbyul Jeong, SeHwan Joo, Minsoo Khang, Eunwon Kim, Minjeong Kim, Sujeong Kim, Yunsu Kim, Hyeonju Lee, Seunghyun Lee, Sukyung Lee, Siyoung Park, Gyungin Shin, Inseo Song, Wonho Song, Seonghoon Yang, Seungyoun Yi, Sanghoon Yoon, Jeonghyun Ko, Seyoung Song, K...
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions 15.01.2026 22:06
🤗 Upvotes: 47 | cs. AI, cs. IR Authors: Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Chenglong Li, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen Title: KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions Arxiv: http://arxiv.org/abs/2601.04745v1 Abstract: Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user...
User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale 15.01.2026 23:26
🤗 Upvotes: 41 | cs. CL Authors: Jungho Cho, Minbyul Jeong, Sungrae Park Title: User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale Arxiv: http://arxiv.org/abs/2601.08225v1 Abstract: The recent paradigm shift toward large reasoning models (LRMs) as autonomous agents has intensified the demand for sophisticated, multi-turn tool-use capabilities. Yet, existing datasets and data-gener...
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands 15.01.2026 22:20
🤗 Upvotes: 37 | cs. CV, cs. AI, cs. HC Authors: Siyuan Hu, Kevin Qinghong Lin, Mike Zheng Shou Title: ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands Arxiv: http://arxiv.org/abs/2512.24965v1 Abstract: Building intelligent agents capable of dexterous manipulation is essential for achieving human-like automation in both robotics and digital environments. However, existing GUI agents...
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking 15.01.2026 26:01
🤗 Upvotes: 37 | cs. LG, cs. AI Authors: Qiang Zhang, Boli Chen, Fanrui Zhang, Ruixue Ding, Shihang Wang, Qiuchen Wang, Yinfeng Huang, Haonan Zhang, Rongxiang Zhu, Pengyong Wang, Ailin Ren, Xin Li, Pengjun Xie, Jiawei Liu, Ning Guo, Jingren Zhou, Zheng-Jun Zha Title: ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking Arxiv: http://arxiv.org/abs/2601.06487v1 Abstract: R...
MemoBrain: Executive Memory as an Agentic Brain for Reasoning 15.01.2026 22:27
🤗 Upvotes: 33 | cs. AI, cs. CL, cs. IR Authors: Hongjin Qian, Zhao Cao, Zheng Liu Title: MemoBrain: Executive Memory as an Agentic Brain for Reasoning Arxiv: http://arxiv.org/abs/2601.08079v1 Abstract: Complex reasoning in tool-augmented agent frameworks is inherently long-horizon, causing reasoning traces and transient tool artifacts to accumulate and strain the bounded working context of large...
Motion Attribution for Video Generation 15.01.2026 20:06
🤗 Upvotes: 26 | cs. CV, cs. AI, cs. LG, cs. MM, cs. RO Authors: Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taixé, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine Title: Motion Attribution for Video Generation Arxiv: http://arxiv.org/abs/2601.08828v1 Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorly unders...
3AM: Segment Anything with Geometric Consistency in Videos 15.01.2026 22:38
🤗 Upvotes: 24 | cs. CV Authors: Yang-Che Sun, Cheng Sun, Chin-Yang Lin, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu Title: 3AM: Segment Anything with Geometric Consistency in Videos Arxiv: http://arxiv.org/abs/2601.08831v1 Abstract: Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes due to reli...
BabyVision: Visual Reasoning Beyond Language 14.01.2026 22:04
🤗 Upvotes: 156 | cs. CV, cs. CL Authors: Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu Liu, Tianyu Liu, Baobao Chang, Xiaobo Hu, Kaiyuan Chen, Yixin Ren, Yang Liu, Yuan Gong, Kuan Li Title:...
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning 14.01.2026 23:17
🤗 Upvotes: 65 | cs. LG Authors: Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Daxin Jiang, Xiangyu Zhang, Heung-Yeung Shum Title: PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning Arxiv: http://arxiv.org/abs/...
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head 14.01.2026 22:26
🤗 Upvotes: 32 | cs. CV, cs. AI Authors: Kewei Zhang, Ye Huang, Yufan Deng, Jincheng Yu, Junsong Chen, Huan Ling, Enze Xie, Daquan Zhou Title: MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head Arxiv: http://arxiv.org/abs/2601.07832v1 Abstract: While the Transformer architecture dominates many fields, its quadratic self-attention complexity hinders its use in large-scale a...
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests 14.01.2026 22:30
🤗 Upvotes: 30 | cs. CL, cs. LG Authors: Jie Wu, Haoling Li, Xin Zhang, Jiani Guo, Jane Luo, Steven Liu, Yangyu Huang, Ruihang Chu, Scarlett Li, Yujiu Yang Title: X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests Arxiv: http://arxiv.org/abs/2601.06953v1 Abstract: Competitive programming presents great challenges for Code LLMs due to its intensive reasoning...
GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts 14.01.2026 20:07
🤗 Upvotes: 26 | cs. AI Authors: Wenhao Zeng, Xuteng Zhang, Yuling Shi, Chao Hu, Yuting Chen, Beijun Shen, Xiaodong Gu Title: GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts Arxiv: http://arxiv.org/abs/2601.05110v1 Abstract: Large Reasoning Models (LRMs) achieve remarkable performance by explicitly generating multi-step chains of thought, but this capability incur...
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors 14.01.2026 23:35
🤗 Upvotes: 24 | cs. AI, cs. CL Authors: Seongyun Lee, Yongrae Jo, Minju Seo, Moontae Lee, Minjoon Seo Title: Lost in the Noise: How Reasoning Models Fail with Contextual Distractors Arxiv: http://arxiv.org/abs/2601.07226v1 Abstract: Recent advances in reasoning models and agentic AI systems have led to an increased reliance on diverse external information. However, this shift introduces input con...
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent 14.01.2026 24:29
🤗 Upvotes: 22 | cs. MA, cs. AI, cs. CL, cs. CV, cs. HC Authors: Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Qingyun Li, Yian Wang, Yu Qiao, Zun Wang, Zichen Ding Title: OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent Arxiv: http://arxiv.org/abs/2601.07779v1 Abstract: While Vision-L...
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization 13.01.2026 26:55
🤗 Upvotes: 131 | cs. CV, cs. AI, cs. CL Authors: Yuxiang Ji, Yong Wang, Ziyu Ma, Yiming Hu, Hailang Huang, Xuecai Hu, Guanhua Chen, Liaoni Wu, Xiangxiang Chu Title: Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization Arxiv: http://arxiv.org/abs/2601.05432v1 Abstract: The image geolocalization task aims to predict the location where an image was taken anywhere on Earth u...
MMFormalizer: Multimodal Autoformalization in the Wild 13.01.2026 21:58
🤗 Upvotes: 94 | cs. CL Authors: Jing Xiong, Qi Han, Yunta Hsieh, Hui Shen, Huajian Xin, Chaofan Tao, Chenyang Zhao, Hengyuan Zhang, Taiqiang Wu, Zhen Zhang, Haochen Wang, Zhongwei Wan, Lingpeng Kong, Ngai Wong Title: MMFormalizer: Multimodal Autoformalization in the Wild Arxiv: http://arxiv.org/abs/2601.03017v1 Abstract: Autoformalization, which translates natural language mathematics into formal...
CaricatureGS: Exaggerating 3D Gaussian Splatting Faces With Gaussian Curvature 13.01.2026 21:52
🤗 Upvotes: 45 | cs. GR, cs. AI, cs. LG Authors: Eldad Matmon, Amit Bracha, Noam Rotstein, Ron Kimmel Title: CaricatureGS: Exaggerating 3D Gaussian Splatting Faces With Gaussian Curvature Arxiv: http://arxiv.org/abs/2601.03319v1 Abstract: A photorealistic and controllable 3D caricaturization framework for faces is introduced. We start with an intrinsic Gaussian curvature-based surface exaggeration...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.