Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning 13.01.2026

🤗 Upvotes: 38 | cs. CL, cs. AI Authors: Qiguang Chen, Yantao Du, Ziniu Li, Jinhao Liu, Songyao Duan, Jiarui Guo, Minghao Liu, Jiaheng Liu, Tong Yang, Ge Zhang, Libo Qin, Wanxiang Che, Wenhao Huang Title: The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning Arxiv: http://arxiv.org/abs/2601.06002v1 Abstract: Large language models (LLMs) often fail to learn eff...

Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards 13.01.2026

🤗 Upvotes: 30 | cs. CL Authors: Jiajie Zhang, Xin Lv, Ling Feng, Lei Hou, Juanzi Li Title: Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards Arxiv: http://arxiv.org/abs/2601.06021v1 Abstract: Reinforcement learning (RL) has emerged as a critical technique for enhancing LLM-based deep search agents. However, existing approaches primarily...

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis 13.01.2026

🤗 Upvotes: 25 | cs. CL, cs. AI, cs. LG Authors: Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Zhicheng Dou, Ji-Rong Wen Title: EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Arxiv: http://arxiv.org/abs/2601.05808v1 Abstract: Large language models (LLMs) are expected to be trained to act as agents in various real-world environments, but this pro...

Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking 13.01.2026

🤗 Upvotes: 22 | cs. CL Authors: Mingxin Li, Yanzhao Zhang, Dingkun Long, Keqin Chen, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, Junyang Lin Title: Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking Arxiv: http://arxiv.org/abs/2601.04720v1 Abstract: In this report, we introduce the Qwen3-VL-Em...

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization 10.01.2026

🤗 Upvotes: 98 | cs. CL, cs. AI, cs. LG Authors: Shih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao, Peter Belcak, Mingjie Liu, Min-Hung Chen, Hongxu Yin, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Yejin Choi, Jan Kautz, Pavlo Molchanov Title: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Arxiv: http://arxiv.org/abs/2601.05242v1 Abstract: As language mod...

Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers 10.01.2026

🤗 Upvotes: 29 | cs. LG Authors: Maksim Velikanov, Ilyas Chahed, Jingwei Zuo, Dhia Eddine Rhaiem, Younes Belkada, Hakim Hacid Title: Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers Arxiv: http://arxiv.org/abs/2601.04890v1 Abstract: Applying weight decay (WD) to matrix layers is standard practice in large-language-model pretraining. Prior work suggests that stochastic gradi...

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes 10.01.2026

🤗 Upvotes: 26 | cs. CV Authors: Yuan-Kang Lee, Kuan-Lin Chen, Chia-Che Chang, Yu-Lun Liu Title: RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes Arxiv: http://arxiv.org/abs/2601.05249v1 Abstract: Nighttime color constancy remains a challenging problem in computational photography due to low-light noise and complex illumination conditions. We pre...

Token-Level LLM Collaboration via FusionRoute 10.01.2026

🤗 Upvotes: 26 | cs. AI, cs. CL, cs. LG Authors: Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao Title: Token-Level LLM Collaboration via FusionRoute Arxiv: http://arxiv.org/abs/2601.05106v1 Abstract: Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a sing...

Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting 09.01.2026

🤗 Upvotes: 67 | cs. LG, cs. AI, cs. CL Authors: Muxi Diao, Lele Yang, Wuxuan Gong, Yutong Zhang, Zhonghao Yan, Yufei Han, Kongming Liang, Weiran Xu, Zhanyu Ma Title: Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting Arxiv: http://arxiv.org/abs/2601.02151v1 Abstract: Supervised Fine-Tuning (SFT) is the standard paradigm for domain adaptation, yet it frequently incu...

Evolving Programmatic Skill Networks 09.01.2026

🤗 Upvotes: 56 | cs. AI, cs. NE Authors: Haochen Shi, Xingdi Yuan, Bang Liu Title: Evolving Programmatic Skill Networks Arxiv: http://arxiv.org/abs/2601.03509v1 Abstract: We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding library of executable skills. We introduce the Programmatic Skill Network (PSN), a framework i...

Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning 09.01.2026

🤗 Upvotes: 31 | cs. CL Authors: Jinyang Wu, Guocheng Zhai, Ruihan Jin, Jiahao Yuan, Yuhao Shen, Shuai Zhang, Zhengqi Wen, Jianhua Tao Title: Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning Arxiv: http://arxiv.org/abs/2601.03872v1 Abstract: The integration of large language models (LLMs) with external tools has significantly expanded the capabilities of AI ag...

Benchmark^2: Systematic Evaluation of LLM Benchmarks 09.01.2026

🤗 Upvotes: 29 | cs. CL Authors: Qi Qian, Chengsong Huang, Jingwen Xu, Changze Lv, Muling Wu, Wenhao Liu, Xiaohua Wang, Zhenghua Wang, Zisu Huang, Muzhao Tian, Jianhan Xu, Kun Hu, He-Da Wang, Yao Hu, Xuanjing Huang, Xiaoqing Zheng Title: Benchmark^2: Systematic Evaluation of LLM Benchmarks Arxiv: http://arxiv.org/abs/2601.03986v1 Abstract: The rapid proliferation of benchmarks for evaluating large...

InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields 08.01.2026

🤗 Upvotes: 73 | cs. CV Authors: Hao Yu, Haotong Lin, Jiawei Wang, Jiaxin Li, Yida Wang, Xueyang Zhang, Yue Wang, Xiaowei Zhou, Ruizhen Hu, Sida Peng Title: InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields Arxiv: http://arxiv.org/abs/2601.03252v1 Abstract: Existing depth estimation methods are fundamentally limited to predicting depth on discrete imag...

LTX-2: Efficient Joint Audio-Visual Foundation Model 08.01.2026

🤗 Upvotes: 47 | cs. CV Authors: Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, Eitan Richardson, Guy Shiran, Itay Chachy, Jonathan Chetboun, Michael Finkelson, Michael Kupchick, Nir Zabari, Nitzan Guetta, Noa Kotler, Ofir Bibi, Ori Gordon, Poriya Panet, Roi Benita, Shahar Armon, Victor Kuli...

MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization 08.01.2026

🤗 Upvotes: 46 | cs. SD, cs. AI, eess. AS Authors: MOSI. AI, Donghua Yu, Zhengyuan Lin, Chen Yang, Yiyang Zhang, Hanfu Chen, Jingqi Chen, Ke Chen, Liwei Fan, Yi Jiang, Jie Zhu, Muchen Li, Wenxuan Wang, Yang Wang, Zhe Xu, Yitian Gong, Yuqian Zhang, Wenbo Zhang, Zhaoye Fei, Qinyuan Cheng, Shimin Li, Xipeng Qiu Title: MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization Arxiv: htt...

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence 08.01.2026

🤗 Upvotes: 28 | cs. AI, cs. CL Authors: Yiheng Wang, Yixin Chen, Shuo Li, Yifan Zhou, Bo Liu, Hengjian Gao, Jiakang Yuan, Jia Bu, Wanghan Xu, Yuhao Zhou, Xiangyu Zhao, Zhiwang Zhou, Fengxiang Wang, Haodong Duan, Songyang Zhang, Jun Yao, Han Deng, Yizhou Wang, Jiabei Xiao, Jiaqi Liu, Encheng Su, Yujie Liu, Weida Wang, Junchi Yao, Shenghe Zheng, Haoran Sun, Runmin Ma, Xiangchao Yan, Bo Zhang, Dongz...

NitroGen: An Open Foundation Model for Generalist Gaming Agents 08.01.2026

🤗 Upvotes: 22 | cs. CV, cs. AI, cs. LG Authors: Loïc Magne, Anas Awadalla, Guanzhi Wang, Yinzhen Xu, Joshua Belofsky, Fengyuan Hu, Joohwan Kim, Ludwig Schmidt, Georgia Gkioxari, Jan Kautz, Yisong Yue, Yejin Choi, Yuke Zhu, Linxi "Jim" Fan Title: NitroGen: An Open Foundation Model for Generalist Gaming Agents Arxiv: http://arxiv.org/abs/2601.02427v1 Abstract: We introduce NitroGen, a vision-action...

Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits 07.01.2026

🤗 Upvotes: 48 | cs. CL Authors: Amirhosein Ghasemabadi, Di Niu Title: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits Arxiv: http://arxiv.org/abs/2512.20578v2 Abstract: Large language models (LLMs) generate fluent and complex outputs but often fail to recognize their own mistakes and hallucinations. Existing approaches typically rely on external judges, multi-sample cons...

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation 07.01.2026

🤗 Upvotes: 45 | cs. CV, cs. AI Authors: Huichao Zhang, Liao Qu, Yiheng Liu, Hang Chen, Yangyang Song, Yongsheng Dong, Shikun Sun, Xian Li, Xu Wang, Yi Jiang, Hu Ye, Bo Chen, Yiming Gao, Peng Liu, Akide Liu, Zhipeng Yang, Qili Deng, Linjie Xing, Jiyang Liu, Zhao Wang, Yang Zhou, Mingcong Liu, Yi Zhang, Qian He, Xiwei Hu, Zhongqi Qi, Jie Shao, Zhiye Fu, Shuai Wang, Fangmin Chen, Xuezhi Chai, Zhihua...

DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer 07.01.2026

🤗 Upvotes: 36 | cs. CV Authors: Xu Guo, Fulong Ye, Xinghui Li, Pengqi Tu, Pengze Zhang, Qichao Sun, Songtao Zhao, Xiangwang Hou, Qian He Title: DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer Arxiv: http://arxiv.org/abs/2601.01425v1 Abstract: Video Face Swapping (VFS) requires seamlessly injecting a source identity into a target video while meti...

VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation 07.01.2026

🤗 Upvotes: 29 | cs. CV, cs. LG Authors: Shikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu, Yangyang Song, Xian Li, Xu Wang, Yi Jiang, Daniel K. Du, Xinglong Wu, Jia Jia Title: VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation Arxiv: http://arxiv.org/abs/2601.02256v1 Abstract: Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion...

GARDO: Reinforcing Diffusion Models without Reward Hacking 07.01.2026

🤗 Upvotes: 23 | cs. LG, cs. AI, cs. CV Authors: Haoran He, Yuxiao Ye, Jie Liu, Jiajun Liang, Zhiyong Wang, Ziyang Yuan, Xintao Wang, Hangyu Mao, Pengfei Wan, Ling Pan Title: GARDO: Reinforcing Diffusion Models without Reward Hacking Arxiv: http://arxiv.org/abs/2512.24138v1 Abstract: Fine-tuning diffusion models via online reinforcement learning (RL) has shown great potential for enhancing text-to...

InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams 07.01.2026

🤗 Upvotes: 22 | cs. CV Authors: Shuai Yuan, Yantai Yang, Xiaotian Yang, Xupeng Zhang, Zhonghao Zhao, Lingming Zhang, Zhipeng Zhang Title: InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams Arxiv: http://arxiv.org/abs/2601.02281v1 Abstract: The grand vision of enabling persistent, large-scale 3D visual geometry understanding is shackled by the irreconcilable demands of scalabil...

VINO: A Unified Visual Generator with Interleaved OmniModal Context 07.01.2026

🤗 Upvotes: 22 | cs. CV Authors: Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye Title: VINO: A Unified Visual Generator with Interleaved OmniModal Context Arxiv: http://arxiv.org/abs/2601.02358v1 Abstract: We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or indepe...

Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization 06.01.2026

🤗 Upvotes: 88 | cs. AI Authors: Yuchen Shi, Yuzheng Cai, Siqi Cai, Zihan Xu, Lichao Chen, Yulei Qin, Zhijian Zhou, Xiang Fei, Chaofan Qiu, Xiaoyu Tan, Gang Li, Zongyi Li, Haojia Lin, Guocan Cai, Yong Mao, Yunsheng Wu, Ke Li, Xing Sun Title: Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization Arxiv: http://arxiv.org/abs/2512.24615v1 Abstract: Existing L...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.