Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

WildDet3D: Scaling Promptable 3D Detection in the Wild 14.04.2026

🤗 Upvotes: 216 | cs. CV Authors: Weikai Huang, Jieyu Zhang, Sijun Li, Taoyang Jia, Jiafei Duan, Yunqian Cheng, Jaemin Cho, Mattew Wallingford, Rustin Soraki, Chris Dongjoo Kim, Donovan Clay, Taira Anderson, Winson Han, Ali Farhadi, Bharath Hariharan, Zhongzheng Ren, Ranjay Krishna Title: WildDet3D: Scaling Promptable 3D Detection in the Wild Arxiv: http://arxiv.org/abs/2604.08626v1 Abstract: Unde...

FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios 14.04.2026

🤗 Upvotes: 84 | cs. CV, cs. AI, cs. LG Authors: Xiangru Jian, Hao Xu, Wei Pang, Xinjian Zhao, Chengyu Tao, Qixin Zhang, Xikun Zhang, Chao Zhang, Guanzhi Deng, Alex Xue, Juan Du, Tianshu Yu, Garth Tarr, Linqi Song, Qiuzhuang Sun, Dacheng Tao Title: FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios Arxiv: http://arxiv.org/abs/2604.07413v2 Abstract: The manufacturing sector is in...

EXAONE 4.5 Technical Report 14.04.2026

🤗 Upvotes: 42 | cs. CL Authors: Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Ahra Jo, Hyunjik Jo, Yeonsik Jo, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Changhun Lee, Haeju Lee, Jinsik Lee, Kyungmin Lee, Sangha Park, Kwangrok Ryoo, Minju Seo, Sejong Yang, Heuiyeen Yeen, Hwan Chang, Stanley Jungkyu Choi, Yejin Choi, Kyubeen Han,...

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details 14.04.2026

🤗 Upvotes: 36 | cs. CV Authors: Dewei Zhou, You Li, Zongxin Yang, Yi Yang Title: RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details Arxiv: http://arxiv.org/abs/2604.06870v1 Abstract: We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal is to re...

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory 14.04.2026

🤗 Upvotes: 36 | cs. CV Authors: Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, Yahui Zhou Title: Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon M...

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability 11.04.2026

🤗 Upvotes: 156 | cs. AI Authors: Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu Title: Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Arxiv: http://arxiv.org/abs/2604.06628v1 Abstract: A prevailing narrative in LLM post-training holds that supervised finetuni...

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver 11.04.2026

🤗 Upvotes: 147 | cs. AI, cs. CL Authors: Ziyu Ma, Shidong Yang, Yuxiang Ji, Xucong Wang, Yong Wang, Yiming Hu, Tongwen Huang, Xiangxiang Chu Title: SkillClaw: Let Skills Evolve Collectively with Agentic Evolver Arxiv: http://arxiv.org/abs/2604.08377v1 Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely stat...

RAGEN-2: Reasoning Collapse in Agentic RL 10.04.2026

🤗 Upvotes: 44 | cs. LG Authors: Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, Manling Li Title: RAGEN-2: Reasoning Collapse in Agentic RL Arxiv: http://arxiv.org/abs/2604.06268v1 Abstract: RL training of multi-turn LLM agents is inherently unstable, and reasoni...

MARS: Enabling Autoregressive Models Multi-Token Generation 10.04.2026

🤗 Upvotes: 25 | cs. CL Authors: Ziqi Jin, Lei Wang, Ziwei Luo, Aixin Sun Title: MARS: Enabling Autoregressive Models Multi-Token Generation Arxiv: http://arxiv.org/abs/2604.07023v1 Abstract: Autoregressive (AR) language models generate text one token at a time, even when consecutive tokens are highly predictable given earlier context. We introduce MARS (Mask AutoRegreSsion), a lightweight fine-tu...

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents 10.04.2026

🤗 Upvotes: 22 | cs. AI, cs. CL, cs. LG Authors: Hanchen Li, Runyuan He, Qizheng Zhang, Changxiu Ji, Qiuyang Mang, Xiaokun Chen, Lakshya A Agrawal, Wei-Liang Liao, Eric Yang, Alvin Cheung, James Zou, Kunle Olukotun, Ion Stoica, Joseph E. Gonzalez Title: Combee: Scaling Prompt Learning for Self-Improving Language Model Agents Arxiv: http://arxiv.org/abs/2604.04247v1 Abstract: Recent advances in pro...

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 09.04.2026

🤗 Upvotes: 201 | cs. CV Authors: Chaoyou Fu, Haozhi Yuan, Yuhao Dong, Yi-Fan Zhang, Yunhang Shen, Xiaoxing Hu, Xueying Li, Jinsen Su, Chengwu Long, Xiaoyao Xie, Yongkang Xie, Xiawu Zheng, Xue Yang, Haoyu Cao, Yunsheng Wu, Ziwei Liu, Xing Sun, Caifeng Shan, Ran He Title: Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Arxiv: http://arxiv.org/abs/2604.05015v...

Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents 09.04.2026

🤗 Upvotes: 98 | cs. AI Authors: Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang Title: Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents Arxiv: http://arxiv.org/abs/2604.06132v1 Abstract: Large language models are increasingly deployed as autonomous agents executing multi-step workflows i...

Learning to Retrieve from Agent Trajectories 09.04.2026

🤗 Upvotes: 55 | cs. IR, cs. AI, cs. CL Authors: Yuqi Zhou, Sunhao Dai, Changle Qu, Liang Pang, Jun Xu, Ji-Rong Wen Title: Learning to Retrieve from Agent Trajectories Arxiv: http://arxiv.org/abs/2604.04949v1 Abstract: Information retrieval (IR) systems have traditionally been designed and trained for human users, with learning-to-rank methods relying heavily on large-scale human interaction logs...

ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation 09.04.2026

🤗 Upvotes: 47 | cs. LG Authors: Hui Sun, Yun-Ji Zhang, Zheng Xie, Ren-Biao Liu, Yali Du, Xin-Ye Li, Ming Li Title: ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation Arxiv: http://arxiv.org/abs/2604.03922v1 Abstract: Selecting LLM-generated code candidates using LLM-generated tests is challenging because the tests themselves may be incorrect. Existing methods either trea...

GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers 09.04.2026

🤗 Upvotes: 37 | cs. SE, cs. AI Authors: Shufan Jiang, Chios Chen, Zhiyang Chen Title: GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers Arxiv: http://arxiv.org/abs/2604.02648v1 Abstract: The autonomous discovery of bugs remains a significant challenge in modern software development. Compared to code generation, the complexity of dynamic runtime environments makes bug disco...

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning 09.04.2026

🤗 Upvotes: 33 | cs. PF, cs. SE Authors: Qisheng Su, Shiting Huang, Zhen Fang, Ziyan Chen, Zehui Chen, Feng Zhao Title: Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning Arxiv: http://arxiv.org/abs/2604.05404v1 Abstract: In real-world Tool-Integrated Reasoning (TIR) scenarios, where LLMs interleave reasoning with external tool calls, a major source of inefficiency is th...

ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement 09.04.2026

🤗 Upvotes: 32 | cs. AI Authors: Difan Jiao, Qianfeng Wen, Blair Yang, Zhenwei Tang, Ashton Anderson Title: ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement Arxiv: http://arxiv.org/abs/2604.01591v2 Abstract: We introduce ThinkTwice, a simple two-phase framework that jointly optimizes LLMs to solve reasoning problems and refine the answers, based on Group Relat...

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision 09.04.2026

🤗 Upvotes: 31 | cs. CV Authors: Hyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul Joo Title: Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision Arxiv: http://arxiv.org/abs/2604.04934v1 Abstract: We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance vide...

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU 09.04.2026

🤗 Upvotes: 26 | cs. CL, cs. DC, cs. OS Authors: Zhengqing Yuan, Hanchi Sun, Lichao Sun, Yanfang Ye Title: MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU Arxiv: http://arxiv.org/abs/2604.05091v1 Abstract: We present MegaTrain, a memory-centric system that efficiently trains 100B+ parameter large language models at full precision on a single GPU. Unlike...

Watch Before You Answer: Learning from Visually Grounded Post-Training 09.04.2026

🤗 Upvotes: 26 | cs. CV, cs. AI, cs. CL Authors: Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen Title: Watch Before You Answer: Learning from Visually Grounded Post-Training Arxiv: http://arxiv.org/abs/2604.05117v1 Abstract: It is critical for vision-language models (VLMs) to comprehensively understa...

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 08.04.2026

🤗 Upvotes: 152 | cs. CV Authors: DataFlow Team, Bohan Zeng, Daili Hua, Kaixin Zhu, Yifan Dai, Bozhou Li, Yuran Wang, Chengzhuo Tong, Yifan Yang, Mingkun Chang, Jianbin Zhao, Zhou Liu, Hao Liang, Xiaochen Ma, Ruichuan An, Junbo Niu, Zimo Meng, Tianyi Bai, Meiyi Qiang, Huanyao Zhang, Zhiyou Xiao, Tianyu Guo, Qinhan Yu, Runhao Zhao, Zhengpin Li, Xinyi Huang, Yisheng Pan, Yiwen Tang, Yang Shi, Yue Di...

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale 08.04.2026

🤗 Upvotes: 91 | cs. CV, cs. CL Authors: Bin Wang, Tianyao He, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Tao Chu, Yuan Qu, Zhenjiang Jin, Weijun Zeng, Ziyang Miao, Bangrui Xu, Junbo Niu, Mengzhang Cai, Jiantao Qiu, Qintong Zhang, Dongsheng Ma, Yuefeng Sun, Hejun Dong, Wenzheng Zhang, Jutao Xiao, Jiayong Shi, Pengyu Liao, Xiaomeng Zhao, Huaping Zhong, Liqun Wei, Jing Yu, Jie Yang, Wei Li, Shasha Wang, Qi...

LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models 08.04.2026

🤗 Upvotes: 72 | cs. LG Authors: Chanyoung Kim, Minwoo Kim, Minseok Kang, Hyunwoo Kim, Dahuin Jung Title: LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models Arxiv: http://arxiv.org/abs/2603.28301v1 Abstract: Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in dow...

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression 08.04.2026

🤗 Upvotes: 70 | cs. CL, cs. CV Authors: Weian Mao, Xi Lin, Wei Huang, Yuxin Xie, Tianfu Fu, Bohan Zhuang, Song Han, Yukang Chen Title: TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Arxiv: http://arxiv.org/abs/2604.04921v1 Abstract: Extended reasoning in large language models (LLMs) creates severe KV cache memory bottlenecks. Leading KV cache compression methods estimate...

Adam's Law: Textual Frequency Law on Large Language Models 08.04.2026

🤗 Upvotes: 46 | cs. CL Authors: Hongyuan Adam Lu, Z. L., Victor Wei, Zefan Zhang, Zhao Hong, Qiqi Xiang, Bowen Cao, Wai Lam Title: Adam's Law: Textual Frequency Law on Large Language Models Arxiv: http://arxiv.org/abs/2604.02176v2 Abstract: While textual frequency has been validated as relevant to human cognition in reading speed, its relatedness to Large Language Models (LLMs) is seldom studied....

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.