Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Learning to Reason under Off-Policy Guidance 23.04.2025

🤗 Upvotes: 59 | cs. LG, cs. AI, cs. CL Authors: Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang, Ganqu Cui, Xiaoye Qu, Yu Cheng, Yue Zhang Title: Learning to Reason under Off-Policy Guidance Arxiv: http://arxiv.org/abs/2504.14945v2 Abstract: Recent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcem...

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models 23.04.2025

🤗 Upvotes: 50 | cs. CV Authors: Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tuomas Rintamaki, Tyler Poon, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu Title: Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models Arxiv: h...

FlowReasoner: Reinforcing Query-Level Meta-Agents 23.04.2025

🤗 Upvotes: 36 | cs. AI Authors: Hongcheng Gao, Yue Liu, Yufei He, Longxu Dou, Chao Du, Zhijie Deng, Bryan Hooi, Min Lin, Tianyu Pang Title: FlowReasoner: Reinforcing Query-Level Meta-Agents Arxiv: http://arxiv.org/abs/2504.15257v1 Abstract: This paper proposes a query-level meta-agent named FlowReasoner to automate the design of query-level multi-agent systems, i.e., one system per user query. Ou...

ToolRL: Reward is All Tool Learning Needs 23.04.2025

🤗 Upvotes: 33 | cs. LG, cs. AI, cs. CL Authors: Cheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan Tur, Heng Ji Title: ToolRL: Reward is All Tool Learning Needs Arxiv: http://arxiv.org/abs/2504.13958v1 Abstract: Current Large Language Models (LLMs) often undergo supervised fine-tuning (SFT) to acquire tool use capabilities. However, SFT struggles to generalize...

X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents 23.04.2025

🤗 Upvotes: 25 | cs. CR, cs. AI, cs. CL, cs. LG, cs. MA Authors: Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, Saadia Gabriel Title: X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents Arxiv: http://arxiv.org/abs/2504.13203v1 Abstract: Multi-turn interactions with language models (LMs) pose c...

StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians 23.04.2025

🤗 Upvotes: 21 | cs. CV Authors: Cailin Zhuang, Yaoqi Hu, Xuanyang Zhang, Wei Cheng, Jiacheng Bao, Shengqi Liu, Yiying Yang, Xianfang Zeng, Gang Yu, Ming Li Title: StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians Arxiv: http://arxiv.org/abs/2504.15281v1 Abstract: 3D Gaussian Splatting (3DGS) excels in photorealistic scene reconstruction but struggles with styliz...

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? 22.04.2025

🤗 Upvotes: 64 | cs. AI, cs. CL, cs. CV Authors: Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang Title: Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? Arxiv: http://arxiv.org/abs/2504.13837v1 Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated notable success in enhancin...

MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space 22.04.2025

🤗 Upvotes: 31 | cs. CL, cs. AI Authors: Yicheng Chen, Yining Li, Kai Hu, Zerun Ma, Haochen Ye, Kai Chen Title: MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space Arxiv: http://arxiv.org/abs/2504.13835v1 Abstract: Data quality and diversity are key to the construction of effective instruction-tuning datasets. % With the increasing availability of...

NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes 22.04.2025

🤗 Upvotes: 30 | cs. AI Authors: Tianyang Xu, Haojie Zheng, Chengze Li, Haoxiang Chen, Yixin Liu, Ruoxi Chen, Lichao Sun Title: NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes Arxiv: http://arxiv.org/abs/2504.11544v1 Abstract: Retrieval-augmented generation (RAG) empowers large language models to access external and private corpus, enabling factually consistent responses in specific...

CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training 19.04.2025

🤗 Upvotes: 69 | cs. CL Authors: Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, Mostofa Patwary, Yingyan, Lin, Jan Kautz, Pavlo Molchanov Title: CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training Arxiv: http://arxiv.org/abs/2504.13161v1 Abstract: Pre-training datasets are typically col...

Antidistillation Sampling 19.04.2025

🤗 Upvotes: 52 | cs. AI, cs. CL Authors: Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter Title: Antidistillation Sampling Arxiv: http://arxiv.org/abs/2504.13146v1 Abstract: Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability,...

Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling 19.04.2025

🤗 Upvotes: 28 | cs. CV Authors: Tsung-Han Wu, Heekyung Lee, Jiaxin Ge, Joseph E. Gonzalez, Trevor Darrell, David M. Chan Title: Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling Arxiv: http://arxiv.org/abs/2504.13169v1 Abstract: Vision-Language Models (VLMs) excel at visual understanding but often suffer from visual hallucinations, where they gen...

Packing Input Frame Context in Next-Frame Prediction Models for Video Generation 19.04.2025

🤗 Upvotes: 24 | cs. CV Authors: Lvmin Zhang, Maneesh Agrawala Title: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation Arxiv: http://arxiv.org/abs/2504.12626v1 Abstract: We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation. The FramePack compresses input frames to make the transformer co...

WORLDMEM: Long-term Consistent World Simulation with Memory 19.04.2025

🤗 Upvotes: 23 | cs. CV Authors: Zeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang, Shuai Yang, Yanhong Zeng, Xingang Pan Title: WORLDMEM: Long-term Consistent World Simulation with Memory Arxiv: http://arxiv.org/abs/2504.12369v1 Abstract: World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited te...

A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis 19.04.2025

🤗 Upvotes: 23 | cs. CL, cs. AI, cs. LG Authors: Xin Gao, Qizhi Pei, Zinan Tang, Yu Li, Honglin Lin, Jiang Wu, Conghui He, Lijun Wu Title: A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis Arxiv: http://arxiv.org/abs/2504.12322v1 Abstract: While data synthesis and distillation are promising strategies to enhance small language models, current approaches heavily...

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness 18.04.2025

🤗 Upvotes: 35 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Yijun Liang, Ming Li, Chenrui Fan, Ziyue Li, Dang Nguyen, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, Tianyi Zhou Title: ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness Arxiv: http://arxiv.org/abs/2504.10514v1 Abstract: Color plays an important...

BitNet b1.58 2B4T Technical Report 18.04.2025

🤗 Upvotes: 35 | cs. CL, cs. LG Authors: Shuming Ma, Hongyu Wang, Shaohan Huang, Xingxing Zhang, Ying Hu, Ting Song, Yan Xia, Furu Wei Title: BitNet b1.58 2B4T Technical Report Arxiv: http://arxiv.org/abs/2504.12285v1 Abstract: We introduce BitNet b1.58 2B4T, the first open-source, native 1-bit Large Language Model (LLM) at the 2-billion parameter scale. Trained on a corpus of 4 trillion tokens, t...

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs 18.04.2025

🤗 Upvotes: 27 | cs. CL, cs. AI Authors: Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang, Yujia Qin, Baoquan Zhong, Chengquan Jiang, Jinxin Chi, Wanjun Zhong Title: ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Arxiv: http://arxiv.org/abs/2504.11536v2 Abstract: While reasoning models (e.g., DeepSeek R1) trained with reinforcement learning (RL), excel in textual reasoning, they str...

xVerify: Efficient Answer Verifier for Reasoning Model Evaluations 17.04.2025

🤗 Upvotes: 63 | cs. CL Authors: Ding Chen, Qingchen Yu, Pengyuan Wang, Wentao Zhang, Bo Tang, Feiyu Xiong, Xinchi Li, Minchuan Yang, Zhiyu Li Title: xVerify: Efficient Answer Verifier for Reasoning Model Evaluations Arxiv: http://arxiv.org/abs/2504.10481v1 Abstract: With the release of the o1 model by OpenAI, reasoning models adopting slow thinking strategies have gradually emerged. As the respon...

Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning 17.04.2025

🤗 Upvotes: 41 | cs. CL, cs. AI, cs. LG Authors: Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, Zhiyong Wu Title: Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning Arxiv: http://arxiv.org/abs/2504.08672v1 Abstract: Advancing LLM reasoning skills has captivated wide interest. However, current post-training te...

How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients 17.04.2025

🤗 Upvotes: 30 | cs. LG, cs. AI, cs. CL Authors: Ming Li, Yanhong Li, Ziyue Li, Tianyi Zhou Title: How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients Arxiv: http://arxiv.org/abs/2504.10766v1 Abstract: As the post-training of large language models (LLMs) advances from instruction-following to complex reasoning tasks, understanding how diffe...

Heimdall: test-time scaling on the generative verification 17.04.2025

🤗 Upvotes: 28 | cs. AI, I.2.7 Authors: Wenlei Shi, Xing Jin Title: Heimdall: test-time scaling on the generative verification Arxiv: http://arxiv.org/abs/2504.10337v2 Abstract: An AI system can create and maintain knowledge only to the extent that it can verify that knowledge itself. Recent work on long Chain-of-Thought reasoning has demonstrated great potential of LLMs on solving competitive pro...

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 17.04.2025

🤗 Upvotes: 23 | cs. CV Authors: Tao Zhang, Xiangtai Li, Zilong Huang, Yanwei Li, Weixian Lei, Xueqing Deng, Shihao Chen, Shunping Ji, Jiashi Feng Title: Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding Arxiv: http://arxiv.org/abs/2504.10465v1 Abstract: Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However,...

TextArena 17.04.2025

🤗 Upvotes: 21 | cs. CL, cs. AI, cs. LG, cs. MA Authors: Leon Guertler, Bobby Cheng, Simon Yu, Bo Liu, Leshem Choshen, Cheston Tan Title: TextArena Arxiv: http://arxiv.org/abs/2504.11442v1 Abstract: TextArena is an open-source collection of competitive text-based games for training and evaluation of agentic behavior in Large Language Models (LLMs). It spans 57+ unique environments (including singl...

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 16.04.2025

🤗 Upvotes: 172 | cs. CV Authors: Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Yuchen Duan, Hao Tian, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Dengnian Chen, Songze Li, Yinan He, Tan Jiang, Jiapeng Luo, Yi Wang, Conghui He, Botian Shi, Xingcheng Zhang, Wenqi Shao, Junju...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.