Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs 05.03.2025 25:44
🤗 Upvotes: 42 | cs. CL, cs. AI, cs. LG Authors: Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, Dong Chen, Dongdong Chen, Junkun Chen, Weizhu Chen, Yen-Chun Chen, Yi-ling Chen, Qi Dai, Xiyang Dai, Ruchao Fan, Mei Gao, Min Gao, Amit Garg, Abhishek Goswami, Junheng Hao, Amr Hendy, Yuxuan Hu, X...
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models 05.03.2025 19:04
🤗 Upvotes: 30 | cs. CV Authors: Jay Zhangjie Wu, Yuxuan Zhang, Haithem Turki, Xuanchi Ren, Jun Gao, Mike Zheng Shou, Sanja Fidler, Zan Gojcic, Huan Ling Title: Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models Arxiv: http://arxiv.org/abs/2503.01774v1 Abstract: Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis tas...
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking 04.03.2025 22:49
🤗 Upvotes: 27 | cs. AI Authors: Zhuoqun Li, Haiyang Yu, Xuanang Chen, Hongyu Lin, Yaojie Lu, Fei Huang, Xianpei Han, Yongbin Li, Le Sun Title: DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking Arxiv: http://arxiv.org/abs/2502.20730v1 Abstract: Designing solutions for complex engineering challenges is crucial in human production activities....
Chain of Draft: Thinking Faster by Writing Less 04.03.2025 22:37
🤗 Upvotes: 27 | cs. CL, I.2.7 Authors: Silei Xu, Wenhao Xie, Lingxiao Zhao, Pengcheng He Title: Chain of Draft: Thinking Faster by Writing Less Arxiv: http://arxiv.org/abs/2502.18600v1 Abstract: Large Language Models (LLMs) have demonstrated remarkable performance in solving complex reasoning tasks through mechanisms like Chain-of-Thought (CoT) prompting, which emphasizes verbose, step-by-step re...
Multi-Turn Code Generation Through Single-Step Rewards 04.03.2025 25:33
🤗 Upvotes: 21 | cs. LG, cs. AI, cs. CL Authors: Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M Rush, Wenting Zhao, Sanjiban Choudhury Title: Multi-Turn Code Generation Through Single-Step Rewards Arxiv: http://arxiv.org/abs/2502.20380v1 Abstract: We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedb...
Self-rewarding correction for mathematical reasoning 01.03.2025 24:30
🤗 Upvotes: 51 | cs. AI, cs. LG Authors: Wei Xiong, Hanning Zhang, Chenlu Ye, Lichang Chen, Nan Jiang, Tong Zhang Title: Self-rewarding correction for mathematical reasoning Arxiv: http://arxiv.org/abs/2502.19613v1 Abstract: We study self-rewarding reasoning large language models (LLMs), which can simultaneously generate step-by-step reasoning and evaluate the correctness of their outputs during t...
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning 01.03.2025 23:19
🤗 Upvotes: 44 | cs. CV, cs. AI Authors: Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, Daniel Rueckert Title: MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning Arxiv: http://arxiv.org/abs/2502.19634v1 Abstract: Reasoning is a critical frontier for advancing medical image analysis,...
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts 01.03.2025 22:25
🤗 Upvotes: 33 | cs. LG Authors: Zhongyang Li, Ziyue Li, Tianyi Zhou Title: R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts Arxiv: http://arxiv.org/abs/2502.20395v1 Abstract: In large multimodal models (LMMs), the perception of non-language modalities (e.g., visual representations) is usually not on par with the large language models (LLMs)' powerful reasoning capabilities, deterr...
LongRoPE2: Near-Lossless LLM Context Window Scaling 01.03.2025 23:05
🤗 Upvotes: 21 | cs. CL Authors: Ning Shang, Li Lyna Zhang, Siyuan Wang, Gaokai Zhang, Gilsinia Lopez, Fan Yang, Weizhu Chen, Mao Yang Title: LongRoPE2: Near-Lossless LLM Context Window Scaling Arxiv: http://arxiv.org/abs/2502.20082v1 Abstract: LongRoPE2 is a novel approach that extends the effective context window of pre-trained large language models (LLMs) to the target length, while preserving...
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving 01.03.2025 26:27
🤗 Upvotes: 19 | cs. CL Authors: Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong Title: FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving Arxiv: http://arxiv.org/abs/2502.20238v1 Abstract: Many challenging reasoning tasks require not just rapid, intuitive responses, but a more deliberate,...
CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale 01.03.2025 21:52
🤗 Upvotes: 15 | cs. CL, cs. AI, cs. SE Authors: Chenlong Wang, Zhaoyang Chu, Zhengxiang Cheng, Xuyi Yang, Kaiyue Qiu, Yao Wan, Zhou Zhao, Xuanhua Shi, Dongping Chen Title: CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale Arxiv: http://arxiv.org/abs/2502.16645v1 Abstract: Large Language Models (LLMs) have exhibited exceptional performance in software engineering y...
UniTok: A Unified Tokenizer for Visual Generation and Understanding 01.03.2025 24:43
🤗 Upvotes: 15 | cs. CV, cs. AI Authors: Chuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang, Xin Yu, Zehuan Yuan, Bingyue Peng, Xiaojuan Qi Title: UniTok: A Unified Tokenizer for Visual Generation and Understanding Arxiv: http://arxiv.org/abs/2502.20321v1 Abstract: The representation disparity between visual generation and understanding imposes a critical gap in integrating these capabilities into a sin...
NeoBERT: A Next-Generation BERT 01.03.2025 23:40
🤗 Upvotes: 11 | cs. CL, cs. AI Authors: Lola Le Breton, Quentin Fournier, Mariam El Mezouar, Sarath Chandar Title: NeoBERT: A Next-Generation BERT Arxiv: http://arxiv.org/abs/2502.19587v1 Abstract: Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models such as LLaMA and DeepS...
Lean and Mean: Decoupled Value Policy Optimization with Global Value Guidance 01.03.2025 21:30
🤗 Upvotes: 9 | cs. LG, cs. AI Authors: Chenghua Huang, Lu Wang, Fangkai Yang, Pu Zhao, Zhixu Li, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, Qi Zhang Title: Lean and Mean: Decoupled Value Policy Optimization with Global Value Guidance Arxiv: http://arxiv.org/abs/2502.16944v1 Abstract: Proximal Policy Optimization (PPO)-based Reinforcement Learning from Human Feedback (RLHF) is essential for ali...
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think 01.03.2025 22:16
🤗 Upvotes: 9 | cs. CV, cs. CL Authors: Liang Chen, Shuai Bai, Wenhao Chai, Weichu Xie, Haozhe Zhao, Leon Vinci, Junyang Lin, Baobao Chang Title: Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think Arxiv: http://arxiv.org/abs/2502.20172v1 Abstract: The field of advanced text-to-image generation is witnessing the emergence of unified fra...
GHOST 2.0: generative high-fidelity one shot transfer of heads 28.02.2025 18:41
🤗 Upvotes: 49 | cs. CV Authors: Alexander Groshev, Anastasiia Iashchenko, Pavel Paramonov, Denis Dimitrov, Andrey Kuznetsov Title: GHOST 2.0: generative high-fidelity one shot transfer of heads Arxiv: http://arxiv.org/abs/2502.18417v3 Abstract: While the task of face swapping has recently gained attention in the research community, a related problem of head swapping remains largely unexplored. In...
Kanana: Compute-efficient Bilingual Language Models 28.02.2025 22:05
🤗 Upvotes: 47 | cs. CL, cs. LG Authors: Kanana LLM Team, Yunju Bak, Hojin Lee, Minho Ryu, Jiyeon Ham, Seungjae Jung, Daniel Wontae Nam, Taegyeong Eo, Donghun Lee, Doohae Jung, Boseop Kim, Nayeon Kim, Jaesun Park, Hyunho Kim, Hyunwoong Ko, Changmin Lee, Kyoung-Woon On, Seulye Baeg, Junrae Cho, Sunghee Jung, Jieun Kang, EungGyun Kim, Eunhwa Kim, Byeongil Ko, Daniel Lee, Minchul Lee, Miok Lee, Shinb...
TheoremExplainAgent: Towards Multimodal Explanations for LLM Theorem Understanding 28.02.2025 22:22
🤗 Upvotes: 32 | cs. AI, cs. CL, cs. CV, cs. MM Authors: Max Ku, Thomas Chong, Jonathan Leung, Krish Shah, Alvin Yu, Wenhu Chen Title: TheoremExplainAgent: Towards Multimodal Explanations for LLM Theorem Understanding Arxiv: http://arxiv.org/abs/2502.19400v1 Abstract: Understanding domain-specific theorems often requires more than just text-based reasoning; effective communication through structur...
Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance 28.02.2025 24:57
🤗 Upvotes: 27 | cs. CL Authors: Xueqing Peng, Triantafillos Papadopoulos, Efstathia Soufleri, Polydoros Giannouris, Ruoyu Xiang, Yan Wang, Lingfei Qian, Jimin Huang, Qianqian Xie, Sophia Ananiadou Title: Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance Arxiv: http://arxiv.org/abs/2502.18772v1 Abstract: Despite Greece's pivotal role in the global economy, large language mod...
Language Models' Factuality Depends on the Language of Inquiry 28.02.2025 22:25
🤗 Upvotes: 19 | cs. CL, cs. AI Authors: Tushar Aggarwal, Kumar Tanmay, Ayush Agrawal, Kumar Ayush, Hamid Palangi, Paul Pu Liang Title: Language Models' Factuality Depends on the Language of Inquiry Arxiv: http://arxiv.org/abs/2502.17955v1 Abstract: Multilingual language models (LMs) are expected to recall factual knowledge consistently across languages, yet they often fail to transfer knowledge b...
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? 28.02.2025 24:04
🤗 Upvotes: 16 | cs. CL Authors: Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, Bo Zheng Title: Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? Arxiv: http://arxiv.org/abs/2502.19361v2 Abstract: Recently, o1-like models have drawn significant attention, where these models produce the l...
Towards an AI co-scientist 28.02.2025 25:03
🤗 Upvotes: 15 | cs. AI, cs. CL, cs. HC, cs. LG, physics.soc-ph, q-bio. OT Authors: Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavi...
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems 28.02.2025 18:21
🤗 Upvotes: 15 | cs. CL, cs. AI Authors: Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Bin Xu, Lei Hou, Juanzi Li Title: Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems Arxiv: http://arxiv.org/abs/2502.19328v1 Abstract: Reward models (RMs) are crucial for the training and inference-time scaling up of large language models (LLMs...
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation 28.02.2025 22:49
🤗 Upvotes: 13 | cs. LG, cs. SE Authors: Shiven Sinha, Shashwat Goel, Ponnurangam Kumaraguru, Jonas Geiping, Matthias Bethge, Ameya Prabhu Title: Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation Arxiv: http://arxiv.org/abs/2502.19414v1 Abstract: There is growing excitement about the potential of Language Models (LMs) to accelerate scientific discovery. Fal...
Rank1: Test-Time Compute for Reranking in Information Retrieval 28.02.2025 19:51
🤗 Upvotes: 11 | cs. IR, cs. CL, cs. LG Authors: Orion Weller, Kathryn Ricci, Eugene Yang, Andrew Yates, Dawn Lawrie, Benjamin Van Durme Title: Rank1: Test-Time Compute for Reranking in Information Retrieval Arxiv: http://arxiv.org/abs/2502.18418v1 Abstract: We introduce Rank1, the first reranking model trained to take advantage of test-time compute. Rank1 demonstrates the applicability within ret...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.