Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start 30.05.2025 21:34
🤗 Upvotes: 31 | cs. CL, cs. AI, cs. CV, cs. LG Authors: Lai Wei, Yuting Li, Kaipeng Zheng, Chen Wang, Yue Wang, Linghe Kong, Lichao Sun, Weiran Huang Title: Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Arxiv: http://arxiv.org/abs/2505.22334v1 Abstract: Recent advancements in large language models (LLMs) have demonstrated impressive chain-of-thought reasoning capabilit...
Fostering Video Reasoning via Next-Event Prediction 30.05.2025 24:54
🤗 Upvotes: 27 | cs. CV, cs. AI, cs. CL Authors: Haonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du, Kenji Kawaguchi, Ye Wang, Tianyu Pang Title: Fostering Video Reasoning via Next-Event Prediction Arxiv: http://arxiv.org/abs/2505.22457v1 Abstract: Next-token prediction serves as the foundational learning task enabling reasoning in LLMs. But what should the learning task be when aiming to equip MLLMs...
RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination 30.05.2025 23:21
🤗 Upvotes: 26 | cs. GR, cs. CV, cs. LG Authors: Chong Zeng, Yue Dong, Pieter Peers, Hongzhi Wu, Xin Tong Title: RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination Arxiv: http://arxiv.org/abs/2505.21925v1 Abstract: We present RenderFormer, a neural rendering pipeline that directly renders an image from a triangle-based representation of a scene with full g...
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows 29.05.2025 22:14
🤗 Upvotes: 85 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, Zhiyong Wu Title: ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Sc...
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs 29.05.2025 21:06
🤗 Upvotes: 73 | cs. AI, cs. CV Authors: Jiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu, Renrui Zhang, Kaituo Feng, Chaoyou Fu, Tao Chen, Lei Bai, Bo Zhang, Xiangyu Yue Title: MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Arxiv: http://arxiv.org/abs/2505.21327v1 Abstract: Logical reasoning is a fundamental aspect of human intelligence and an essential capability for...
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers 29.05.2025 17:55
🤗 Upvotes: 73 | cs. CV, cs. AI, cs. CL, cs. MA Authors: Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, Philip Torr Title: Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers Arxiv: http://arxiv.org/abs/2505.21497v1 Abstract: Academic poster generation is a crucial yet challenging task in scientific communication, requiring the compression of long-context interleaved docu...
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data 29.05.2025 24:24
🤗 Upvotes: 57 | cs. CV Authors: Yiren Song, Cheng Liu, Mike Zheng Shou Title: OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data Arxiv: http://arxiv.org/abs/2505.18445v1 Abstract: Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, compositio...
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation 29.05.2025 19:53
🤗 Upvotes: 49 | cs. CV, cs. AI Authors: Shenghai Yuan, Xianyi He, Yufan Deng, Yang Ye, Jinfa Huang, Bin Lin, Jiebo Luo, Li Yuan Title: OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Arxiv: http://arxiv.org/abs/2505.20292v3 Abstract: Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanc...
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond 29.05.2025 21:52
🤗 Upvotes: 43 | cs. AI, cs. CL Authors: Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Han Ding, Yongyi Hu, Chi Zhang, Yiqi Shi, Shitong Weng, Aili Chen, Shiqi Chen, Yunan Huang, Mozhi Zhang, Pengyu Zhao, Junjie Yan, Junxian He Title: SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond Arxiv: http://arxiv.org/abs/2505.19641v3 Abstract: Recent advances such...
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning 29.05.2025 18:54
🤗 Upvotes: 41 | cs. CL, cs. AI Authors: Michael Hassid, Gabriel Synnaeve, Yossi Adi, Roy Schwartz Title: Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning Arxiv: http://arxiv.org/abs/2505.17813v1 Abstract: Reasoning large language models (LLMs) heavily rely on scaling test-time compute to perform complex reasoning tasks by generating extensive "thinking" chains. Wh...
Exploring the Latent Capacity of LLMs for One-Step Text Generation 29.05.2025 20:35
🤗 Upvotes: 40 | cs. CL, cs. AI, cs. LG Authors: Gleb Mezentsev, Ivan Oseledets Title: Exploring the Latent Capacity of LLMs for One-Step Text Generation Arxiv: http://arxiv.org/abs/2505.21189v1 Abstract: A recent study showed that large language models (LLMs) can reconstruct surprisingly long texts - up to thousands of tokens - via autoregressive generation from just one specially trained input e...
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence 29.05.2025 22:23
🤗 Upvotes: 39 | cs. CL, cs. AI Authors: Amirhosein Ghasemabadi, Keith G. Mills, Baochun Li, Di Niu Title: Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Arxiv: http://arxiv.org/abs/2505.20325v1 Abstract: Test-Time Scaling (TTS) methods for enhancing Large Language Model (LLM) reasoning often incur substantial computational costs, primarily due to extensive relianc...
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization 29.05.2025 20:58
🤗 Upvotes: 35 | cs. CL, cs. CV Authors: Yunxin Li, Xinyu Chen, Zitao Li, Zhenyu Liu, Longyue Wang, Wenhan Luo, Baotian Hu, Min Zhang Title: VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization Arxiv: http://arxiv.org/abs/2505.19000v1 Abstract: Applying Reinforcement Learning (RL) to Video Large Language Models (Video-LLMs) shows significant promise fo...
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model 28.05.2025 20:46
🤗 Upvotes: 178 | cs. CL, cs. AI Authors: Khalil Hennara, Muhammad Hreden, Mohamed Motaism Hamed, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan Title: Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model Arxiv: http://arxiv.org/abs/2505.17894v1 Abstract: We introduce Mutarjim, a compact yet powerful language model for bidirectional Arabic-English translation. W...
Shifting AI Efficiency From Model-Centric to Data-Centric Compression 28.05.2025 22:14
🤗 Upvotes: 124 | cs. CL, cs. AI, cs. CV Authors: Xuyang Liu, Zichen Wen, Shaobo Wang, Junjie Chen, Zhishan Tao, Yubo Wang, Xiangqi Jin, Chang Zou, Yiyu Wang, Chenfei Liao, Xu Zheng, Honggang Chen, Weijia Li, Xuming Hu, Conghui He, Linfeng Zhang Title: Shifting AI Efficiency From Model-Centric to Data-Centric Compression Arxiv: http://arxiv.org/abs/2505.19147v1 Abstract: The rapid advancement of l...
Alchemist: Turning Public Text-to-Image Data into Generative Gold 28.05.2025 19:19
🤗 Upvotes: 58 | cs. CV Authors: Valerii Startsev, Alexander Ustyuzhanin, Alexey Kirillov, Dmitry Baranchuk, Sergey Kastryulin Title: Alchemist: Turning Public Text-to-Image Data into Generative Gold Arxiv: http://arxiv.org/abs/2505.19297v1 Abstract: Pre-training equips text-to-image (T2I) models with broad world knowledge, but this alone is often insufficient to achieve high aesthetic quality and...
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs 28.05.2025 23:32
🤗 Upvotes: 56 | cs. AI, cs. CE, cs. CL Authors: Guilong Lu, Xuntao Guo, Rongjunchen Zhang, Wenqiao Zhu, Ji Liu Title: BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs Arxiv: http://arxiv.org/abs/2505.19457v1 Abstract: Large language models excel in general tasks, yet assessing their reliability in logic-heavy, precision-critical domains like finance, law, and heal...
PATS: Process-Level Adaptive Thinking Mode Switching 28.05.2025 21:12
🤗 Upvotes: 44 | cs. CL Authors: Yi Wang, Junxiao Liu, Shimao Zhang, Jiajun Chen, Shujian Huang Title: PATS: Process-Level Adaptive Thinking Mode Switching Arxiv: http://arxiv.org/abs/2505.19250v1 Abstract: Current large-language models (LLMs) typically adopt a fixed reasoning strategy, either simple or complex, for all questions, regardless of their difficulty. This neglect of variation in task a...
Embodied Agents Meet Personalization: Exploring Memory Utilization for Personalized Assistance 28.05.2025 20:57
🤗 Upvotes: 42 | cs. CL Authors: Taeyoon Kwon, Dongwook Choi, Sunghwan Kim, Hyojun Kim, Seungjun Moon, Beong-woo Kwak, Kuan-Hao Huang, Jinyoung Yeo Title: Embodied Agents Meet Personalization: Exploring Memory Utilization for Personalized Assistance Arxiv: http://arxiv.org/abs/2505.16348v1 Abstract: Embodied agents empowered by large language models (LLMs) have shown strong performance in househol...
ARM: Adaptive Reasoning Model 28.05.2025 22:44
🤗 Upvotes: 40 | cs. CL Authors: Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, Yanghua Xiao Title: ARM: Adaptive Reasoning Model Arxiv: http://arxiv.org/abs/2505.20258v1 Abstract: While large reasoning models demonstrate strong performance on complex tasks, they lack the ability to adjust reasoning token usage based on task difficulty. This often leads to the "overthinking" problem...
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles 28.05.2025 21:22
🤗 Upvotes: 33 | cs. CL, cs. AI Authors: Jiangjie Chen, Qianyu He, Siyu Yuan, Aili Chen, Zhicheng Cai, Weinan Dai, Hongli Yu, Qiying Yu, Xuefeng Li, Jiaze Chen, Hao Zhou, Mingxuan Wang Title: Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles Arxiv: http://arxiv.org/abs/2505.19914v1 Abstract: Large Language Models (LLMs), such as OpenAI's o1 and DeepSeek...
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective 28.05.2025 21:22
🤗 Upvotes: 33 | cs. CL, cs. AI Authors: Junnan Liu, Hongwei Liu, Linchen Xiao, Shudong Liu, Taolin Zhang, Zihan Ma, Songyang Zhang, Kai Chen Title: Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective Arxiv: http://arxiv.org/abs/2505.19815v1 Abstract: We propose a novel framework for comprehending the reasoning capabilities of large language models (LLMs) through the perspectiv...
B-score: Detecting biases in large language models using response history 28.05.2025 23:16
🤗 Upvotes: 25 | cs. LG, cs. CL Authors: An Vo, Mohammad Reza Taesiri, Daeyoung Kim, Anh Totti Nguyen Title: B-score: Detecting biases in large language models using response history Arxiv: http://arxiv.org/abs/2505.18545v1 Abstract: Large language models (LLMs) often exhibit strong biases, e.g, against women or in favor of the number 7. We investigate whether LLMs would be able to output less bia...
TabSTAR: A Foundation Tabular Model With Semantically Target-Aware Representations 27.05.2025 20:40
🤗 Upvotes: 95 | cs. LG, cs. CL Authors: Alan Arazi, Eilam Shapira, Roi Reichart Title: TabSTAR: A Foundation Tabular Model With Semantically Target-Aware Representations Arxiv: http://arxiv.org/abs/2505.18125v1 Abstract: While deep learning has achieved remarkable success across many domains, it has historically underperformed on tabular learning tasks, which remain dominated by gradient boosting...
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning 27.05.2025 24:07
🤗 Upvotes: 60 | cs. CL Authors: Fanqi Wan, Weizhou Shen, Shengyi Liao, Yingcheng Shi, Chenliang Li, Ziyi Yang, Ji Zhang, Fei Huang, Jingren Zhou, Ming Yan Title: QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.17667v1 Abstract: Recent large reasoning models (LRMs) have demonstrated strong reasoning capabilities through reinforc...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.