Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Emerging Properties in Unified Multimodal Pretraining 22.05.2025 22:46
🤗 Upvotes: 87 | cs. CV Authors: Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan Title: Emerging Properties in Unified Multimodal Pretraining Arxiv: http://arxiv.org/abs/2505.14683v1 Abstract: Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary syste...
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training 22.05.2025 21:11
🤗 Upvotes: 48 | cs. LG, cs. AI, cs. AR, cs. CV, cs. PF Authors: Jintao Zhang, Jia Wei, Pengle Zhang, Xiaoming Xu, Haofeng Huang, Haoxu Wang, Kai Jiang, Jun Zhu, Jianfei Chen Title: SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training Arxiv: http://arxiv.org/abs/2505.11594v1 Abstract: The efficiency of attention is important due to its quadratic time comple...
Optimizing Anytime Reasoning via Budget Relative Policy Optimization 22.05.2025 21:59
🤗 Upvotes: 30 | cs. LG, cs. AI, cs. CL Authors: Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin Title: Optimizing Anytime Reasoning via Budget Relative Policy Optimization Arxiv: http://arxiv.org/abs/2505.13438v1 Abstract: Scaling test-time compute is crucial for enhancing the reasoning capabilities of large language models (LLMs). Existing approaches typically employ reinforce...
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank 22.05.2025 20:39
🤗 Upvotes: 28 | cs. CV Authors: Tianhe Wu, Jian Zou, Jie Liang, Lei Zhang, Kede Ma Title: VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank Arxiv: http://arxiv.org/abs/2505.14460v1 Abstract: DeepSeek-R1 has demonstrated remarkable effectiveness in incentivizing reasoning and generalization capabilities of large language models (LLMs) through reinforce...
Visual Agentic Reinforcement Fine-Tuning 22.05.2025 23:31
🤗 Upvotes: 26 | cs. CV, cs. AI Authors: Ziyu Liu, Yuhang Zang, Yushan Zou, Zijian Liang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang Title: Visual Agentic Reinforcement Fine-Tuning Arxiv: http://arxiv.org/abs/2505.14246v1 Abstract: A key trend in Large Reasoning Models (e.g., OpenAI's o3) is the native agentic ability to use external tools such as web browsers for searching and w...
Neurosymbolic Diffusion Models 22.05.2025 23:44
🤗 Upvotes: 25 | cs. LG Authors: Emile van Krieken, Pasquale Minervini, Edoardo Ponti, Antonio Vergari Title: Neurosymbolic Diffusion Models Arxiv: http://arxiv.org/abs/2505.13138v1 Abstract: Neurosymbolic (NeSy) predictors combine neural perception with symbolic reasoning to solve tasks like visual reasoning. However, standard NeSy predictors assume conditional independence between the symbols th...
Chain-of-Model Learning for Language Model 21.05.2025 23:37
🤗 Upvotes: 70 | cs. CL Authors: Kaitao Song, Xiaohua Wang, Xu Tan, Huiqiang Jiang, Chengruidong Zhang, Yongliang Shen, Cen LU, Zihao Li, Zifan Song, Caihua Shan, Yansen Wang, Kan Ren, Xiaoqing Zheng, Tao Qin, Yuqing Yang, Dongsheng Li, Lili Qiu Title: Chain-of-Model Learning for Language Model Arxiv: http://arxiv.org/abs/2505.11820v1 Abstract: In this paper, we propose a novel learning paradigm,...
AdaptThink: Reasoning Models Can Learn When to Think 21.05.2025 20:31
🤗 Upvotes: 58 | cs. CL, cs. AI, cs. LG Authors: Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, Juanzi Li Title: AdaptThink: Reasoning Models Can Learn When to Think Arxiv: http://arxiv.org/abs/2505.13417v1 Abstract: Recently, large reasoning models have achieved impressive performance on various tasks by employing human-like deep thinking. However, the lengthy thinking process substantially increa...
AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning 21.05.2025 20:57
🤗 Upvotes: 46 | cs. LG, cs. AI Authors: Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, Shuangzhi Wu Title: AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning Arxiv: http://arxiv.org/abs/2505.11896v1 Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities but often face challenges with tas...
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction 21.05.2025 20:30
🤗 Upvotes: 39 | cs. LG Authors: Jeffrey Willette, Heejun Lee, Sung Ju Hwang Title: Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction Arxiv: http://arxiv.org/abs/2505.11254v1 Abstract: The attention mechanism of a transformer has a quadratic complexity, leading to high inference costs and latency for long sequences. However, attention matrices are mostly sparse, whi...
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis 21.05.2025 22:05
🤗 Upvotes: 34 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Tianbao Xie, Jiaqi Deng, Xiaochuan Li, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang, Yuhui Xu, Zekun Wang, Yiheng Xu, Junli Wang, Doyen Sahoo, Tao Yu, Caiming Xiong Title: Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis Arxiv: http://arxiv.org/abs/2505.13227v1 Abstract: Graphical user interface...
Faster Video Diffusion with Trainable Sparse Attention 21.05.2025 25:40
🤗 Upvotes: 29 | cs. CV Authors: Peiyuan Zhang, Haofeng Huang, Yongqi Chen, Will Lin, Zhengzhong Liu, Ion Stoica, Eric P. Xing, Hao Zhang Title: Faster Video Diffusion with Trainable Sparse Attention Arxiv: http://arxiv.org/abs/2505.13389v1 Abstract: Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a smal...
Thinkless: LLM Learns When to Think 21.05.2025 17:59
🤗 Upvotes: 28 | cs. CL, cs. AI Authors: Gongfan Fang, Xinyin Ma, Xinchao Wang Title: Thinkless: LLM Learns When to Think Arxiv: http://arxiv.org/abs/2505.13379v1 Abstract: Reasoning Language Models, capable of extended chain-of-thought reasoning, have demonstrated remarkable performance on tasks requiring complex logical inference. However, applying elaborate reasoning for all queries often resul...
Model Merging in Pre-training of Large Language Models 21.05.2025 23:02
🤗 Upvotes: 27 | cs. CL, cs. LG Authors: Yunshui Li, Yiyuan Ma, Shen Yan, Chaoyi Zhang, Jing Liu, Jianqiao Lu, Ziwen Xu, Mengzhao Chen, Minrui Wang, Shiyi Zhan, Jin Ma, Xunhao Lai, Yao Luo, Xingyan Bin, Hongbin Ren, Mingji Han, Wenhao Hao, Bairen Yi, LingJun Liu, Bole Ma, Xiaoying Jia, Zhou Xun, Siyuan Qiao, Liang Xiang, Yonghui Wu Title: Model Merging in Pre-training of Large Language Models Arxi...
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space 21.05.2025 24:52
🤗 Upvotes: 23 | cs. LG, cs. AI, cs. CL Authors: Hengli Li, Chenxi Li, Tong Wu, Xuekai Zhu, Yuxuan Wang, Zhaoxin Yu, Eric Hanchen Jiang, Song-Chun Zhu, Zixia Jia, Ying Nian Wu, Zilong Zheng Title: Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space Arxiv: http://arxiv.org/abs/2505.13308v1 Abstract: Reasoning ability, a core component of human intelligence, cont...
Qwen3 Technical Report 20.05.2025 21:31
🤗 Upvotes: 117 | cs. CL Authors: An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Li...
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning 20.05.2025 23:50
🤗 Upvotes: 43 | cs. AI, cs. CR Authors: Yue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang, Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi Title: GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning Arxiv: http://arxiv.org/abs/2505.11049v1 Abstract: To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model...
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly 20.05.2025 19:26
🤗 Upvotes: 42 | cs. CV, cs. CL Authors: Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman Title: MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly Arxiv: http://arxiv.org/abs/2505.10610v1 Abstract: The rapid extension of context windows in large vision-l...
Visual Planning: Let's Think Only with Images 20.05.2025 21:56
🤗 Upvotes: 33 | cs. LG, cs. AI, cs. CL, cs. CV Authors: Yi Xu, Chengzu Li, Han Zhou, Xingchen Wan, Caiqi Zhang, Anna Korhonen, Ivan Vulić Title: Visual Planning: Let's Think Only with Images Arxiv: http://arxiv.org/abs/2505.11409v1 Abstract: Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse task...
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models 17.05.2025 21:52
🤗 Upvotes: 76 | cs. CL Authors: Zhiyuan Hu, Yibo Wang, Hanze Dong, Yuhui Xu, Amrita Saha, Caiming Xiong, Bryan Hooi, Junnan Li Title: Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models Arxiv: http://arxiv.org/abs/2505.10554v1 Abstract: Large reasoning models (LRMs) already possess a latent capacity for long chain-of-thought reasoning. Prior work has shown that out...
System Prompt Optimization with Meta-Learning 17.05.2025 21:40
🤗 Upvotes: 48 | cs. CL, cs. AI, cs. LG Authors: Yumin Choi, Jinheon Baek, Sung Ju Hwang Title: System Prompt Optimization with Meta-Learning Arxiv: http://arxiv.org/abs/2505.09666v1 Abstract: Large Language Models (LLMs) have shown remarkable capabilities, with optimizing their input prompts playing a pivotal role in maximizing their performance. However, while LLM prompts consist of both the tas...
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 16.05.2025 19:27
🤗 Upvotes: 49 | cs. CV, cs. AI Authors: Jiuhai Chen, Zhiyang Xu, Xichen Pan, Yushi Hu, Can Qin, Tom Goldstein, Lifu Huang, Tianyi Zhou, Saining Xie, Silvio Savarese, Le Xue, Caiming Xiong, Ran Xu Title: BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset Arxiv: http://arxiv.org/abs/2505.09568v1 Abstract: Unifying image understanding and generation has gain...
DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception 16.05.2025 19:01
🤗 Upvotes: 36 | cs. CV Authors: Junjie Wang, Bin Chen, Yulin Li, Bin Kang, Yichi Chen, Zhuotao Tian Title: DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception Arxiv: http://arxiv.org/abs/2505.04410v1 Abstract: Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are un...
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures 16.05.2025 21:48
🤗 Upvotes: 26 | cs. DC, cs. AI, cs. AR Authors: Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Huazuo Gao, Jiashi Li, Liyue Zhang, Panpan Huang, Shangyan Zhou, Shirong Ma, Wenfeng Liang, Ying He, Yuqing Wang, Yuxuan Liu, Y. X. Wei Title: Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures Arxiv: http://arxiv.org/abs/2505.09343v1 Abstract: The rapid...
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder 15.05.2025 21:57
🤗 Upvotes: 83 | eess. AS, cs. SD Authors: Bowen Zhang, Congchao Guo, Geng Yang, Hang Yu, Haozhe Zhang, Heidi Lei, Jialong Mai, Junjie Yan, Kaiyue Yang, Mingqi Yang, Peikai Huang, Ruiyang Jin, Sitan Jiang, Weihua Cheng, Yawei Li, Yichen Xiao, Yiying Zhou, Yongmao Zhang, Yuan Lu, Yucen He Title: MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Arxiv: http://arxiv....
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.