Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model 05.02.2025 23:34
🤗 Upvotes: 25 | cs. CR, cs. AI, cs. IR Authors: Xun Liang, Simin Niu, Zhiyu Li, Sensen Zhang, Hanyu Wang, Feiyu Xiong, Jason Zhaoxin Fan, Bo Tang, Shichao Song, Mengwei Wang, Jiawei Yang Title: SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model Arxiv: http://arxiv.org/abs/2501.18636v1 Abstract: The indexing-retrieval-generation paradigm of retrieval-augmented...
Preference Leakage: A Contamination Problem in LLM-as-a-judge 05.02.2025 21:56
🤗 Upvotes: 25 | cs. LG, cs. AI, cs. CL Authors: Dawei Li, Renliang Sun, Yue Huang, Ming Zhong, Bohan Jiang, Jiawei Han, Xiangliang Zhang, Wei Wang, Huan Liu Title: Preference Leakage: A Contamination Problem in LLM-as-a-judge Arxiv: http://arxiv.org/abs/2502.01534v1 Abstract: Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annota...
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models 05.02.2025 25:02
🤗 Upvotes: 19 | cs. CV, cs. GR, cs. LG Authors: Rohit Gandikota, Zongze Wu, Richard Zhang, David Bau, Eli Shechtman, Nick Kolkin Title: SliderSpace: Decomposing the Visual Capabilities of Diffusion Models Arxiv: http://arxiv.org/abs/2502.01639v1 Abstract: We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-under...
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models 05.02.2025 24:39
🤗 Upvotes: 15 | cs. AI, cs. CV Authors: Huanqia Cai, Yijun Yang, Winston Hu Title: MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models Arxiv: http://arxiv.org/abs/2502.00698v1 Abstract: IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or do...
AIN: The Arabic INclusive Large Multimodal Model 05.02.2025 20:32
🤗 Upvotes: 12 | cs. CV, cs. AI, cs. CL, cs. HC, cs. LG Authors: Ahmed Heakl, Sara Ghaboura, Omkar Thawkar, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan Title: AIN: The Arabic INclusive Large Multimodal Model Arxiv: http://arxiv.org/abs/2502.00094v1 Abstract: Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), si...
s1: Simple test-time scaling 04.02.2025 22:36
🤗 Upvotes: 54 | cs. CL, cs. AI, cs. LG Authors: Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto Title: s1: Simple test-time scaling Arxiv: http://arxiv.org/abs/2501.19393v1 Abstract: Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to...
Reward-Guided Speculative Decoding for Efficient LLM Reasoning 04.02.2025 21:49
🤗 Upvotes: 28 | cs. CL, cs. AI Authors: Baohao Liao, Yuhui Xu, Hanze Dong, Junnan Li, Christof Monz, Silvio Savarese, Doyen Sahoo, Caiming Xiong Title: Reward-Guided Speculative Decoding for Efficient LLM Reasoning Arxiv: http://arxiv.org/abs/2501.19324v1 Abstract: We introduce Reward-Guided Speculative Decoding (RSD), a novel framework aimed at improving the efficiency of inference in large lang...
Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models 04.02.2025 21:27
🤗 Upvotes: 12 | cs. CL, cs. AI Authors: Qika Lin, Tianzhe Zhao, Kai He, Zhen Peng, Fangzhi Xu, Ling Huang, Jingying Ma, Mengling Feng Title: Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models Arxiv: http://arxiv.org/abs/2501.18119v1 Abstract: Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural...
PixelWorld: Towards Perceiving Everything as Pixels 04.02.2025 20:07
🤗 Upvotes: 10 | cs. CV, cs. CL Authors: Zhiheng Lyu, Xueguang Ma, Wenhu Chen Title: PixelWorld: Towards Perceiving Everything as Pixels Arxiv: http://arxiv.org/abs/2501.19339v1 Abstract: Existing foundation models typically process visual input as pixels and textual input as tokens, a paradigm that contrasts with human perception, where both modalities are processed in a unified manner. With the...
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning 04.02.2025 20:19
🤗 Upvotes: 8 | cs. RO, cs. AI Authors: Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto Title: DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning Arxiv: http://arxiv.org/abs/2411.04983v2 Abstract: The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remains chal...
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming 04.02.2025 20:55
🤗 Upvotes: 6 | cs. CL, cs. AI, cs. CR, cs. LG Authors: Mrinank Sharma, Meg Tong, Jesse Mu, Jerry Wei, Jorrit Kruthoff, Scott Goodfriend, Euan Ong, Alwin Peng, Raj Agarwal, Cem Anil, Amanda Askell, Nathan Bailey, Joe Benton, Emma Bluemke, Samuel R. Bowman, Eric Christiansen, Hoagy Cunningham, Andy Dau, Anjali Gopal, Rob Gilson, Logan Graham, Logan Howard, Nimit Kalra, Taesung Lee, Kevin Lin, Peter...
Scalable-Softmax Is Superior for Attention 04.02.2025 23:31
🤗 Upvotes: 6 | cs. CL, cs. AI, cs. LG Authors: Ken M. Nakanishi Title: Scalable-Softmax Is Superior for Attention Arxiv: http://arxiv.org/abs/2501.19399v1 Abstract: The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to...
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training 04.02.2025 21:51
🤗 Upvotes: 3 | cs. LG, math. OC, stat. ML Authors: Fabian Schaipp, Alexander Hägele, Adrien Taylor, Umut Simsekli, Francis Bach Title: The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Arxiv: http://arxiv.org/abs/2501.18965v1 Abstract: We show that learning-rate schedules for large model training behave surprisingly similar to a perf...
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders 04.02.2025 20:09
🤗 Upvotes: 3 | cs. LG, cs. AI Authors: Bartosz Cywiński, Kamil Deja Title: SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders Arxiv: http://arxiv.org/abs/2501.18052v2 Abstract: Diffusion models, while powerful, can inadvertently generate harmful or undesirable content, raising significant ethical and safety concerns. Recent machine unlearning approaches offer p...
GuardReasoner: Towards Reasoning-based LLM Safeguards 01.02.2025 21:00
🤗 Upvotes: 46 | cs. CR, cs. AI, cs. LG Authors: Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, Bryan Hooi Title: GuardReasoner: Towards Reasoning-based LLM Safeguards Arxiv: http://arxiv.org/abs/2501.18492v1 Abstract: As LLMs increasingly impact safety-critical applications, ensuring their safety using guardrails remains a key c...
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs 01.02.2025 23:01
🤗 Upvotes: 22 | cs. CL Authors: Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, Dong Yu Title: Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs Arxiv: http://arxiv.org/abs/2501.18585v1 Abstract: Large language models (LLMs) such as OpenAI's o1 have demonstrated remarkable...
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch 01.02.2025 23:36
🤗 Upvotes: 15 | cs. CL Authors: Arthur Douillard, Yanislav Donchev, Keith Rush, Satyen Kale, Zachary Charles, Zachary Garrett, Gabriel Teston, Dave Lacey, Ross McIlroy, Jiajun Shen, Alexandre Ramé, Arthur Szlam, Marc'Aurelio Ranzato, Paul Barham Title: Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch Arxiv: http://arxiv.org/abs/2501.18512v1 Abstract: Training of l...
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 01.02.2025 19:26
🤗 Upvotes: 15 | cs. AI, cs. CL, cs. CV, cs. LG Authors: Yuxin Zuo, Shang Qu, Yifei Li, Zhangren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, Bowen Zhou Title: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Arxiv: http://arxiv.org/abs/2501.18362v1 Abstract: We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical...
Large Language Models Think Too Fast To Explore Effectively 01.02.2025 25:52
🤗 Upvotes: 10 | cs. AI, q-bio. NC Authors: Lan Pan, Hanbo Xie, Robert C. Wilson Title: Large Language Models Think Too Fast To Explore Effectively Arxiv: http://arxiv.org/abs/2501.18009v1 Abstract: Large Language Models have emerged many intellectual capacities. While numerous benchmarks assess their intelligence, limited attention has been given to their ability to explore, an essential capacity...
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 01.02.2025 20:15
🤗 Upvotes: 10 | cs. LG, cs. CL Authors: Benjamin Feuer, Chinmay Hegde Title: WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training Arxiv: http://arxiv.org/abs/2501.18511v1 Abstract: Language model (LLM) post-training, from DPO to distillation, can refine behaviors and unlock new skills, but the open science supporting these post-training techniques is still in its infancy. On...
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding 01.02.2025 24:32
🤗 Upvotes: 10 | cs. CV, cs. AI, cs. CL, cs. LG, cs. RO Authors: Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, Yue Wang Title: PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding Arxiv: http://arxiv.org/abs/2501.16411v2 Abstract: Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to per...
o3-mini vs DeepSeek-R1: Which One is Safer? 01.02.2025 20:01
🤗 Upvotes: 6 | cs. SE, cs. AI Authors: Aitor Arrieta, Miriam Ugarte, Pablo Valle, José Antonio Parejo, Sergio Segura Title: o3-mini vs DeepSeek-R1: Which One is Safer? Arxiv: http://arxiv.org/abs/2501.18438v1 Abstract: The irruption of DeepSeek-R1 constitutes a turning point for the AI industry in general and the LLMs in particular. Its capabilities have demonstrated outstanding performance in se...
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 01.02.2025 21:14
🤗 Upvotes: 1 | cs. AI, cs. CL, cs. HC Authors: Faria Huq, Zora Zhiruo Wang, Frank F. Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P. Bigham, Graham Neubig Title: CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation Arxiv: http://arxiv.org/abs/2501.16609v1 Abstract: While much work on web agents emphasizes the promise of autonomously performing tasks on behalf of users, in rea...
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 31.01.2025 22:30
🤗 Upvotes: 28 | cs. CL Authors: Yubo Wang, Xiang Yue, Wenhu Chen Title: Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate Arxiv: http://arxiv.org/abs/2501.17703v2 Abstract: Supervised Fine-Tuning (SFT) is commonly used to train language models to imitate annotated responses for given instructions. In this paper, we challenge this paradigm and propose Critique F...
Atla Selene Mini: A General Purpose Evaluation Model 31.01.2025 25:28
🤗 Upvotes: 24 | cs. CL, cs. AI Authors: Andrei Alexandru, Antonia Calvi, Henry Broomfield, Jackson Golden, Kyle Dai, Mathias Leys, Maurice Burger, Max Bartolo, Roman Engeler, Sashank Pisupati, Toby Drane, Young Sun Park Title: Atla Selene Mini: A General Purpose Evaluation Model Arxiv: http://arxiv.org/abs/2501.17195v1 Abstract: We introduce Atla Selene Mini, a state-of-the-art small language mod...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.