Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation 15.02.2025

🤗 Upvotes: 28 | cs. LG, cs. AI, cs. CV Authors: Hoigi Seo, Wongi Jeong, Jae-sun Seo, Se Young Chun Title: Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation Arxiv: http://arxiv.org/abs/2502.08690v1 Abstract: Large-scale text encoders in text-to-image (T2I) diffusion models have demonstrated exceptional performance in generating high-quality images from textual...

SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models 15.02.2025

🤗 Upvotes: 22 | cs. CL, cs. AI, cs. LG Authors: Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James Glass, Shang-Wen Li, Wen-tau Yih Title: SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models Arxiv: http://arxiv.org/abs/2502.09604v1 Abstract: We introduce SelfCite, a novel self-supervised approach that aligns LLM...

Can this Model Also Recognize Dogs? Zero-Shot Model Search from Weights 15.02.2025

🤗 Upvotes: 21 | cs. LG, cs. CV Authors: Jonathan Kahana, Or Nathan, Eliahu Horwitz, Yedid Hoshen Title: Can this Model Also Recognize Dogs? Zero-Shot Model Search from Weights Arxiv: http://arxiv.org/abs/2502.09619v1 Abstract: With the increasing numbers of publicly available models, there are probably pretrained, online models for most tasks users require. However, current model search methods a...

An Open Recipe: Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging 15.02.2025

🤗 Upvotes: 21 | cs. CL, cs. AI Authors: Kunat Pipatanakul, Pittawat Taveekitworachai, Potsawee Manakul, Kasima Tharnpipitchai Title: An Open Recipe: Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging Arxiv: http://arxiv.org/abs/2502.09056v1 Abstract: This paper investigates data selection and model merging methodologies aimed at incorporating advanced reasoning capa...

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents 15.02.2025

🤗 Upvotes: 20 | cs. AI, cs. CL, cs. CV Authors: Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, Heng Ji, Huan Zhang, Tong Zhang Title: EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents Arxiv: http://arxiv.org/abs/2502.09560v1 Abstract: Leveraging...

Exploring the Potential of Encoder-free Architectures in 3D LMMs 15.02.2025

🤗 Upvotes: 17 | cs. CV, cs. AI, cs. CL Authors: Yiwen Tang, Zoey Guo, Zhuhao Wang, Ray Zhang, Qizhi Chen, Junli Liu, Delin Qu, Zhigang Wang, Dong Wang, Xuelong Li, Bin Zhao Title: Exploring the Potential of Encoder-free Architectures in 3D LMMs Arxiv: http://arxiv.org/abs/2502.09620v1 Abstract: Encoder-free architectures have been preliminarily explored in the 2D visual domain, yet it remains an...

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles 15.02.2025

🤗 Upvotes: 16 | cs. CL, cs. AI Authors: Xintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan, Rui Xu, Jen-tse Huang, Siyu Yuan, Haoran Guo, Jiangjie Chen, Wei Wang, Yanghua Xiao, Shuchang Zhou Title: CoSER: Coordinating LLM-Based Persona Simulation of Established Roles Arxiv: http://arxiv.org/abs/2502.09082v1 Abstract: Role-playing language agents (RPLAs) have emerged as promising applications of la...

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models 15.02.2025

🤗 Upvotes: 15 | cs. CV, cs. AI Authors: Yangguang Li, Zi-Xin Zou, Zexiang Liu, Dehu Wang, Yuan Liang, Zhipeng Yu, Xingchao Liu, Yuan-Chen Guo, Ding Liang, Wanli Ouyang, Yan-Pei Cao Title: TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models Arxiv: http://arxiv.org/abs/2502.06608v1 Abstract: Recent advancements in diffusion techniques have propelled image and video gen...

Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance 14.02.2025

🤗 Upvotes: 40 | cs. CL Authors: Lingfei Qian, Weipeng Zhou, Yan Wang, Xueqing Peng, Jimin Huang, Qianqian Xie Title: Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance Arxiv: http://arxiv.org/abs/2502.08127v1 Abstract: Recent advancements in large language models (LLMs) have shown strong general reasoning abilities, yet their effectiveness in financial reasoning remains underexpl...

TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation 14.02.2025

🤗 Upvotes: 35 | cs. CV Authors: Alex Jinpeng Wang, Dongxing Mao, Jiawei Zhang, Weiming Han, Zhuobai Dong, Linjie Li, Yiqi Lin, Zhengyuan Yang, Libo Qin, Fuwei Zhang, Lijuan Wang, Min Li Title: TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation Arxiv: http://arxiv.org/abs/2502.07870v1 Abstract: Text-conditioned image generation has gained significant attention in recent years and a...

BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models 14.02.2025

🤗 Upvotes: 35 | cs. CL Authors: Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, Fei Yuan Title: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Arxiv: http://arxiv.org/abs/2502.07346v1 Abstract: Previous multilingual benchmarks focus primarily on simple understanding tasks, but for large language models(LLMs), we emphasize proficiency in instru...

CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation 14.02.2025

🤗 Upvotes: 29 | cs. CV Authors: Qinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia, Huchuan Lu, Tianfan Xue, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai Title: CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation Arxiv: http://arxiv.org/abs/2502.08639v1 Abstract: In this work, we present CineMaster, a novel framework for 3D-aware and controllable text-to-video generati...

Distillation Scaling Laws 14.02.2025

🤗 Upvotes: 26 | cs. LG, cs. AI, cs. CL, stat. ML Authors: Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb Title: Distillation Scaling Laws Arxiv: http://arxiv.org/abs/2502.08606v1 Abstract: We provide a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings...

TransMLA: Multi-Head Latent Attention Is All You Need 14.02.2025

🤗 Upvotes: 25 | cs. LG, cs. AI Authors: Fanxu Meng, Zengwei Yao, Muhan Zhang Title: TransMLA: Multi-Head Latent Attention Is All You Need Arxiv: http://arxiv.org/abs/2502.07864v2 Abstract: Modern large language models (LLMs) often encounter communication bottlenecks on current hardware, rather than purely computational constraints. Multi-head Latent Attention (MLA) tackles this challenge by using...

WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation 14.02.2025

🤗 Upvotes: 21 | cs. AI, cs. MA Authors: Henry Hengyuan Zhao, Difei Gao, Mike Zheng Shou Title: WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation Arxiv: http://arxiv.org/abs/2502.08047v1 Abstract: Current GUI agents have achieved outstanding performance in GUI element grounding. However, planning remains highly challenging, especially due to sensitivity to the initial state of the...

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid 14.02.2025

🤗 Upvotes: 19 | cs. LG, cs. AI, cs. CL Authors: Weigao Sun, Disen Lan, Yiran Zhong, Xiaoye Qu, Yu Cheng Title: LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Arxiv: http://arxiv.org/abs/2502.07563v1 Abstract: Linear sequence modeling approaches, such as linear attention, provide advantages like linear-time training and constant-memory inference over sequence lengths....

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning 14.02.2025

🤗 Upvotes: 11 | cs. CL, cs. LG Authors: Jean Vassoyan, Nathanaël Beau, Roman Plaud Title: Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Arxiv: http://arxiv.org/abs/2502.06533v1 Abstract: The ability to achieve long-term goals is a key challenge in the current development of large language models (LLMs). To address this, pre-trained LLMs can be fine-tuned...

Expect the Unexpected: FailSafe Long Context QA for Finance 13.02.2025

🤗 Upvotes: 105 | cs. CL Authors: Kiran Kamble, Melisa Russak, Dmytro Mozolevskyi, Muayad Ali, Mateusz Russak, Waseem AlShikh Title: Expect the Unexpected: FailSafe Long Context QA for Finance Arxiv: http://arxiv.org/abs/2502.06329v1 Abstract: We propose a new long-context financial benchmark, FailSafeQA, designed to test the robustness and context-awareness of LLMs against six variations in human...

Competitive Programming with Large Reasoning Models 13.02.2025

🤗 Upvotes: 42 | cs. LG, cs. AI, cs. CL Authors: OpenAI, :, Ahmed El-Kishky, Alexander Wei, Andre Saraiva, Borys Minaev, Daniel Selsam, David Dohan, Francis Song, Hunter Lightman, Ignasi Clavera, Jakub Pachocki, Jerry Tworek, Lorenz Kuhn, Lukasz Kaiser, Mark Chen, Max Schwarzer, Mostafa Rohaninejad, Nat McAleese, o3 contributors, Oleg Mürk, Rhythm Garg, Rui Shu, Szymon Sidor, Vineet Kosaraju, Wend...

Enhancing Financial Time-Series Forecasting with Retrieval-Augmented Large Language Models 13.02.2025

🤗 Upvotes: 25 | cs. CL Authors: Mengxi Xiao, Zihao Jiang, Lingfei Qian, Zhengyu Chen, Yueru He, Yijing Xu, Yuecheng Jiang, Dong Li, Ruey-Ling Weng, Min Peng, Jimin Huang, Sophia Ananiadou, Qianqian Xie Title: Enhancing Financial Time-Series Forecasting with Retrieval-Augmented Large Language Models Arxiv: http://arxiv.org/abs/2502.05878v2 Abstract: Stock movement prediction, a critical task in fi...

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction 13.02.2025

🤗 Upvotes: 23 | cs. CL, cs. AI Authors: Junlong Li, Daya Guo, Dejian Yang, Runxin Xu, Yu Wu, Junxian He Title: CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction Arxiv: http://arxiv.org/abs/2502.07316v2 Abstract: Reasoning is a fundamental capability of Large Language Models. While prior research predominantly focuses on enhancing narrow skills like math or code generation, i...

Magic 1-For-1: Generating One Minute Video Clips within One Minute 13.02.2025

🤗 Upvotes: 20 | cs. CV Authors: Hongwei Yi, Shitong Shao, Tian Ye, Jiantong Zhao, Qingyu Yin, Michael Lingelbach, Li Yuan, Yonghong Tian, Enze Xie, Daquan Zhou Title: Magic 1-For-1: Generating One Minute Video Clips within One Minute Arxiv: http://arxiv.org/abs/2502.07701v1 Abstract: In this technical report, we present Magic 1-For-1 (Magic141), an efficient video generation model with optimized...

LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters! 13.02.2025

🤗 Upvotes: 20 | cs. AI Authors: Dacheng Li, Shiyi Cao, Tyler Griggs, Shu Liu, Xiangxi Mo, Shishir G. Patil, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica Title: LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters! Arxiv: http://arxiv.org/abs/2502.07374v1 Abstract: Large reasoning models (LRMs) tackle complex reasoning problems by following long chain-of-tho...

Teaching Language Models to Critique via Reinforcement Learning 13.02.2025

🤗 Upvotes: 16 | cs. LG, cs. AI, cs. CL Authors: Zhihui Xie, Jie chen, Liyu Chen, Weichao Mao, Jingjing Xu, Lingpeng Kong Title: Teaching Language Models to Critique via Reinforcement Learning Arxiv: http://arxiv.org/abs/2502.03492v1 Abstract: Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundament...

Scaling Pre-training to One Hundred Billion Data for Vision Language Models 13.02.2025

🤗 Upvotes: 15 | cs. CV Authors: Xiao Wang, Ibrahim Alabdulmohsin, Daniel Salz, Zhe Li, Keran Rong, Xiaohua Zhai Title: Scaling Pre-training to One Hundred Billion Data for Vision Language Models Arxiv: http://arxiv.org/abs/2502.07617v1 Abstract: We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We find th...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.