Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Enhance-A-Video: Better Generated Video for Free 13.02.2025

🤗 Upvotes: 14 | cs. CV Authors: Yang Luo, Xuanlei Zhao, Mengzhao Chen, Kaipeng Zhang, Wenqi Shao, Kai Wang, Zhangyang Wang, Yang You Title: Enhance-A-Video: Better Generated Video for Free Arxiv: http://arxiv.org/abs/2502.07508v1 Abstract: DiT-based video generation has achieved remarkable results, but research into enhancing existing models remains relatively unexplored. In this work, we introdu...

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 12.02.2025

🤗 Upvotes: 71 | cs. CL Authors: Runze Liu, Junqi Gao, Jian Zhao, Kaiyan Zhang, Xiu Li, Biqing Qi, Wanli Ouyang, Bowen Zhou Title: Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Arxiv: http://arxiv.org/abs/2502.06703v1 Abstract: Test-Time Scaling (TTS) is an important method for improving the performance of Large Language Models (LLMs) by using additional computation dur...

SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators 12.02.2025

🤗 Upvotes: 71 | cs. CL Authors: Daniil Moskovskiy, Nikita Sushko, Sergey Pletenev, Elena Tutubalina, Alexander Panchenko Title: SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators Arxiv: http://arxiv.org/abs/2502.06394v1 Abstract: Existing approaches to multilingual text detoxification are hampered by the scarcity of parallel multilingual datasets. In this work, we intro...

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 12.02.2025

🤗 Upvotes: 36 | cs. CL, cs. LG Authors: Chengqi Lyu, Songyang Gao, Yuzhe Gu, Wenwei Zhang, Jianfei Gao, Kuikun Liu, Ziyi Wang, Shuaibin Li, Qian Zhao, Haian Huang, Weihan Cao, Jiangning Liu, Hongwei Liu, Junnan Liu, Songyang Zhang, Dahua Lin, Kai Chen Title: Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Arxiv: http://arxiv.org/abs/2502.06781v1 Abstract: Reasoning abili...

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning 12.02.2025

🤗 Upvotes: 22 | cs. AI, cs. CL, cs. LG, cs. MA Authors: Bidipta Sarkar, Warren Xia, C. Karen Liu, Dorsa Sadigh Title: Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Arxiv: http://arxiv.org/abs/2502.06060v1 Abstract: Communicating in natural language is a powerful tool in multi-agent settings, as it enables independent agents to share information in partially...

CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging 12.02.2025

🤗 Upvotes: 17 | cs. CL, cs. AI Authors: Md. Ashraful Islam, Mohammed Eunus Ali, Md Rizwan Parvez Title: CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging Arxiv: http://arxiv.org/abs/2502.05664v1 Abstract: Large Language Models (LLMs) have made significant strides in code generation and problem solving. Current approaches employ external tool...

LM2: Large Memory Models 12.02.2025

🤗 Upvotes: 16 | cs. CL, cs. AI Authors: Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J. Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, Andy Toulis Title: LM2: Large Memory Models Arxiv: http://arxiv.org/abs/2502.06049v1 Abstract: This paper introduces the Large Memory Model (LM2), a decoder-only Transformer architecture enhanced with an auxiliary memory module that aims to address the...

Matryoshka Quantization 12.02.2025

🤗 Upvotes: 13 | cs. LG, cs. AI Authors: Pranav Nair, Puranjay Datta, Jeff Dean, Prateek Jain, Aditya Kusupati Title: Matryoshka Quantization Arxiv: http://arxiv.org/abs/2502.06786v1 Abstract: Quantizing model weights is critical for reducing the communication and inference costs of large models. However, quantizing models -- especially to low precisions like int4 or int2 -- requires a trade-off i...

Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation 12.02.2025

🤗 Upvotes: 13 | cs. CV, cs. AI Authors: Chenkai Xu, Xu Wang, Zhenyi Liao, Yishun Li, Tianqi Hou, Zhijie Deng Title: Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation Arxiv: http://arxiv.org/abs/2502.05415v1 Abstract: There has been increasing research interest in building unified multimodal understanding and generation models, among which Show-o stands as a notable...

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding 12.02.2025

🤗 Upvotes: 12 | cs. CL Authors: Sukmin Cho, Sangjin Choi, Taeho Hwang, Jeongyeon Seo, Soyeong Jeong, Huije Lee, Hoyun Song, Jong C. Park, Youngjin Kwon Title: Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Arxiv: http://arxiv.org/abs/2502.05609v1 Abstract: Accelerating inference in Large Language Models (LLMs) is critic...

ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates 12.02.2025

🤗 Upvotes: 11 | cs. CL Authors: Ling Yang, Zhaochen Yu, Bin Cui, Mengdi Wang Title: ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates Arxiv: http://arxiv.org/abs/2502.06772v1 Abstract: We present that hierarchical LLM reasoning via scaling thought templates can effectively optimize the reasoning search space and outperform the mathematical reasoning capabilities of powerful LLM...

VideoRoPE: What Makes for Good Video Rotary Position Embedding? 11.02.2025

🤗 Upvotes: 52 | cs. CV Authors: Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Jian Tong, Haodong Duan, Qipeng Guo, Jiaqi Wang, Xipeng Qiu, Dahua Lin Title: VideoRoPE: What Makes for Good Video Rotary Position Embedding? Arxiv: http://arxiv.org/abs/2502.05173v1 Abstract: While Rotary Position Embedding (RoPE) and its variants are widely adopted for their long-context cap...

Fast Video Generation with Sliding Tile Attention 11.02.2025

🤗 Upvotes: 39 | cs. CV Authors: Peiyuan Zhang, Yongqi Chen, Runlong Su, Hangliang Ding, Ion Stoica, Zhenghong Liu, Hao Zhang Title: Fast Video Generation with Sliding Tile Attention Arxiv: http://arxiv.org/abs/2502.04507v1 Abstract: Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-sec...

Goku: Flow Based Video Generative Foundation Models 11.02.2025

🤗 Upvotes: 39 | cs. CV Authors: Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, Ting-Che Lin, Shilong Zhang, Fu Li, Chuan Li, Xing Wang, Yanghua Peng, Peize Sun, Ping Luo, Yi Jiang, Zehuan Yuan, Bingyue Peng, Xiaobing Liu Title: Goku: Flow Based Video Generative Foundation Models Arxiv: http://arxiv.org/abs/2502.04896v2 Abstra...

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations 11.02.2025

🤗 Upvotes: 32 | cs. LG Authors: Andrei Panferov, Jiale Chen, Soroush Tabesh, Roberto L. Castro, Mahdi Nikdan, Dan Alistarh Title: QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Arxiv: http://arxiv.org/abs/2502.05003v1 Abstract: One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment....

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach 11.02.2025

🤗 Upvotes: 30 | cs. LG, cs. CL Authors: Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein Title: Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Arxiv: http://arxiv.org/abs/2502.05171v1 Abstract: We study a novel language model architecture that is capable of scaling test...

AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting 11.02.2025

🤗 Upvotes: 23 | cs. CV Authors: Chung-Ho Wu, Yang-Jung Chen, Ying-Huan Chen, Jie-Ying Lee, Bo-Hsu Ke, Chun-Wei Tuan Mu, Yi-Chuan Huang, Chin-Yang Lin, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu Title: AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting Arxiv: http://arxiv.org/abs/2502.05176v1 Abstract: Three-dimensional scene inpainting is crucial for a...

DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails 11.02.2025

🤗 Upvotes: 18 | cs. CL, cs. LG Authors: Yihe Deng, Yu Yang, Junkai Zhang, Wei Wang, Bo Li Title: DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails Arxiv: http://arxiv.org/abs/2502.05163v1 Abstract: The rapid advancement of large language models (LLMs) has increased the need for guardrail models to ensure responsible use, particularly in detecting unsafe and illegal conten...

Agency Is Frame-Dependent 11.02.2025

🤗 Upvotes: 15 | cs. AI Authors: David Abel, André Barreto, Michael Bowling, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh Title: Agency Is Frame-Dependent Arxiv: http://arxiv.org/abs/2502.04403v1 Abstract: Agency is a system's capacity to steer outco...

FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation 11.02.2025

🤗 Upvotes: 14 | cs. CV Authors: Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge, Peize Sun, Yida Zhang, Yi Jiang, Zehuan Yuan, Binyue Peng, Ping Luo Title: FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation Arxiv: http://arxiv.org/abs/2502.05179v1 Abstract: DiT diffusion models have achieved great success in text-to-video generation, leveraging their scalabili...

Generating Symbolic World Models via Test-time Scaling of Large Language Models 11.02.2025

🤗 Upvotes: 13 | cs. AI Authors: Zhouliang Yu, Yuhuan Yuan, Tim Z. Xiao, Fuxiang Frank Xia, Jie Fu, Ge Zhang, Ge Lin, Weiyang Liu Title: Generating Symbolic World Models via Test-time Scaling of Large Language Models Arxiv: http://arxiv.org/abs/2502.04728v1 Abstract: Solving complex planning problems requires Large Language Models (LLMs) to explicitly model the state transition to avoid rule viola...

Analyze Feature Flow to Enhance Interpretation and Steering in Language Models 08.02.2025

🤗 Upvotes: 41 | cs. LG, cs. CL Authors: Daniil Laptev, Nikita Balagansky, Yaroslav Aksenov, Daniil Gavrilov Title: Analyze Feature Flow to Enhance Interpretation and Steering in Language Models Arxiv: http://arxiv.org/abs/2502.03032v2 Abstract: We introduce a new approach to systematically map features discovered by sparse autoencoder across consecutive layers of large language models, extending...

UltraIF: Advancing Instruction Following from the Wild 08.02.2025

🤗 Upvotes: 15 | cs. CL, cs. AI Authors: Kaikai An, Li Sheng, Ganqu Cui, Shuzheng Si, Ning Ding, Yu Cheng, Baobao Chang Title: UltraIF: Advancing Instruction Following from the Wild Arxiv: http://arxiv.org/abs/2502.04153v1 Abstract: Instruction-following made modern large language models (LLMs) helpful assistants. However, the key to taming LLMs on complex instructions remains mysterious, for that...

Great Models Think Alike and this Undermines AI Oversight 08.02.2025

🤗 Upvotes: 14 | cs. LG, cs. AI, cs. CL Authors: Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, Jonas Geiping Title: Great Models Think Alike and this Undermines AI Oversight Arxiv: http://arxiv.org/abs/2502.04313v1 Abstract: As Language Model (LM) capabilities advance, evaluating and supervising them at sca...

Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2 08.02.2025

🤗 Upvotes: 14 | cs. AI, cs. LG Authors: Yuri Chervonyi, Trieu H. Trinh, Miroslav Olšák, Xiaomeng Yang, Hoang Nguyen, Marcelo Menegali, Junehyuk Jung, Vikas Verma, Quoc V. Le, Thang Luong Title: Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2 Arxiv: http://arxiv.org/abs/2502.03544v1 Abstract: We present AlphaGeometry2, a significantly improved version of AlphaGeometry in...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.