Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 08.02.2025

🤗 Upvotes: 14 | cs. CV, cs. CL, cs. MM, cs. SD, eess. AS, eess. IV Authors: Zuyan Liu, Yuhao Dong, Jiahui Wang, Ziwei Liu, Winston Hu, Jiwen Lu, Yongming Rao Title: Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment Arxiv: http://arxiv.org/abs/2502.04328v1 Abstract: Recent advances in large language models, particularly following GPT-4o, have sparked incre...

MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm 08.02.2025

🤗 Upvotes: 13 | cs. CV Authors: Ziyan Guo, Zeyu Hu, Na Zhao, De Wen Soh Title: MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm Arxiv: http://arxiv.org/abs/2502.02358v3 Abstract: Human motion generation and editing are key components of computer graphics and vision. However, current approaches in this field tend to offer isolated solutions tailored t...

MAGA: MAssive Genre-Audience Reformulation to Pretraining Corpus Expansion 08.02.2025

🤗 Upvotes: 13 | cs. CL Authors: Xintong Hao, Ke Shen, Chenggang Li Title: MAGA: MAssive Genre-Audience Reformulation to Pretraining Corpus Expansion Arxiv: http://arxiv.org/abs/2502.04235v1 Abstract: Despite the remarkable capabilities of large language models across various tasks, their continued scaling faces a critical challenge: the scarcity of high-quality pretraining data. While model archi...

ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization 08.02.2025

🤗 Upvotes: 12 | cs. CL Authors: Yinjie Wang, Ling Yang, Guohao Li, Mengdi Wang, Bryon Aragam Title: ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Arxiv: http://arxiv.org/abs/2502.04306v1 Abstract: Recent research has leveraged large language model multi-agent systems for complex problem-solving while trying to reduce the manual effort required to build them, dri...

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis 08.02.2025

🤗 Upvotes: 10 | eess. AS, cs. AI, cs. CL, cs. MM, cs. SD Authors: Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi DAI, Hongzhan Lin, Jianyi Chen, Xingjian Du, Liumeng Xue, Yunlin Chen, Zhifei Li, Lei Xie, Qiuqiang Kong, Yike Guo, Wei Xue Title: Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis Arxiv: http...

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model 07.02.2025

🤗 Upvotes: 90 | cs. CL Authors: Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, Thomas...

TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets 07.02.2025

🤗 Upvotes: 27 | cs. CE, cs. CY Authors: Yuzhe Yang, Yifei Zhang, Minghao Wu, Kaidi Zhang, Yunmiao Zhang, Honghai Yu, Yan Hu, Benyou Wang Title: TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets Arxiv: http://arxiv.org/abs/2502.01506v2 Abstract: The study of social emergence has long been a central focus in social science. Traditional modeling approaches, such as rule-b...

Demystifying Long Chain-of-Thought Reasoning in LLMs 07.02.2025

🤗 Upvotes: 26 | cs. CL, cs. LG Authors: Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, Xiang Yue Title: Demystifying Long Chain-of-Thought Reasoning in LLMs Arxiv: http://arxiv.org/abs/2502.03373v1 Abstract: Scaling inference compute enhances reasoning in large language models (LLMs), with long chains-of-thought (CoTs) enabling strategies like backtracking and error correction. Reinforcement...

LIMO: Less is More for Reasoning 07.02.2025

🤗 Upvotes: 24 | cs. CL, cs. AI Authors: Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu Title: LIMO: Less is More for Reasoning Arxiv: http://arxiv.org/abs/2502.03387v1 Abstract: We present a fundamental discovery that challenges our understanding of how complex reasoning emerges in large language models. While conventional wisdom suggests that sophisticated reasoning tasks...

Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking 07.02.2025

🤗 Upvotes: 10 | cs. CL Authors: Jinyang Wu, Mingkuan Feng, Shuai Zhang, Ruihan Jin, Feihu Che, Zengqi Wen, Jianhua Tao Title: Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking Arxiv: http://arxiv.org/abs/2502.02339v1 Abstract: Multimodal large language models (MLLMs) exhibit impressive capabilities but still face challenges in complex visual reasoning. While recent efforts att...

LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer 07.02.2025

🤗 Upvotes: 7 | cs. CV Authors: Yiren Song, Danze Chen, Mike Zheng Shou Title: LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer Arxiv: http://arxiv.org/abs/2502.01105v1 Abstract: Generating cognitive-aligned layered SVGs remains challenging due to existing methods' tendencies toward either oversimplified single-layer outputs or optimization-induced shape redundancies....

On Teacher Hacking in Language Model Distillation 07.02.2025

🤗 Upvotes: 6 | cs. LG, cs. AI, cs. CL, stat. ML Authors: Daniil Tiapkin, Daniele Calandriello, Johan Ferret, Sarah Perrin, Nino Vieillard, Alexandre Ramé, Mathieu Blondel Title: On Teacher Hacking in Language Model Distillation Arxiv: http://arxiv.org/abs/2502.02671v1 Abstract: Post-training of language models (LMs) increasingly relies on the following two stages: (i) knowledge distillation, wher...

A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods 07.02.2025

🤗 Upvotes: 5 | cs. LG, cs. AI Authors: Isha Puri, Shivchander Sudalairaj, Guangxuan Xu, Kai Xu, Akash Srivastava Title: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods Arxiv: http://arxiv.org/abs/2502.01618v2 Abstract: Large language models (LLMs) have achieved significant performance gains via scaling up model sizes and/or data. Howev...

Jailbreaking with Universal Multi-Prompts 07.02.2025

🤗 Upvotes: 4 | cs. CL, cs. AI, cs. CR, cs. LG Authors: Yu-Ling Hsu, Hsuan Su, Shang-Tse Chen Title: Jailbreaking with Universal Multi-Prompts Arxiv: http://arxiv.org/abs/2502.01154v1 Abstract: Large language models (LLMs) have seen rapid development in recent years, revolutionizing various applications and significantly enhancing convenience and productivity. However, alongside their impressive c...

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models 06.02.2025

🤗 Upvotes: 29 | cs. CV Authors: Hila Chefer, Uriel Singer, Amit Zohar, Yuval Kirstain, Adam Polyak, Yaniv Taigman, Lior Wolf, Shelly Sheynin Title: VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Arxiv: http://arxiv.org/abs/2502.02492v1 Abstract: Despite tremendous recent progress, generative video models still struggle to capture real-world motion...

Inverse Bridge Matching Distillation 06.02.2025

🤗 Upvotes: 22 | cs. LG, cs. CV Authors: Nikita Gushchin, David Li, Daniil Selikhanovych, Evgeny Burnaev, Dmitry Baranchuk, Alexander Korotin Title: Inverse Bridge Matching Distillation Arxiv: http://arxiv.org/abs/2502.01362v1 Abstract: Learning diffusion bridge models is easy; making them fast and practical is an art. Diffusion bridge models (DBMs) are a promising extension of diffusion models fo...

ACECODER: Acing Coder RL via Automated Test-Case Synthesis 06.02.2025

🤗 Upvotes: 16 | cs. SE, cs. AI, cs. CL Authors: Huaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie, Xiaotong Chen, Wenhu Chen Title: ACECODER: Acing Coder RL via Automated Test-Case Synthesis Arxiv: http://arxiv.org/abs/2502.01718v1 Abstract: Most progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexpl...

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search 06.02.2025

🤗 Upvotes: 12 | cs. LG, cs. AI Authors: Zongyu Lin, Yao Tang, Xingcheng Yao, Da Yin, Ziniu Hu, Yizhou Sun, Kai-Wei Chang Title: QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search Arxiv: http://arxiv.org/abs/2502.02584v1 Abstract: Language agents have become a promising solution to complex interactive tasks. One of the key ingredients to the success of language agents is the rew...

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search 06.02.2025

🤗 Upvotes: 12 | cs. CL, cs. AI Authors: Maohao Shen, Guangtao Zeng, Zhenting Qi, Zhang-Wei Hong, Zhenfang Chen, Wei Lu, Gregory Wornell, Subhro Das, David Cox, Chuang Gan Title: Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Arxiv: http://arxiv.org/abs/2502.02508v1 Abstract: Large language models (LLMs) have demonstrated remarkable rea...

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? 06.02.2025

🤗 Upvotes: 7 | cs. CL, cs. LG Authors: Wenzhe Li, Yong Lin, Mengzhou Xia, Chi Jin Title: Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? Arxiv: http://arxiv.org/abs/2502.00674v1 Abstract: Ensembling outputs from diverse sources is a straightforward yet effective approach to boost performance. Mixture-of-Agents (MoA) is one such popular ensemble method that aggr...

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation 06.02.2025

🤗 Upvotes: 7 | cs. CV Authors: Xueqing Deng, Qihang Yu, Ali Athar, Chenglin Yang, Linjie Yang, Xiaojie Jin, Xiaohui Shen, Liang-Chieh Chen Title: COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation Arxiv: http://arxiv.org/abs/2502.02589v1 Abstract: This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation...

The Differences Between Direct Alignment Algorithms are a Blur 05.02.2025

🤗 Upvotes: 84 | cs. LG Authors: Alexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii, Alexey Malakhov, Daniil Gavrilov Title: The Differences Between Direct Alignment Algorithms are a Blur Arxiv: http://arxiv.org/abs/2502.01237v1 Abstract: Direct Alignment Algorithms (DAAs) simplify language model alignment by replacing reinforcement learning (RL) and reward modeling (RM) in Reinforcement Lea...

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models 05.02.2025

🤗 Upvotes: 83 | cs. CV Authors: Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, Chao Liang Title: OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models Arxiv: http://arxiv.org/abs/2502.01061v1 Abstract: End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing metho...

Process Reinforcement through Implicit Rewards 05.02.2025

🤗 Upvotes: 44 | cs. LG, cs. AI, cs. CL Authors: Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, Qixin Xu, Weize Chen, Jiarui Yuan, Huayu Chen, Kaiyan Zhang, Xingtai Lv, Shuo Wang, Yuan Yao, Xu Han, Hao Peng, Yu Cheng, Zhiyuan Liu, Maosong Sun, Bowen Zhou, Ning Ding Title: Process Reinforcement through Implicit Rewards Arxiv: http://arxiv.org/abs/2502...

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Understanding 05.02.2025

🤗 Upvotes: 25 | cs. CL Authors: Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar Title: Align...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.