Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Seed1.5-VL Technical Report 14.05.2025

🤗 Upvotes: 86 | cs. CV, cs. AI Authors: Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, Jingji Chen, Jingjia Huang, Kang Lei, Liping Yuan, Lishu Luo, Pengfei Liu, Qinghao Ye, Rui Qian, Shen Yan, Shixiong Zhao, Shuai Peng, Shuangye Li, Sihang Yuan, Sijin Wu, Tianheng Cheng, Weiwei Liu, Wenqian Wang, Xianhan Zeng, Xiao Liu, Xiaob...

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining 14.05.2025

🤗 Upvotes: 53 | cs. CL, cs. AI, cs. LG Authors: Xiaomi LLM-Core Team, :, Bingquan Xia, Bowen Shen, Cici, Dawei Zhu, Di Zhang, Gang Wang, Hailin Zhang, Huaqiu Liu, Jiebao Xiao, Jinhao Dong, Liang Zhao, Peidian Li, Peng Wang, Shihua Yu, Shimao Chen, Weikun Wang, Wenhan Ma, Xiangwei Deng, Yi Huang, Yifan Song, Zihan Jiang, Bowen Ye, Can Cai, Chenhong He, Dong Zhang, Duo Zhang, Guoan Wang, Hao Tian,...

Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets 14.05.2025

🤗 Upvotes: 48 | cs. CV Authors: Weiyu Li, Xuanyang Zhang, Zheng Sun, Di Qi, Hao Li, Wei Cheng, Weiwei Cai, Shihao Wu, Jiarui Liu, Zihao Wang, Xiao Chen, Feipeng Tian, Jianxiong Pan, Zeming Li, Gang Yu, Xiangyu Zhang, Daxin Jiang, Ping Tan Title: Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets Arxiv: http://arxiv.org/abs/2505.07747v1 Abstract: While generative ar...

Learning from Peers in Reasoning Models 14.05.2025

🤗 Upvotes: 34 | cs. CL Authors: Tongxu Luo, Wenyu Du, Jiaxi Bi, Stephen Chung, Zhengyang Tang, Hao Yang, Min Zhang, Benyou Wang Title: Learning from Peers in Reasoning Models Arxiv: http://arxiv.org/abs/2505.07787v1 Abstract: Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning proc...

Unified Continuous Generative Models 14.05.2025

🤗 Upvotes: 32 | cs. LG, cs. AI, cs. CV Authors: Peng Sun, Yi Jiang, Tao Lin Title: Unified Continuous Generative Models Arxiv: http://arxiv.org/abs/2505.07447v1 Abstract: Recent advances in continuous generative models, including multi-step approaches like diffusion and flow-matching (typically requiring 8-1000 sampling steps) and few-step methods such as consistency models (typically 1-8 steps),...

REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback 14.05.2025

🤗 Upvotes: 26 | cs. CL Authors: Aniruddha Roy, Pretam Ray, Abhilash Nandy, Somak Aditya, Pawan Goyal Title: REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback Arxiv: http://arxiv.org/abs/2505.06548v1 Abstract: Instruction-based Large Language Models (LLMs) have proven effective in numerous few-shot or...

Bielik v3 Small: Technical Report 13.05.2025

🤗 Upvotes: 53 | cs. LG, cs. AI, cs. CL, 68T50, I.2.7 Authors: Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel, Adrian Gwoździej Title: Bielik v3 Small: Technical Report Arxiv: http://arxiv.org/abs/2505.02550v2 Abstract: We introduce Bielik v3, a series of parameter-efficient generative text models (1.5B and 4.5B) optimized for Polish language processing. These models demonstrate...

Bielik 11B v2 Technical Report 13.05.2025

🤗 Upvotes: 44 | cs. CL, cs. AI, 68T50, I.2.7 Authors: Krzysztof Ociepa, Łukasz Flis, Krzysztof Wróbel, Adrian Gwoździej, Remigiusz Kinas Title: Bielik 11B v2 Technical Report Arxiv: http://arxiv.org/abs/2505.02410v2 Abstract: We present Bielik 11B v2, a state-of-the-art language model optimized for Polish text processing. Built on the Mistral 7B v0.2 architecture and scaled to 11B parameters usin...

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models 10.05.2025

🤗 Upvotes: 79 | cs. CV, cs. CL Authors: Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, Shouzheng Huang, Xinping Zhao, Borui Jiang, Lanqing Hong, Longyue Wang, Zhuotao Tian, Baoxing Huai, Wenhan Luo, Weihua Luo, Zheng Zhang, Baotian Hu, Min Zhang Title: Perception, Reason, Think, and Plan: A Survey on Large Multimodal...

On Path to Multimodal Generalist: General-Level and General-Bench 10.05.2025

🤗 Upvotes: 55 | cs. CV Authors: Hao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li, Qingshan Xu, Bobo Li, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Qingyu Shi, Zhiyuan Zhou, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Weiming Wu, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang, Tianjie Ju, Zixiang Meng, Shilin Xu, Liyu Jia, Wentao Hu, Meng Luo, Jiebo Luo, Tat-Seng Chua, Shu...

Flow-GRPO: Training Flow Matching Models via Online RL 10.05.2025

🤗 Upvotes: 36 | cs. CV, cs. AI Authors: Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, Wanli Ouyang Title: Flow-GRPO: Training Flow Matching Models via Online RL Arxiv: http://arxiv.org/abs/2505.05470v1 Abstract: We propose Flow-GRPO, the first method integrating online reinforcement learning (RL) into flow matching models. Our approach uses two...

Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities 09.05.2025

🤗 Upvotes: 57 | cs. CV Authors: Xinjie Zhang, Jintao Guo, Shanshan Zhao, Minghao Fu, Lunhao Duan, Guo-Hua Wang, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang Title: Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities Arxiv: http://arxiv.org/abs/2505.02567v2 Abstract: Recent years have seen remarkable progress in both multimodal understanding models an...

ZeroSearch: Incentivize the Search Capability of LLMs without Searching 09.05.2025

🤗 Upvotes: 35 | cs. CL Authors: Hao Sun, Zile Qiao, Jiayan Guo, Xuanbo Fan, Yingyan Hou, Yong Jiang, Pengjun Xie, Fei Huang, Yan Zhang Title: ZeroSearch: Incentivize the Search Capability of LLMs without Searching Arxiv: http://arxiv.org/abs/2505.04588v1 Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs)....

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning 08.05.2025

🤗 Upvotes: 67 | cs. CV Authors: Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang Title: Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning Arxiv: http://arxiv.org/abs/2505.03318v1 Abstract: Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human...

Absolute Zero: Reinforced Self-play Reasoning with Zero Data 08.05.2025

🤗 Upvotes: 63 | cs. LG, cs. AI, cs. CL Authors: Andrew Zhao, Yiran Wu, Yang Yue, Tong Wu, Quentin Xu, Yang Yue, Matthieu Lin, Shenzhi Wang, Qingyun Wu, Zilong Zheng, Gao Huang Title: Absolute Zero: Reinforced Self-play Reasoning with Zero Data Arxiv: http://arxiv.org/abs/2505.03335v2 Abstract: Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capab...

RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale 08.05.2025

🤗 Upvotes: 23 | cs. CL, cs. AI, cs. LG, I.2.7 Authors: Daniel Goldstein, Eric Alcaide, Janna Lu, Eugene Cheah Title: RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale Arxiv: http://arxiv.org/abs/2505.03005v1 Abstract: We present Rapid Attention Distillation to Linear Attention Decoders at Scale (RADLADS), a protocol for rapidly converting softmax attention transformers i...

FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios 08.05.2025

🤗 Upvotes: 21 | cs. CV, cs. AI, cs. MM Authors: Shiyi Zhang, Junhao Zhuang, Zhaoyang Zhang, Ying Shan, Yansong Tang Title: FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios Arxiv: http://arxiv.org/abs/2505.03730v1 Abstract: Action customization involves generating videos where the subject performs actions dictated by input control signals. Current methods use pose-guided or glo...

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play 07.05.2025

🤗 Upvotes: 56 | cs. AI, cs. CL, cs. SD Authors: Yemin Shi, Yu Shu, Siwei Dong, Guangyi Liu, Jaward Sesay, Jingwen Li, Zhiting Hu Title: Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Arxiv: http://arxiv.org/abs/2505.02707v1 Abstract: A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and...

RM-R1: Reward Modeling as Reasoning 07.05.2025

🤗 Upvotes: 48 | cs. CL, cs. AI, cs. LG Authors: Xiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin, Cheng Qian, Yu Wang, Hongru Wang, Yu Zhang, Denghui Zhang, Tong Zhang, Hanghang Tong, Heng Ji Title: RM-R1: Reward Modeling as Reasoning Arxiv: http://arxiv.org/abs/2505.02387v1 Abstract: Reward modeling is essential for aligning large language models (LLMs) with human preferences, especially through rei...

Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers 07.05.2025

🤗 Upvotes: 44 | cs. CL, cs. AI, cs. LG, I.2.7; I.2.6; I.2.3; I.7 Authors: Roman Abramov, Felix Steinbauer, Gjergji Kasneci Title: Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers Arxiv: http://arxiv.org/abs/2504.20752v1 Abstract: Transformers have achieved great success in numerous NLP tasks but continue to exhibit notable gaps in multi-step factual rea...

Practical Efficiency of Muon for Pretraining 07.05.2025

🤗 Upvotes: 30 | cs. LG, stat. ML Authors: Essential AI, :, Ishaan Shah, Anthony M. Polloreno, Karl Stratos, Philip Monk, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Ashish Tanwer, Darsh J Shah, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Peter Rushton, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani, Ashish...

PixelHacker: Image Inpainting with Structural and Semantic Consistency 06.05.2025

🤗 Upvotes: 24 | cs. CV Authors: Ziyang Xu, Kangsheng Duan, Xiaolei Shen, Zhifeng Ding, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang Title: PixelHacker: Image Inpainting with Structural and Semantic Consistency Arxiv: http://arxiv.org/abs/2504.20438v2 Abstract: Image inpainting is a fundamental research area between image editing and image generation. Recent state-of-the-art (SOTA) methods...

A Survey of Interactive Generative Video 03.05.2025

🤗 Upvotes: 31 | cs. CV Authors: Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Hao Chen, Xihui Liu Title: A Survey of Interactive Generative Video Arxiv: http://arxiv.org/abs/2504.21853v1 Abstract: Interactive Generative Video (IGV) has emerged as a crucial technology in response to the growing demand for high-quality, interactive video content across v...

DeepCritic: Deliberate Critique with Large Language Models 03.05.2025

🤗 Upvotes: 27 | cs. CL, cs. AI, cs. LG Authors: Wenkai Yang, Jingwen Chen, Yankai Lin, Ji-Rong Wen Title: DeepCritic: Deliberate Critique with Large Language Models Arxiv: http://arxiv.org/abs/2505.00662v1 Abstract: As Large Language Models (LLMs) are rapidly evolving, providing accurate feedback and scalable oversight on their outputs becomes an urgent and critical problem. Leveraging LLMs as cr...

Sadeed: Advancing Arabic Diacritization Through Small Language Model 02.05.2025

🤗 Upvotes: 44 | cs. CL, cs. AI Authors: Zeina Aldallal, Sara Chrouf, Khalil Hennara, Mohamed Motaism Hamed, Muhammad Hreden, Safwan AlModhayan Title: Sadeed: Advancing Arabic Diacritization Through Small Language Model Arxiv: http://arxiv.org/abs/2504.21635v1 Abstract: Arabic text diacritization remains a persistent challenge in natural language processing due to the language's morphological rich...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.