Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation 13.06.2025

🤗 Upvotes: 36 | cs. CV, cs. AI, cs. LG Authors: Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao, Lu Jiang Title: Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation Arxiv: http://arxiv.org/abs/2506.09350v1 Abstract: Existing large-scale video generation models are computationally intensive, preventing adoption in real-t...

ComfyUI-R1: Exploring Reasoning Models for Workflow Generation 13.06.2025

🤗 Upvotes: 34 | cs. CL, cs. CV, cs. SE Authors: Zhenran Xu, Yiyu Wang, Xue Yang, Longyue Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang Title: ComfyUI-R1: Exploring Reasoning Models for Workflow Generation Arxiv: http://arxiv.org/abs/2506.09790v1 Abstract: AI-generated content has evolved from monolithic models to modular workflows, particularly on platforms like ComfyUI, enabling customiza...

PlayerOne: Egocentric World Simulator 13.06.2025

🤗 Upvotes: 26 | cs. CV Authors: Yuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai, Fan Wang, Hengshuang Zhao Title: PlayerOne: Egocentric World Simulator Arxiv: http://arxiv.org/abs/2506.09995v1 Abstract: We introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric scene image from the u...

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation 13.06.2025

🤗 Upvotes: 26 | cs. SD, cs. AI, cs. LG, eess. AS Authors: Or Tal, Felix Kreuk, Yossi Adi Title: Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation Arxiv: http://arxiv.org/abs/2506.08570v2 Abstract: Recent progress in text-to-music generation has enabled models to synthesize high-quality musical segments, full compositions, and even respond to...

Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models 12.06.2025

🤗 Upvotes: 54 | cs. CL Authors: Mikhail Salnikov, Dmitrii Korzh, Ivan Lazichny, Elvir Karimov, Artyom Iudin, Ivan Oseledets, Oleg Y. Rogov, Alexander Panchenko, Natalia Loukachevitch, Elena Tutubalina Title: Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models Arxiv: http://arxiv.org/abs/2506.06751v1 Abstract: This paper evaluates geop...

Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better 12.06.2025

🤗 Upvotes: 26 | cs. CV, cs. AI, cs. CL Authors: Dianyi Wang, Wei Song, Yikun Wang, Siyuan Wang, Kaicheng Yu, Zhongyu Wei, Jiaqi Wang Title: Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better Arxiv: http://arxiv.org/abs/2506.09040v1 Abstract: Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporatin...

RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling 12.06.2025

🤗 Upvotes: 25 | cs. CL, cs. AI, cs. LG Authors: Yang Liu, Jiaqi Li, Zilong Zheng Title: RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling Arxiv: http://arxiv.org/abs/2506.08672v1 Abstract: Rule-based reasoning has been acknowledged as one of the fundamental problems in reasoning, while deviations in rule formats, types, and complexity in real-world applications pose...

Reinforcement Pre-Training 11.06.2025

🤗 Upvotes: 150 | cs. CL Authors: Qingxiu Dong, Li Dong, Yao Tang, Tianzhu Ye, Yutao Sun, Zhifang Sui, Furu Wei Title: Reinforcement Pre-Training Arxiv: http://arxiv.org/abs/2506.08007v1 Abstract: In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token prediction as a reason...

Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance 11.06.2025

🤗 Upvotes: 62 | cs. LG, cs. AI, cs. CR Authors: Ruizhong Qiu, Gaotang Li, Tianxin Wei, Jingrui He, Hanghang Tong Title: Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance Arxiv: http://arxiv.org/abs/2506.06444v1 Abstract: Existing safety assurance research has primarily focused on training-phase alignment to instill safe behaviors into LLMs. However, recent studies have exp...

MiniCPM4: Ultra-Efficient LLMs on End Devices 11.06.2025

🤗 Upvotes: 60 | cs. CL, cs. AI Authors: MiniCPM Team, Chaojun Xiao, Yuxuan Li, Xu Han, Yuzhuo Bai, Jie Cai, Haotian Chen, Wentong Chen, Xin Cong, Ganqu Cui, Ning Ding, Shengdan Fan, Yewei Fang, Zixuan Fu, Wenyu Guan, Yitong Guan, Junshao Guo, Yufeng Han, Bingxiang He, Yuxiang Huang, Cunliang Kong, Qiuzuo Li, Siyuan Li, Wenhao Li, Yanghao Li, Yishan Li, Zhen Li, Dan Liu, Biyuan Lin, Yankai Lin, Xi...

SpatialLM: Training Large Language Models for Structured Indoor Modeling 11.06.2025

🤗 Upvotes: 31 | cs. CV Authors: Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng, Rui Tang, Hao Zhu, Ping Tan, Zihan Zhou Title: SpatialLM: Training Large Language Models for Structured Indoor Modeling Arxiv: http://arxiv.org/abs/2506.07491v1 Abstract: SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs...

Image Reconstruction as a Tool for Feature Analysis 11.06.2025

🤗 Upvotes: 27 | cs. CV, 68T10, 68T30, 68T45, I.2.10 Authors: Eduard Allakhverdov, Dmitrii Tarasov, Elizaveta Goncharova, Andrey Kuznetsov Title: Image Reconstruction as a Tool for Feature Analysis Arxiv: http://arxiv.org/abs/2506.07803v1 Abstract: Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite the...

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning 11.06.2025

🤗 Upvotes: 25 | cs. RO, cs. AI Authors: Sheng Chen, Peiyu He, Jiaxin Hu, Ziyang Liu, Yansheng Wang, Tao Xu, Chi Zhang, Chongchong Zhang, Chao An, Shiyu Cai, Duo Cao, Kangping Chen, Shuai Chu, Tianwei Chu, Mingdi Dan, Min Du, Weiwei Fang, Pengyou Fu, Junkai Hu, Xiaowei Jiang, Zhaodi Jiang, Fuxuan Li, Jun Li, Minghui Li, Mingyao Li, Yanchang Li, Zhibin Li, Guangming Liu, Kairui Liu, Lihao Liu, Weiz...

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA 10.06.2025

🤗 Upvotes: 83 | cs. CL Authors: Sergey Pletenev, Maria Marina, Nikolay Ivanov, Daria Galimzianova, Nikita Krayko, Mikhail Salnikov, Vasily Konovalov, Alexander Panchenko, Viktor Moskvoretskii Title: Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA Arxiv: http://arxiv.org/abs/2505.21115v1 Abstract: Large Language Models (LLMs) often hallucina...

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion 10.06.2025

🤗 Upvotes: 27 | cs. SD, cs. AI, eess. AS Authors: Shunian Chen, Xinyuan Xie, Zheshu Chen, Liyan Zhao, Owen Lee, Zhan Su, Qilin Sun, Benyou Wang Title: FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Arxiv: http://arxiv.org/abs/2506.01111v1 Abstract: High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automa...

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning 10.06.2025

🤗 Upvotes: 26 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa, Hyunwoo Jae, Keenan Powell, Minghui Liu, Neel Jay, Sungbin Oh, Xiyao Wang, Yongyuan Liang, Tom Goldstein, Furong Huang Title: MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning Arxiv: http://arxiv.org/abs/2506.05523v1 Abstract: Despite rapid...

Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs 10.06.2025

🤗 Upvotes: 25 | cs. CL Authors: Ananth Muppidi, Abhilash Nandy, Sambaran Bandyopadhyay Title: Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs Arxiv: http://arxiv.org/abs/2506.05629v1 Abstract: The performance of large language models in domain-specific tasks necessitates fine-tuning, which is computationally expensive and technically challenging. This paper focuses on paramet...

SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training 07.06.2025

🤗 Upvotes: 39 | cs. CV Authors: Jianyi Wang, Shanchuan Lin, Zhijie Lin, Yuxi Ren, Meng Wei, Zongsheng Yue, Shangchen Zhou, Hao Chen, Yang Zhao, Ceyuan Yang, Xuefeng Xiao, Chen Change Loy, Lu Jiang Title: SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training Arxiv: http://arxiv.org/abs/2506.05301v1 Abstract: Recent advances in diffusion-based video restoration (VR) demonstrat...

ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development 07.06.2025

🤗 Upvotes: 38 | cs. CL, cs. CV Authors: Zhenran Xu, Xue Yang, Yiyu Wang, Qingli Hu, Zijiao Wu, Longyue Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang Title: ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development Arxiv: http://arxiv.org/abs/2506.05010v1 Abstract: We introduce ComfyUI-Copilot, a large language model-powered plugin designed to enhance the usability and ef...

Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts 07.06.2025

🤗 Upvotes: 32 | cs. LG, cs. CL Authors: Danil Sivtsov, Ivan Rodkin, Gleb Kuzmin, Yuri Kuratov, Ivan Oseledets Title: Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts Arxiv: http://arxiv.org/abs/2506.05229v1 Abstract: Transformer models struggle with long-context inference due to their quadratic time and linear memory complexity. Recurrent Memory Transformer...

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics 07.06.2025

🤗 Upvotes: 32 | cs. RO, cs. AI, cs. CV Authors: Enshen Zhou, Jingkun An, Cheng Chi, Yi Han, Shanyu Rong, Chi Zhang, Pengwei Wang, Zhongyuan Wang, Tiejun Huang, Lu Sheng, Shanghang Zhang Title: RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics Arxiv: http://arxiv.org/abs/2506.04308v1 Abstract: Spatial referring is a fundamental capability of embodied robots...

Video World Models with Long-term Spatial Memory 07.06.2025

🤗 Upvotes: 30 | cs. CV Authors: Tong Wu, Shuai Yang, Ryan Po, Yinghao Xu, Ziwei Liu, Dahua Lin, Gordon Wetzstein Title: Video World Models with Long-term Spatial Memory Arxiv: http://arxiv.org/abs/2506.05284v1 Abstract: Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited tempora...

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights 07.06.2025

🤗 Upvotes: 27 | cs. AI Authors: Mathieu Andreux, Breno Baldas Skuk, Hamza Benchekroun, Emilien Biré, Antoine Bonnet, Riaz Bordie, Matthias Brunel, Pierre-Louis Cedoz, Antoine Chassang, Mickaël Chen, Alexandra D. Constantinou, Antoine d'Andigné, Hubert de La Jonquière, Aurélien Delfosse, Ludovic Denoyer, Alexis Deprez, Augustin Derupti, Michael Eickenberg, Mathïs Federico, Charles Kantor, Xavier K...

Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models 07.06.2025

🤗 Upvotes: 24 | cs. CL Authors: Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, Jingren Zhou Title: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models Arxiv: http://arxiv.org/abs/2506.05176v1 Abstract: In this work, we introduce the Qwen3 Embedding series, a significant advanceme...

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models 07.06.2025

🤗 Upvotes: 23 | cs. CV Authors: Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Yan, Yu Cheng Title: VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Arxiv: http://arxiv.org/abs/2505.23656v1 Abstract: Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthe...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.