Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Diffusion Language Models are Super Data Learners 07.11.2025

🤗 Upvotes: 67 | cs. LG Authors: Jinjie Ni, Qian Liu, Longxu Dou, Chao Du, Zili Wang, Hang Yan, Tianyu Pang, Michael Qizhe Shieh Title: Diffusion Language Models are Super Data Learners Arxiv: http://arxiv.org/abs/2511.03276v1 Abstract: Under strictly controlled pre-training settings, we observe a Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoreg...

LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation 07.11.2025

🤗 Upvotes: 39 | cs. CL Authors: Gyeom Hwangbo, Hyungjoo Chae, Minseok Kang, Hyeonjong Ju, Soohyun Oh, Jinyoung Yeo Title: LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation Arxiv: http://arxiv.org/abs/2511.03001v1 Abstract: Despite recent progress in using Large Language Models (LLMs) for automatically generating 3D scenes, generated scenes...

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions 07.11.2025

🤗 Upvotes: 39 | cs. CV Authors: Guozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng, Youliang Zhang, Yi Chen, Yuan Zhou, Qinglin Lu, Limin Wang Title: UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions Arxiv: http://arxiv.org/abs/2511.03334v1 Abstract: Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhi...

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization 06.11.2025

🤗 Upvotes: 71 | cs. LG, cs. AI, cs. RO Authors: Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov Title: Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization Arxiv: http://arxiv.org/abs/2510.25616v1 Abstract: The growing success of Vision-Language-Action (VLA) models stems from the promise that pretrained Vision-Language Models (VLMs...

VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation 06.11.2025

🤗 Upvotes: 65 | cs. CV, cs. CL Authors: Kevin Qinghong Lin, Yuhao Zheng, Hangyu Ran, Dantong Zhu, Dongxing Mao, Linjie Li, Philip Torr, Alex Jinpeng Wang Title: VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation Arxiv: http://arxiv.org/abs/2511.02778v1 Abstract: Code has emerged as a precise and executable medium for reasoning and action in the agent era. Yet, progres...

When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought 06.11.2025

🤗 Upvotes: 42 | cs. CV Authors: Yiyang Zhou, Haoqin Tu, Zijun Wang, Zeyu Wang, Niklas Muennighoff, Fan Nie, Yejin Choi, James Zou, Chaorui Deng, Shen Yan, Haoqi Fan, Cihang Xie, Huaxiu Yao, Qinghao Ye Title: When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought Arxiv: http://arxiv.org/abs/2511.02779v1 Abstract: We propose MIRA, a new benchmark designed to...

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation 05.11.2025

🤗 Upvotes: 61 | cs. CL, cs. AI Authors: Ling-Team, Ang Li, Ben Liu, Binbin Hu, Bing Li, Bingwei Zeng, Borui Ye, Caizhi Tang, Changxin Tian, Chao Huang, Chao Zhang, Chen Qian, Chenchen Ju, Chenchen Li, Chengfu Tang, Chili Fu, Chunshao Ren, Chunwei Wu, Cong Zhang, Cunyin Peng, Dafeng Xu, Daixin Wang, Dalong Zhang, Dingnan Jin, Dingyuan Zhu, Dongke Hu, Fangzheng Zhao, Feifan Wu, Feng Zhu, Gangshan W...

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph 05.11.2025

🤗 Upvotes: 34 | cs. LG, cs. AI, cs. CL, I.2.7 Authors: Fali Wang, Jihai Chen, Shuhua Yang, Runxue Bao, Tianxiang Zhao, Zhiwei Zhang, Xianfeng Tang, Hui Liu, Qi He, Suhang Wang Title: Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph Arxiv: http://arxiv.org/abs/2511.00086v1 Abstract: Test-Time Scaling (TTS) improves large language models (LLMs) by allocating additional computa...

The Underappreciated Power of Vision Models for Graph Structural Understanding 05.11.2025

🤗 Upvotes: 31 | cs. CV, cs. AI, cs. LG Authors: Xinjian Zhao, Wei Pang, Zhongkai Xue, Xiangru Jian, Lei Zhang, Yaoyao Xu, Xiaozhuang Song, Shu Wu, Tianshu Yu Title: The Underappreciated Power of Vision Models for Graph Structural Understanding Arxiv: http://arxiv.org/abs/2510.24788v1 Abstract: Graph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visu...

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback 05.11.2025

🤗 Upvotes: 27 | cs. CV Authors: Ropeway Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing, Jinwang Wang, Rui Zhao, Yan Xing, Weihua Chen, Fan Wang Title: UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback Arxiv: http://arxiv.org/abs/2511.01678v1 Abstract: Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown s...

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation 05.11.2025

🤗 Upvotes: 24 | cs. CV Authors: Yongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma, Xiyao Wang, Jiageng Mao, Jiuhai Chen, Jiatao Gu, Yue Wang, Furong Huang Title: ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation Arxiv: http://arxiv.org/abs/2511.01163v1 Abstract: Unified multimodal models (UMMs) have emerged as a powerful paradigm for seamlessly unifying text and image under...

PHUMA: Physically-Grounded Humanoid Locomotion Dataset 05.11.2025

🤗 Upvotes: 23 | cs. RO Authors: Kyungmin Lee, Sibeen Kim, Minho Park, Hyunseung Kim, Dongyoon Hwang, Hojoon Lee, Jaegul Choo Title: PHUMA: Physically-Grounded Humanoid Locomotion Dataset Arxiv: http://arxiv.org/abs/2510.26236v1 Abstract: Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on high-quality...

UniREditBench: A Unified Reasoning-based Image Editing Benchmark 05.11.2025

🤗 Upvotes: 22 | cs. CV Authors: Feng Han, Yibin Wang, Chenglin Li, Zheming Liang, Dianyi Wang, Yang Jiao, Zhipeng Wei, Chao Gong, Cheng Jin, Jingjing Chen, Jiaqi Wang Title: UniREditBench: A Unified Reasoning-based Image Editing Benchmark Arxiv: http://arxiv.org/abs/2511.01295v1 Abstract: Recent advances in multi-modal generative models have driven substantial improvements in image editing. Howev...

World Simulation with Video Foundation Models for Physical AI 05.11.2025

🤗 Upvotes: 22 | cs. CV, cs. AI, cs. LG, cs. RO Authors: NVIDIA, :, Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, Prithvijit Chattopadhyay, Mike Chen, Yongxin Chen, Yu Chen, Shuai Cheng, Yin Cui, Jenna Diamond, Yifan Ding, Jiaojiao Fan, Linxi Fan, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Ruiyuan Ga...

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning 04.11.2025

🤗 Upvotes: 56 | cs. CV Authors: Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng Title: ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning Arxiv: http://arxiv.org/abs/2510.27492v1 Abstract: Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what co...

INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats 04.11.2025

🤗 Upvotes: 49 | cs. LG, cs. AI Authors: Mengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan, Jing Liu, Chaoyi Zhang, Yunshui Li, Jie Huang, Jin Ma, Zeyue Xue, Zhiheng Liu, Xingyan Bin, Ping Luo Title: INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats Arxiv: http://arxiv.org/abs/2510.25602v1 Abstract: Modern AI hardware, such as Nvidia's Blackwell architecture, is increasin...

Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning 04.11.2025

🤗 Upvotes: 21 | cs. CV, cs. AI Authors: Yuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao, Long Xing, Xiaoyi Dong, Haodong Duan, Dahua Lin, Jiaqi Wang Title: Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning Arxiv: http://arxiv.org/abs/2510.27606v1 Abstract: Spatial understanding remains a weakness of Large Vision-Language Models (LVLMs). Existing supervise...

The End of Manual Decoding: Towards Truly End-to-End Language Models 01.11.2025

🤗 Upvotes: 70 | cs. CL, cs. AI Authors: Zhichao Wang, Dongyang Ma, Xinting Huang, Deng Cai, Tian Lan, Jiahao Xu, Haitao Mi, Xiaoying Tang, Yan Wang Title: The End of Manual Decoding: Towards Truly End-to-End Language Models Arxiv: http://arxiv.org/abs/2510.26697v1 Abstract: The "end-to-end" label for LLMs is a misnomer. In practice, they depend on a non-differentiable decoding process that requir...

Kimi Linear: An Expressive, Efficient Attention Architecture 01.11.2025

🤗 Upvotes: 40 | cs. CL, cs. LG Authors: Kimi Team, Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, Chengyin Liu, Xin Men, Songlin Yang, Zhiyuan Li, Wentao Li, Enzhe Lu, Weizhou Liu, Yanru Chen, Weixin Xu, Longhui Yu, Yejie Wang, Yu Fan, Longguang Zhong, Enming Yuan, Dehao Zhang, Yizhi Zhang, T. Y. Liu, Haiming Wang, Shengjun Fang, Weiran He, Shaowei Liu, Yiwei Li, Jianlin Su, Jiezhon...

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents 01.11.2025

🤗 Upvotes: 28 | cs. AI Authors: Mathieu Andreux, Märt Bakler, Yanael Barbier, Hamza Benchekroun, Emilien Biré, Antoine Bonnet, Riaz Bordie, Nathan Bout, Matthias Brunel, Aleix Cambray, Pierre-Louis Cedoz, Antoine Chassang, Gautier Cloix, Ethan Connelly, Alexandra Constantinou, Ramzi De Coster, Hubert de la Jonquiere, Aurélien Delfosse, Maxime Delpit, Alexis Deprez, Augustin Derupti, Mathieu Diaz,...

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark 01.11.2025

🤗 Upvotes: 27 | cs. CV, cs. AI, cs. CL Authors: Ziyu Guo, Xinyan Chen, Renrui Zhang, Ruichuan An, Yu Qi, Dongzhi Jiang, Xiangtai Li, Manyuan Zhang, Hongsheng Li, Pheng-Ann Heng Title: Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark Arxiv: http://arxiv.org/abs/2510.26802v1 Abstract: Recent video generation models can produce high-fidelity, temporally co...

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation 01.11.2025

🤗 Upvotes: 25 | cs. CV Authors: Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang, Guorui Song, Ailing Zeng, Xian Liu, Chen Wei, Wanqi Yin, Qingping Sun, Zhongang Cai, Lei Yang, Ziwei Liu Title: The Quest for Generalizable Motion Generation: Data, Model, and Evaluation Arxiv: http://arxiv.org/abs/2510.26794v1 Abstract: Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks...

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations 29.10.2025

🤗 Upvotes: 147 | cs. CV Authors: Yujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao Title: Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations Arxiv: http://arxiv.org/abs/2510.23607v1 Abstract: Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a singl...

Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning 24.10.2025

🤗 Upvotes: 79 | cs. LG, cs. AI, cs. CL Authors: Ling Team, Bin Han, Caizhi Tang, Chen Liang, Donghao Zhang, Fan Yuan, Feng Zhu, Jie Gao, Jingyu Hu, Longfei Li, Meng Li, Mingyang Zhang, Peijie Jiang, Peng Jiao, Qian Zhao, Qingyuan Yang, Wenbo Shen, Xinxing Yang, Yalin Zhang, Yankun Ren, Yao Zhao, Yibo Cao, Yixuan Sun, Yue Zhang, Yuchen Fang, Zibin Lin, Zixuan Cheng, Jun Zhou Title: Every Attention...

BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping 24.10.2025

🤗 Upvotes: 68 | cs. LG, cs. AI, cs. CL Authors: Zhiheng Xi, Xin Guo, Yang Nan, Enyu Zhou, Junrui Shen, Wenxiang Chen, Jiaqi Liu, Jixuan Huang, Zhihao Zhang, Honglin Guo, Xun Deng, Zhikai Lei, Miao Zheng, Guoteng Wang, Shuo Zhang, Peng Sun, Rui Zheng, Hang Yan, Tao Gui, Qi Zhang, Xuanjing Huang Title: BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization wit...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.