Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing 21.03.2026

🤗 Upvotes: 56 | cs. CV Authors: Xinyao Zhang, Wenkai Dong, Yuxin Song, Bo Fang, Qi Zhang, Jing Wang, Fan Chen, Hui Zhang, Haocheng Feng, Yu Lu, Hang Zhou, Chun Yuan, Jingdong Wang Title: SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing Arxiv: http://arxiv.org/abs/2603.19228v1 Abstract: Current instruction-guided video editing models struggle to simulta...

FASTER: Rethinking Real-Time Flow VLAs 21.03.2026

🤗 Upvotes: 41 | cs. RO, cs. CV Authors: Yuxiang Lu, Zhe Liu, Xianzhe Fan, Zhenya Yang, Jinghua Hou, Junyi Li, Kaixin Ding, Hengshuang Zhao Title: FASTER: Rethinking Real-Time Flow VLAs Arxiv: http://arxiv.org/abs/2603.19199v1 Abstract: Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize...

3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model 21.03.2026

🤗 Upvotes: 40 | cs. CV Authors: Hyun-kyu Ko, Jihyeon Park, Younghyun Kim, Dongheok Park, Eunbyung Park Title: 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model Arxiv: http://arxiv.org/abs/2603.18524v1 Abstract: Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual produ...

Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer 21.03.2026

🤗 Upvotes: 34 | cs. CV Authors: Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, Ziwei Liu Title: Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer Arxiv: http://arxiv.org/abs/2603.19227v1 Abstract: Prior motion generation largely follows two paradigms: continuous diffusion models that excel at kinematic control, and discrete token-based gen...

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction 21.03.2026

🤗 Upvotes: 28 | cs. CV Authors: Haitian Li, Haozhe Xie, Junxiang Xu, Beichen Wen, Fangzhou Hong, Ziwei Liu Title: MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction Arxiv: http://arxiv.org/abs/2603.19231v1 Abstract: Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from lim...

Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 21.03.2026

🤗 Upvotes: 28 | cs. CL, cs. AI, cs. LG Authors: Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, Renjie Pi, Grace Lam, Nayeon Lee, Alexander Bukharin, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping Title: Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Arxiv: http://arxiv.o...

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens 21.03.2026

🤗 Upvotes: 26 | cs. CV Authors: Yuqing Wang, Chuofan Ma, Zhijie Lin, Yao Teng, Lijun Yu, Shuai Wang, Jiaming Han, Jiashi Feng, Yi Jiang, Xihui Liu Title: Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens Arxiv: http://arxiv.org/abs/2603.19232v1 Abstract: Visual generation with discrete tokens has gained significant attention as it enables a unified tok...

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs 21.03.2026

🤗 Upvotes: 26 | cs. CV Authors: Keda Tao, Yuhua Zheng, Jia Xu, Wenjie Du, Kele Shao, Hesong Wang, Xueyi Chen, Xin Jin, Junhan Zhu, Bohan Yu, Weiqiang Wang, Jian Liu, Can Qin, Yulun Zhang, Ming-Hsuan Yang, Huan Wang Title: LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs Arxiv: http://arxiv.org/abs/2603.19217v1 Abstract: Recent advancements in omnimodal large la...

Memento-Skills: Let Agents Design Agents 21.03.2026

🤗 Upvotes: 24 | cs. AI, cs. CL, cs. LG Authors: Huichi Zhou, Siyuan Guo, Anjie Liu, Zhongwei Yu, Ziqin Gong, Bowen Zhao, Zhixun Chen, Menglong Zhang, Yihang Chen, Jinsong Li, Runyu Yang, Qiangbin Liu, Xinlei Yu, Jianmin Zhou, Na Wang, Chunyang Sun, Jun Wang Title: Memento-Skills: Let Agents Design Agents Arxiv: http://arxiv.org/abs/2603.18743v1 Abstract: We introduce \emph{Memento-Skills}, a gene...

MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild 20.03.2026

🤗 Upvotes: 97 | cs. LG Authors: Peng Xia, Jianwen Chen, Xinyu Yang, Haoqin Tu, Jiaqi Liu, Kaiwen Xiong, Siwei Han, Shi Qiu, Haonian Ji, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Huaxiu Yao Title: MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild Arxiv: http://arxiv.org/abs/2603.17187v1 Abstract: Large language model (LLM) agents are increasingly used for complex tasks, yet deploy...

Video-CoE: Reinforcing Video Event Prediction via Chain of Events 20.03.2026

🤗 Upvotes: 85 | cs. CV Authors: Qile Su, Jing Tang, Rui Chen, Lei Sun, Xiangxiang Chu Title: Video-CoE: Reinforcing Video Event Prediction via Chain of Events Arxiv: http://arxiv.org/abs/2603.14935v1 Abstract: Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal...

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models 20.03.2026

🤗 Upvotes: 72 | cs. CV Authors: Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg Title: MosaicMem: Hybrid Spatial Memory for Controllable Video World Models Arxiv: http://arxiv.org/abs/2603.17117v1 Abstract: Video diffusion models are moving beyond short, plausible clips toward world simulators that...

Alignment Makes Language Models Normative, Not Descriptive 20.03.2026

🤗 Upvotes: 36 | cs. CL, cs. AI, cs. GT Authors: Eilam Shapira, Moshe Tennenholtz, Roi Reichart Title: Alignment Makes Language Models Normative, Not Descriptive Arxiv: http://arxiv.org/abs/2603.17218v1 Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed human behavior. We compare 120 base-aligned...

Complementary Reinforcement Learning 20.03.2026

🤗 Upvotes: 28 | cs. LG, cs. CL Authors: Dilxat Muhtar, Jiashun Liu, Wei Gao, Weixun Wang, Shaopan Xiong, Ju Huang, Siran Yang, Wenbo Su, Jiamang Wang, Ling Pan, Bo Zheng Title: Complementary Reinforcement Learning Arxiv: http://arxiv.org/abs/2603.17621v1 Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for training LLM-based agents, yet remains limited by low sample effici...

When AI Navigates the Fog of War 20.03.2026

🤗 Upvotes: 22 | cs. AI, cs. CL, cs. CY Authors: Ming Li, Xirui Li, Tianyi Zhou Title: When AI Navigates the Fog of War Arxiv: http://arxiv.org/abs/2603.16642v1 Abstract: Can AI reason about a war before its trajectory becomes historically obvious? Analyzing this capability is difficult because retrospective geopolitical prediction is heavily confounded by training-data leakage. We address this ch...

MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification 19.03.2026

🤗 Upvotes: 150 | cs. CL, cs. AI, cs. IR, cs. LG Authors: MiroMind Team, S. Bai, L. Bing, L. Lei, R. Li, X. Li, X. Lin, E. Min, L. Su, B. Wang, L. Wang, L. Wang, S. Wang, X. Wang, Y. Zhang, Z. Zhang, G. Chen, L. Chen, Z. Cheng, Y. Deng, Z. Huang, D. Ng, J. Ni, Q. Ren, X. Tang, B. L. Wang, H. Wang, N. Wang, C. Wei, Q. Wu, J. Xia, Y. Xiao, H. Xu, X. Xu, C. Xue, Z. Yang, Z. Yang, F. Ye, H. Ye, J. Yu,...

InCoder-32B: Code Foundation Model for Industrial Scenarios 19.03.2026

🤗 Upvotes: 144 | cs. SE, cs. AI Authors: Jian Yang, Wei Zhang, Jiajun Wu, Junhang Cheng, Shawn Guo, Haowen Wang, Weicheng Gu, Yaxin Du, Joseph Li, Fanglin Xu, Yizhi Li, Lin Jing, Yuanbo Wang, Yuhan Gao, Ruihao Gong, Chuan Hao, Ran Tao, Aishan Liu, Tuney Zheng, Ganqu Cui, Zhoujun Li, Mingjie Tang, Chenghua Lin, Wayne Xin Zhao, Xianglong Liu, Ming Zhou, Bryan Dai, Weifeng Lv Title: InCoder-32B: Cod...

Qianfan-OCR: A Unified End-to-End Model for Document Intelligence 19.03.2026

🤗 Upvotes: 118 | cs. CV Authors: Daxiang Dong, Mingming Zheng, Dong Xu, Chunhua Luo, Bairong Zhuang, Yuxuan Li, Ruoyun He, Haoran Wang, Wenyu Zhang, Wenbo Wang, Yicheng Wang, Xue Xiong, Ayong Zheng, Xiaoying Zuo, Ziwei Ou, Jingnan Gu, Quanhao Guo, Jianmin Wu, Dawei Yin, Dou Shen Title: Qianfan-OCR: A Unified End-to-End Model for Document Intelligence Arxiv: http://arxiv.org/abs/2603.13398v1 Abstr...

Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding 19.03.2026

🤗 Upvotes: 83 | cs. CV, cs. AI Authors: Zhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi, Feilong Tang, Ming Hu, Shiyan Su, Xiaocheng Zou, Wei Feng, Dwarikanath Mahapatra, Yifan Peng, Mingquan Lin, Zongyuan Ge Title: Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding Arxiv: http://arxiv.org/abs/2603.13366v1 Abstract: Recent advancements in multimodal...

Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation 19.03.2026

🤗 Upvotes: 64 | cs. RO, cs. CV Authors: Mutian Xu, Tianbao Zhang, Tianqi Liu, Zhaoxi Chen, Xiaoguang Han, Ziwei Liu Title: Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation Arxiv: http://arxiv.org/abs/2603.16669v1 Abstract: Simulating robot-world interactions is a cornerstone of Embodied AI. Recently, a few works have shown promise in leveraging video generations to tra...

Demystifing Video Reasoning 19.03.2026

🤗 Upvotes: 58 | cs. CV, cs. AI Authors: Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, Ziqi Huang, Hokin Deng, Dahua Lin, Ziwei Liu, Lei Yang Title: Demystifing Video Reasoning Arxiv: http://arxiv.org/abs/2603.16870v1 Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit...

WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation 19.03.2026

🤗 Upvotes: 49 | cs. CV Authors: Jisu Nam, Yicong Hong, Chun-Hao Paul Huang, Feng Liu, JoungBin Lee, Jiyoung Kim, Siyoon Jin, Yunsung Lee, Jaeyoon Jung, Suhwan Choi, Seungryong Kim, Yang Zhou Title: WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation Arxiv: http://arxiv.org/abs/2603.16871v1 Abstract: Recent advances in video diffusion trans...

TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas 19.03.2026

🤗 Upvotes: 44 | cs. AI Authors: Ai Jian, Xiaoyun Zhang, Wanrou Du, Jingqing Ruan, Jiangbo Pei, Weipeng Zhang, Ke Zeng, Xunliang Cai Title: TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas Arxiv: http://arxiv.org/abs/2603.16448v2 Abstract: Text-to-SQL parsing has achieved remarkable progress under the Full Schema Assumption. However, this premise fa...

Online Experiential Learning for Language Models 19.03.2026

🤗 Upvotes: 39 | cs. CL Authors: Tianzhu Ye, Li Dong, Qingxiu Dong, Xun Wu, Shaohan Huang, Furu Wei Title: Online Experiential Learning for Language Models Arxiv: http://arxiv.org/abs/2603.16856v1 Abstract: The prevailing paradigm for improving large language models relies on offline training with human annotations or simulated environments, leaving the rich experience accumulated during real-worl...

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use 19.03.2026

🤗 Upvotes: 32 | cs. AI Authors: Jiaxuan Lu, Kong Wang, Yemin Wang, Qingmei Tang, Hongwei Zeng, Xiang Chen, Jiahao Pi, Shujian Deng, Lingzhi Chen, Yi Fu, Kehua Yang, Xiao Sun Title: FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Arxiv: http://arxiv.org/abs/2603.08262v1 Abstract: The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm s...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.