Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Latent Implicit Visual Reasoning 27.12.2025

🤗 Upvotes: 34 | cs. CV Authors: Kelvin Li, Chuyi Shang, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig Title: Latent Implicit Visual Reasoning Arxiv: http://arxiv.org/abs/2512.21218v1 Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are limited in t...

Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning 27.12.2025

🤗 Upvotes: 26 | cs. LG, cs. AI Authors: Seijin Kobayashi, Yanick Schimpf, Maximilian Schlegel, Angelika Steger, Maciej Wolczyk, Johannes von Oswald, Nino Scherrer, Kaitlin Maile, Guillaume Lajoie, Blake A. Richards, Rif A. Saurous, James Manyika, Blaise Agüera y Arcas, Alexander Meulemans, João Sacramento Title: Emergent temporal abstractions in autoregressive models enable hierarchical reinforce...

TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times 26.12.2025

🤗 Upvotes: 51 | cs. CV, cs. AI, cs. LG Authors: Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang, Ion Stoica, Joseph E. Gonzalez, Jianfei Chen, Jun Zhu Title: TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times Arxiv: http://arxiv.org/abs/2512.16093v1 Abstract: We introduce TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generatio...

Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models 26.12.2025

🤗 Upvotes: 42 | cs. CV Authors: Shengchao Zhou, Yuxin Chen, Yuying Ge, Wei Huang, Jiehong Lin, Ying Shan, Xiaojuan Qi Title: Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models Arxiv: http://arxiv.org/abs/2512.20557v1 Abstract: Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the ev...

DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation 26.12.2025

🤗 Upvotes: 26 | cs. CV Authors: Jiawei Liu, Junqiao Li, Jiangfan Deng, Gen Li, Siyu Zhou, Zetao Fang, Shanshan Lao, Zengde Deng, Jianing Zhu, Tingting Ma, Jiayi Li, Yunqiu Wang, Qian He, Xinglong Wu Title: DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation Arxiv: http://arxiv.org/abs/2512.21252v1 Abstract: The "one-shot" technique represents a distinct and sophisticated aesthetic in fi...

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation 26.12.2025

🤗 Upvotes: 23 | cs. CV Authors: Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang, Yuanxing Zhang, Jialu Chen, Miao Deng, Jiahao Wang, Yubin Guo, Chenxi Liao, Yize Zhang, Zhaoxiang Zhang, Jiaheng Liu Title: T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation Arxiv: http://arxiv.org/abs/2512.21094v1 Abstract: Text-to-Audio-Video (T2AV) generation aims to synthesize temporally c...

SemanticGen: Video Generation in Semantic Space 25.12.2025

🤗 Upvotes: 78 | cs. CV Authors: Jianhong Bai, Xiaoshi Wu, Xintao Wang, Xiao Fu, Yuanxing Zhang, Qinghe Wang, Xiaoyu Shi, Menghan Xia, Zuozhu Liu, Haoji Hu, Pengfei Wan, Kun Gai Title: SemanticGen: Video Generation in Semantic Space Arxiv: http://arxiv.org/abs/2512.20619v2 Abstract: State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map...

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies 25.12.2025

🤗 Upvotes: 49 | cs. LG, cs. AI, cs. CL Authors: Yuqiao Tan, Minzheng Wang, Shizhu He, Huanxuan Liao, Chengfeng Zhao, Qiunan Lu, Tian Liang, Jun Zhao, Kang Liu Title: Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Arxiv: http://arxiv.org/abs/2512.19673v1 Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a...

LongVideoAgent: Multi-Agent Reasoning with Long Videos 25.12.2025

🤗 Upvotes: 38 | cs. AI, cs. CV, cs. LG, cs. MA Authors: Runtao Liu, Ziyi Liu, Jiaqi Tang, Yue Ma, Renjie Pi, Jipeng Zhang, Qifeng Chen Title: LongVideoAgent: Multi-Agent Reasoning with Long Videos Arxiv: http://arxiv.org/abs/2512.20618v1 Abstract: Recent advances in multimodal LLMs and systems that use tools for long-video QA point to the promise of reasoning over hour-long episodes. However, man...

SpatialTree: How Spatial Abilities Branch Out in MLLMs 25.12.2025

🤗 Upvotes: 35 | cs. CV Authors: Yuxi Xiao, Longfei Li, Shen Yan, Xinhang Liu, Sida Peng, Yunchao Wei, Xiaowei Zhou, Bingyi Kang Title: SpatialTree: How Spatial Abilities Branch Out in MLLMs Arxiv: http://arxiv.org/abs/2512.20617v1 Abstract: Cognitive science suggests that spatial ability develops progressively-from perception to reasoning and interaction. Yet in multimodal LLMs (MLLMs), this hier...

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI 24.12.2025

🤗 Upvotes: 159 | cs. LG, cs. CL Authors: Hao Liang, Xiaochen Ma, Zhou Liu, Zhen Hao Wong, Zhengyang Zhao, Zimo Meng, Runming He, Chengyu Shen, Qifeng Cai, Zhaoyang Han, Meiyi Qiang, Yalin Feng, Tianyi Bai, Zewei Pan, Ziyi Guo, Yizhen Jiang, Jingwen Deng, Qijie You, Peichao Lai, Tianyu Guo, Chi Hsu Tsai, Hengyi Feng, Rui Hu, Wenkai Yu, Junbo Niu, Bohan Zeng, Ruichuan An, Lu Ma, Jihao Huang, Yaowei...

The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding 24.12.2025

🤗 Upvotes: 53 | cs. CV Authors: Weichen Fan, Haiwen Diao, Quan Wang, Dahua Lin, Ziwei Liu Title: The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding Arxiv: http://arxiv.org/abs/2512.19693v1 Abstract: Deep representations across modalities are inherently intertwined. In this paper, we systematically analyze the spectral characteristics of various semantic...

Region-Constraint In-Context Generation for Instructional Video Editing 24.12.2025

🤗 Upvotes: 40 | cs. CV, cs. MM Authors: Zhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu, Wu Liu, Ting Yao, Tao Mei Title: Region-Constraint In-Context Generation for Instructional Video Editing Arxiv: http://arxiv.org/abs/2512.17650v1 Abstract: The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality....

QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation 24.12.2025

🤗 Upvotes: 26 | cs. CL, cs. IR Authors: Dehai Min, Kailin Zhang, Tongtong Wu, Lu Cheng Title: QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation Arxiv: http://arxiv.org/abs/2512.19134v1 Abstract: Dynamic Retrieval-Augmented Generation adaptively determines when to retrieve during generation to mitigate hallucinations in large language models...

Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation 24.12.2025

🤗 Upvotes: 26 | cs. CV Authors: Min-Jung Kim, Jeongho Kim, Hoiyeong Jin, Junha Hyung, Jaegul Choo Title: Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation Arxiv: http://arxiv.org/abs/2512.17040v1 Abstract: Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide cre...

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction 24.12.2025

🤗 Upvotes: 21 | cs. CL, cs. AI, cs. CY Authors: Ming Li, Han Chen, Yunze Xiao, Jian Chen, Hong Jiao, Tianyi Zhou Title: Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction Arxiv: http://arxiv.org/abs/2512.18880v1 Abstract: Accurate estimation of item (question or task) difficulty is critical for educational assessment but s...

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows 23.12.2025

🤗 Upvotes: 78 | cs. AI, cs. CL, cs. LG Authors: Wanghan Xu, Yuhao Zhou, Yifan Zhou, Qinglong Cao, Shuo Li, Jia Bu, Bo Liu, Yixin Chen, Xuming He, Xiangyu Zhao, Xiang Zhuang, Fengxiang Wang, Zhiwang Zhou, Qiantai Feng, Wenxuan Huang, Jiaqi Wei, Hao Wu, Yuejin Yang, Guangshuai Wang, Sheng Xu, Ziyan Huang, Xinyao Liu, Jiyao Liu, Cheng Tang, Wei Li, Ying Chen, Junzhi Ning, Pengfei Jiang, Chenglong Ma...

PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence 23.12.2025

🤗 Upvotes: 64 | cs. RO Authors: Xiaopeng Lin, Shijie Lian, Bin Yu, Ruoqi Yang, Changti Wu, Yuzhuo Miao, Yurun Jin, Yukun Shi, Cong Huang, Bojun Cheng, Kai Chen Title: PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence Arxiv: http://arxiv.org/abs/2512.16793v1 Abstract: Robotic generalization relies on physical intelligence: the ability to reason about...

When Reasoning Meets Its Laws 23.12.2025

🤗 Upvotes: 48 | cs. AI, cs. CL Authors: Junyu Zhang, Yifan Sun, Tianang Leng, Jingyan Shen, Liu Ziyin, Paul Pu Liang, Huan Zhang Title: When Reasoning Meets Its Laws Arxiv: http://arxiv.org/abs/2512.17901v1 Abstract: Despite the superior performance of Large Reasoning Models (LRMs), their reasoning behaviors are often counterintuitive, leading to suboptimal reasoning capabilities. To theoreticall...

Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience 23.12.2025

🤗 Upvotes: 40 | cs. CL Authors: Jiangjie Chen, Wenxiang Chen, Jiacheng Du, Jinyi Hu, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Wenlei Shi, Zhihong Wang, Mingxuan Wang, Chenrui Wei, Shufa Wei, Huajian Xin, Fan Yang, Weihao Gao, Zheng Yuan, Tianyang Zhan, Zeyu Zheng, Tianxi Zhou, Thomas Hanwen Zhu Title: Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learni...

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation 23.12.2025

🤗 Upvotes: 31 | cs. CV Authors: Chiao-An Yang, Ryo Hachiuma, Sifei Liu, Subhashree Radhakrishnan, Raymond A. Yeh, Yu-Chiang Frank Wang, Min-Hung Chen Title: 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation Arxiv: http://arxiv.org/abs/2512.17012v1 Abstract: Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains...

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing 23.12.2025

🤗 Upvotes: 30 | cs. CV Authors: Shilong Zhang, He Zhang, Zhifei Zhang, Chongjian Ge, Shuchen Xue, Shaoteng Liu, Mengwei Ren, Soo Ye Kim, Yuqian Zhou, Qing Liu, Daniil Pakhomov, Kai Zhang, Zhe Lin, Ping Luo Title: Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing Arxiv: http://arxiv.org/abs/2512.17909v1 Abstract: Modern Latent D...

Are We on the Right Way to Assessing LLM-as-a-Judge? 23.12.2025

🤗 Upvotes: 24 | cs. CL, cs. AI Authors: Yuanning Feng, Sinan Wang, Zhengxiang Cheng, Yao Wan, Dongping Chen Title: Are We on the Right Way to Assessing LLM-as-a-Judge? Arxiv: http://arxiv.org/abs/2512.16041v1 Abstract: LLM-as-a-Judge has been widely adopted as an evaluation method and served as supervised rewards in model training. However, existing benchmarks for LLM-as-a-Judge are mainly relyin...

Kling-Omni Technical Report 20.12.2025

🤗 Upvotes: 112 | cs. CV Authors: Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, Zipeng Feng, Kun Gai, Sainan Guo, Feng Han, Jingbin He, Kang He, Xiao Hu, Xiaohua Hu, Boyuan Jiang, Fangyuan Kong, Hang Li, Jie Li, Qingyu Li, Shen Li, Xiaohan Li, Yan Li, Jiajun Liang, Borui Liao, Yiqiao Liao, Weihong Lin, Quande Liu, Xiaokun Liu, Yilun Liu, Yuliang Liu, Shun Lu, Hangyu Mao, Yunyao Mao, Haodong Ou...

Adaptation of Agentic AI 20.12.2025

🤗 Upvotes: 59 | cs. AI, cs. CL Authors: Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi, Zifeng Wang, Luxi He, Yichen Wu, Ming Zhong, Peiyang Song, Qizheng Zhang, Heng Wang, Xueqiang Xu, Hanwen Xu, Pengrui Han, Dylan Zhang, Jiashuo Sun, Chaoqi Yang, Kun Qian, Tian Wang, Changran Hu, Manling Li, Quanzheng Li, Hao Peng, Sheng Wang, Jingbo Shang, Chao Zhang, Jiaxuan You, Liyuan Liu, Pan Lu, Yu Zhang, Heng...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.