Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos 06.01.2026 22:46
🤗 Upvotes: 86 | cs. CV Authors: Yuxue Yang, Lue Fan, Ziqi Shi, Junran Peng, Feng Wang, Zhaoxiang Zhang Title: NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos Arxiv: http://arxiv.org/abs/2601.00393v1 Abstract: In this paper, we propose NeoVerse, a versatile 4D world model that is capable of 4D reconstruction, novel-trajectory video generation, and rich downstream applications....
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation 06.01.2026 22:58
🤗 Upvotes: 44 | cs. LG, cs. AI, cs. CV, cs. HC, cs. MM Authors: Taekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, Sung Ju Hwang Title: Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation Arxiv: http://arxiv.org/abs/2601.00664v1 Abstract: Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. How...
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation 06.01.2026 26:51
🤗 Upvotes: 38 | cs. CV, cs. AI Authors: Zhe Huang, Hao Wen, Aiming Hao, Bingze Song, Meiqi Wu, Jiahong Wu, Xiangxiang Chu, Sheng Lu, Haoqian Wang Title: Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation Arxiv: http://arxiv.org/abs/2512.24271v1 Abstract: Multimodal Large Language Models (MLLMs) have made remarkable progress in video understanding. Howev...
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning 06.01.2026 27:56
🤗 Upvotes: 30 | cs. CV Authors: Yong Xien Chng, Tao Hu, Wenwen Tong, Xueheng Li, Jiandong Chen, Haojia Yu, Jiefan Lu, Hewei Guo, Hanming Deng, Chengjun Xie, Gao Huang, Dahua Lin, Lewei Lu Title: SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning Arxiv: http://arxiv.org/abs/2512.24330v1 Abstract: While Vision-Language Models (VLMs) can solve complex tasks...
Deep Delta Learning 06.01.2026 20:34
🤗 Upvotes: 22 | cs. LG, cs. AI, cs. CL, cs. CV Authors: Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu Title: Deep Delta Learning Arxiv: http://arxiv.org/abs/2601.00417v1 Abstract: The efficacy of deep residual networks is fundamentally predicated on the identity shortcut connection. While this mechanism effectively mitigates the vanishing gradient problem, it imposes a strictly additive induc...
AdaGaR: Adaptive Gabor Representation for Dynamic Scene Reconstruction 06.01.2026 23:09
🤗 Upvotes: 22 | cs. CV Authors: Jiewen Chan, Zhenjun Zhao, Yu-Lun Liu Title: AdaGaR: Adaptive Gabor Representation for Dynamic Scene Reconstruction Arxiv: http://arxiv.org/abs/2601.00796v1 Abstract: Reconstructing dynamic 3D scenes from monocular videos requires simultaneously capturing high-frequency appearance details and temporally continuous motion. Existing methods using single Gaussian prim...
Nested Learning: The Illusion of Deep Learning Architectures 06.01.2026 23:45
🤗 Upvotes: 22 | cs. LG, cs. AI Authors: Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni Title: Nested Learning: The Illusion of Deep Learning Architectures Arxiv: http://arxiv.org/abs/2512.24695v1 Abstract: Despite the recent progresses, particularly in developing Language Models, there are fundamental challenges and unanswered questions about how such models can continually learn/me...
Improving Multi-step RAG with Hypergraph-based Memory for Long-Context Complex Relational Modeling 03.01.2026 22:36
🤗 Upvotes: 44 | cs. CL, cs. AI, cs. LG Authors: Chulun Zhou, Chunkang Zhang, Guoxin Yu, Fandong Meng, Jie Zhou, Wai Lam, Mo Yu Title: Improving Multi-step RAG with Hypergraph-based Memory for Long-Context Complex Relational Modeling Arxiv: http://arxiv.org/abs/2512.23959v1 Abstract: Multi-step retrieval-augmented generation (RAG) has become a widely adopted strategy for enhancing large language m...
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space 03.01.2026 25:22
🤗 Upvotes: 25 | cs. LG, cs. AI Authors: Xingwei Qu, Shaowen Wang, Zihao Huang, Kai Hua, Fan Yin, Rui-Jie Zhu, Jundong Zhou, Qiyang Min, Zihao Wang, Yizhi Li, Tianyu Zhang, He Xing, Zheng Zhang, Yuxuan Song, Tianyu Zheng, Zhiyuan Zeng, Chenghua Lin, Ge Zhang, Wenhao Huang Title: Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space Arxiv: http://arxiv.org/abs/2512.24617v1 Ab...
mHC: Manifold-Constrained Hyper-Connections 02.01.2026 20:57
🤗 Upvotes: 73 | cs. CL, cs. AI, cs. LG Authors: Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Liang Zhao, Shangyan Zhou, Zhean Xu, Zhengyan Zhang, Wangding Zeng, Shengding Hu, Yuqing Wang, Jingyang Yuan, Lean Wang, Wenfeng Liang Title: mHC: Manifold-Constrained Hyper-Connections Arxiv: http://arxiv.org/abs/2512.24880v1 Abstract: R...
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models 02.01.2026 28:35
🤗 Upvotes: 45 | cs. CL Authors: Junru Lu, Jiarui Qin, Lingfeng Qiao, Yinghui Li, Xinyi Dai, Bo Ke, Jianfeng He, Ruizhi Qiao, Di Yin, Xing Sun, Yunsheng Wu, Yinsong Liu, Shuangyin Liu, Mingkong Tang, Haodong Lin, Jiayi Kuang, Fanxu Meng, Xiaojuan Tang, Yunjia Xi, Junjie Huang, Haotong Yang, Zhenyi Shen, Yangning Li, Qianwen Zhang, Yifei Yu, Siyu An, Junnan Dong, Qiufeng Wang, Jie Wang, Keyu Chen,...
Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem 02.01.2026 25:58
🤗 Upvotes: 33 | cs. AI, cs. CL Authors: Weixun Wang, XiaoXiao Xu, Wanhe An, Fangwen Dai, Wei Gao, Yancheng He, Ju Huang, Qiang Ji, Hanqi Jin, Xiaoyang Li, Yang Li, Zhongwen Li, Shirong Lin, Jiashun Liu, Zenan Liu, Tao Luo, Dilxat Muhtar, Yuanbin Qu, Jiaqiang Shi, Qinghui Sun, Yingshui Tan, Hao Tang, Runze Wang, Yi Wang, Zhaoguo Wang, Yanan Wu, Shaopan Xiong, Binchen Xu, Xander Xu, Yuchi Xu, Qipen...
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction 02.01.2026 22:28
🤗 Upvotes: 22 | cs. CV Authors: Yi-Chuan Huang, Hao-Jen Chien, Chin-Yang Lin, Ying-Huan Chen, Yu-Lun Liu Title: GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction Arxiv: http://arxiv.org/abs/2512.25073v1 Abstract: Recent advances in 3D reconstruction have achieved remarkable progress in high-quality scene capture from dense multi-view imagery, yet struggle whe...
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss 31.12.2025 24:49
🤗 Upvotes: 72 | cs. CL, cs. LG Authors: Ang Lv, Jin Ma, Yiyuan Ma, Siyuan Qiao Title: Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss Arxiv: http://arxiv.org/abs/2512.23447v1 Abstract: Mixture-of-Experts (MoE) models lack explicit constraints to ensure the router's decisions align well with the experts' capabilities, which ultimately limits model performance. To address t...
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation 31.12.2025 23:16
🤗 Upvotes: 51 | cs. CV Authors: Ethan Chern, Zhulin Hu, Bohao Tang, Jiadi Su, Steffi Chern, Zhijie Deng, Pengfei Liu Title: LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation Arxiv: http://arxiv.org/abs/2512.23576v1 Abstract: Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, th...
Yume-1.5: A Text-Controlled Interactive World Generation Model 31.12.2025 25:01
🤗 Upvotes: 50 | cs. CV Authors: Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, Kaining Ying, Tong He, Jiangmiao Pang, Yu Qiao, Kaipeng Zhang Title: Yume-1.5: A Text-Controlled Interactive World Generation Model Arxiv: http://arxiv.org/abs/2512.22096v1 Abstract: Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most o...
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 31.12.2025 24:01
🤗 Upvotes: 33 | cs. CL, cs. AI, cs. CV, cs. LG, cs. MA Authors: Shaofei Cai, Yulei Qin, Haojia Lin, Zihan Xu, Gang Li, Yuchen Shi, Zongyi Li, Yong Mao, Siqi Cai, Xiaoyu Tan, Yitao Liang, Ke Li, Xing Sun Title: SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents Arxiv: http://arxiv.org/abs/2512.22322v1 Abstract: Agentic reinforcement learning (RL) holds great promise for the developmen...
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation 31.12.2025 25:32
🤗 Upvotes: 32 | cs. CV Authors: Shaocong Xu, Songlin Wei, Qizhe Wei, Zheng Geng, Hong Li, Licheng Shen, Qianpu Sun, Shu Han, Bin Ma, Bohan Li, Chongjie Ye, Yuhang Zheng, Nan Wang, Saining Zhang, Hao Zhao Title: Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation Arxiv: http://arxiv.org/abs/2512.23705v1 Abstract: Transparent objects remain n...
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion 31.12.2025 25:06
🤗 Upvotes: 30 | cs. CV Authors: Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang, Chi-Wei Hsiao, Po-Fan Yu, Yu-Chih Chen, Yu-Lun Liu Title: Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion Arxiv: http://arxiv.org/abs/2512.23709v1 Abstract: Diffusion-based video super-resolution (VSR) methods achieve strong perceptual quality but remain impractical for laten...
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone 31.12.2025 23:48
🤗 Upvotes: 28 | cs. CV, cs. CL Authors: Jiacheng Ye, Shansan Gong, Jiahui Gao, Junming Fan, Shuang Wu, Wei Bi, Haoli Bai, Lifeng Shang, Lingpeng Kong Title: Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone Arxiv: http://arxiv.org/abs/2512.22615v1 Abstract: While autoregressive Large Vision-Language Models (VLMs) have achieved...
SpotEdit: Selective Region Editing in Diffusion Transformers 31.12.2025 22:44
🤗 Upvotes: 27 | cs. CV, cs. AI Authors: Zhibin Qin, Zhenxiong Tan, Zeqing Wang, Songhua Liu, Xinchao Wang Title: SpotEdit: Selective Region Editing in Diffusion Transformers Arxiv: http://arxiv.org/abs/2512.22323v1 Abstract: Diffusion Transformer models have significantly advanced image editing by encoding conditional images and integrating them into transformer layers. However, most edits involv...
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 31.12.2025 22:03
🤗 Upvotes: 21 | cs. CV Authors: Bozhou Li, Sihan Yang, Yushuo Guan, Ruichuan An, Xinlong Chen, Yang Shi, Pengfei Wan, Wentao Zhang, Yuanxing zhang Title: GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models Arxiv: http://arxiv.org/abs/2512.15560v2 Abstract: The text encoder is a critical component of text-to-image and text-to-video diffusion models, fundamentally...
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion 30.12.2025 23:11
🤗 Upvotes: 74 | cs. CV, cs. AI Authors: Hoiyeong Jin, Hyojin Jang, Jeongho Kim, Junha Hyung, Kinam Kim, Dongjin Kim, Huijin Choi, Hyeonji Kim, Jaegul Choo Title: InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion Arxiv: http://arxiv.org/abs/2512.17504v1 Abstract: Recent advances in diffusion-based video generation have opened new possibilities for...
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding 30.12.2025 21:17
🤗 Upvotes: 70 | cs. CL Authors: Yuqing Li, Jiangnan Li, Zheng Lin, Ziyan Zhou, Junjie Wu, Weiping Wang, Jie Zhou, Mo Yu Title: Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding Arxiv: http://arxiv.org/abs/2512.17220v1 Abstract: Humans understand long and complex texts by relying on a holistic semantic representation of the content. This global view helps organ...
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents 30.12.2025 24:59
🤗 Upvotes: 21 | cs. CV Authors: Hanzhang Zhou, Xu Zhang, Panrong Tong, Jianan Zhang, Liangyu Chen, Quyu Kong, Chenglin Cai, Chen Liu, Yue Wang, Jingren Zhou, Steven Hoi Title: MAI-UI Technical Report: Real-World Centric Foundation GUI Agents Arxiv: http://arxiv.org/abs/2512.22047v1 Abstract: The development of GUI agents could revolutionize the next generation of human-computer interaction. Motiv...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.