Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics 17.12.2025 22:42
🤗 Upvotes: 34 | cs. LG Authors: Jingdi Lei, Di Zhang, Soujanya Poria Title: Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics Arxiv: http://arxiv.org/abs/2512.12602v1 Abstract: Linear-time attention and State Space Models (SSMs) promise to solve the quadratic cost bottleneck in long-context language models employing softmax attention. We introduce Error-Fre...
KlingAvatar 2.0 Technical Report 17.12.2025 24:12
🤗 Upvotes: 31 | cs. CV Authors: Kling Team, Jialu Chen, Yikang Ding, Zhixue Fang, Kun Gai, Yuan Gao, Kang He, Jingyun Hua, Boyuan Jiang, Mingming Lao, Xiaohan Li, Hui Liu, Jiwen Liu, Xiaoqiang Liu, Yuan Liu, Shun Lu, Yongsen Mao, Yingchao Shao, Huafeng Shi, Xiaoyu Shi, Peiqin Sun, Songlin Tang, Pengfei Wan, Chao Wang, Xuebo Wang, Haoxian Zhang, Yuanxing Zhang, Yan Zhou Title: KlingAvatar 2.0 Tech...
MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment 17.12.2025 23:46
🤗 Upvotes: 22 | cs. CL Authors: Mengxi Xiao, Kailai Yang, Pengde Zhao, Enze Zhang, Ziyan Kuang, Zhiwei Liu, Weiguang Han, Shu Liao, Lianting Huang, Jinpeng Hu, Min Peng, Qianqian Xie, Sophia Ananiadou Title: MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment Arxiv: http://arxiv.org/abs/2512.09636v2 Abstract: Mental health disorders affect hundreds of milli...
EgoX: Egocentric Video Generation from a Single Exocentric Video 16.12.2025 21:54
🤗 Upvotes: 48 | cs. CV Authors: Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park, Junha Hyung, Jaegul Choo Title: EgoX: Egocentric Video Generation from a Single Exocentric Video Arxiv: http://arxiv.org/abs/2512.08269v1 Abstract: Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into ego...
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry 16.12.2025 18:26
🤗 Upvotes: 38 | cs. CV, cs. AI, cs. CL Authors: Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng, Yanchao Li, Jingyi Liang, Junying Chen, Yunjin Yang, Jiajun You, Shuzhi Deng, Tongfei Wang, Wanting Chen, Chunxiu Hao, Ruiqi Xie, Zhenwei Wen, Xiangyi Feng, Zou Ting, Jin Zou Lin, Jianquan Li, Guangjun Yu, Liangyi Chen, Junwen Wang, Shan Jiang, Benyou Wang Title: DentalGPT: Incentivizing Multimoda...
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder 16.12.2025 22:17
🤗 Upvotes: 33 | cs. CV Authors: Minglei Shi, Haolin Wang, Borui Zhang, Wenzhao Zheng, Bohan Zeng, Ziyang Yuan, Xiaoshi Wu, Yuanxing Zhang, Huan Yang, Xintao Wang, Pengfei Wan, Kun Gai, Jie Zhou, Jiwen Lu Title: SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder Arxiv: http://arxiv.org/abs/2512.11749v1 Abstract: Visual generation grounded in Visual Foundation...
V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties 16.12.2025 23:26
🤗 Upvotes: 25 | cs. CV Authors: Ye Fang, Tong Wu, Valentin Deschaintre, Duygu Ceylan, Iliyan Georgiev, Chun-Hao Paul Huang, Yiwei Hu, Xuelin Chen, Tuanfeng Yang Wang Title: V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties Arxiv: http://arxiv.org/abs/2512.11799v1 Abstract: Large-scale video generation models have shown remarkable potential in modeling photorealistic appearanc...
T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground 13.12.2025 22:52
🤗 Upvotes: 60 | cs. CL Authors: Dmitrii Stoianov, Danil Taranets, Olga Tsymboi, Ramil Latypov, Almaz Dautov, Vladislav Kruglikov, Nikita Surkov, German Abramov, Pavel Gein, Dmitry Abulkhanov, Mikhail Gashkov, Viktor Zelenkovskiy, Artem Batalov, Aleksandr Medvedev, Anatolii Potapov Title: T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground Arxiv: http://arxiv.org/abs/2512.10430v1...
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving 13.12.2025 23:06
🤗 Upvotes: 37 | cs. CL, cs. AI Authors: Songyang Gao, Yuzhe Gu, Zijian Wu, Lingkai Kong, Wenwei Zhang, Zhongrui Cai, Fan Zheng, Tianyou Ma, Junhao Shen, Haiteng Zhao, Duanyang Zhang, Huilun Zhang, Kuikun Liu, Chengqi Lyu, Yanhui Duan, Chiyu Chen, Ningsheng Ma, Jianfei Gao, Han Lyu, Dahua Lin, Kai Chen Title: Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving Arxiv: http:...
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation 13.12.2025 28:53
🤗 Upvotes: 36 | cs. CV, cs. AI, cs. CL Authors: Yiwen Tang, Zoey Guo, Kaixin Zhu, Ray Zhang, Qizhi Chen, Dongzhi Jiang, Junli Liu, Bohan Zeng, Haoming Song, Delin Qu, Tianyi Bai, Dan Xu, Wentao Zhang, Bin Zhao Title: Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation Arxiv: http://arxiv.org/abs/2512.10949v1 Abstract: Reinforcement learning (RL), earlier proven to be effecti...
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification 13.12.2025 24:56
🤗 Upvotes: 31 | cs. CL, cs. LG Authors: Zijian Wu, Lingkai Kong, Wenwei Zhang, Songyang Gao, Yuzhe Gu, Zhongrui Cai, Tianyou Ma, Yuhong Liu, Zhi Wang, Runyuan Ma, Guangyu Wang, Wei Li, Conghui He, Dahua Lin, Kai Chen Title: OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification Arxiv: http://arxiv.org/abs/2512.10756v1 Abstract: Large language models (LLMs) have achie...
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning 13.12.2025 23:41
🤗 Upvotes: 25 | cs. AI Authors: Haiteng Zhao, Junhao Shen, Yiming Zhang, Songyang Gao, Kuikun Liu, Tianyou Ma, Fan Zheng, Dahua Lin, Wenwei Zhang, Kai Chen Title: Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning Arxiv: http://arxiv.org/abs/2512.10534v1 Abstract: Large language model (LLM) agents exhibit strong mathematical problem-solving...
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation 12.12.2025 21:49
🤗 Upvotes: 48 | cs. CV Authors: Ke Xing, Xiaojie Jin, Longfei Li, Yuyang Yin, Hanwen Liang, Guixun Luo, Chen Fang, Jue Wang, Konstantinos N. Plataniotis, Yao Zhao, Yunchao Wei Title: StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation Arxiv: http://arxiv.org/abs/2512.09363v2 Abstract: The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its p...
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain 12.12.2025 22:44
🤗 Upvotes: 32 | cs. CV Authors: Navve Wasserman, Matias Cosarinsky, Yuval Golbari, Aude Oliva, Antonio Torralba, Tamar Rott Shaham, Michal Irani Title: BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain Arxiv: http://arxiv.org/abs/2512.08560v1 Abstract: Understanding how the human brain represents visual concepts, and in which brain regions these repres...
OmniPSD: Layered PSD Generation with Diffusion Transformer 12.12.2025 26:08
🤗 Upvotes: 27 | cs. CV Authors: Cheng Liu, Yiren Song, Haofan Wang, Mike Zheng Shou Title: OmniPSD: Layered PSD Generation with Diffusion Transformer Arxiv: http://arxiv.org/abs/2512.09247v1 Abstract: Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains highly challenging....
Composing Concepts from Images and Videos via Concept-prompt Binding 12.12.2025 23:04
🤗 Upvotes: 24 | cs. CV, cs. AI, cs. MM Authors: Xianghao Kong, Zeyu Zhang, Yuwei Guo, Zhuoran Zhao, Songchun Zhang, Anyi Rao Title: Composing Concepts from Images and Videos via Concept-prompt Binding Arxiv: http://arxiv.org/abs/2512.09824v1 Abstract: Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls sh...
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 11.12.2025 23:28
🤗 Upvotes: 94 | cs. CV Authors: Ruihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang, Xiaogang Xu, Bin Xia, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu, Yingya Zhang, Yujiu Yang Title: Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance Arxiv: http://arxiv.org/abs/2512.08765v1 Abstract: We present Wan-Move, a simple and scalable framework that brings motion...
Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform 11.12.2025 25:38
🤗 Upvotes: 66 | cs. CV, cs. AI, cs. GR Authors: Yuning Gong, Yifei Liu, Yifan Zhan, Muyao Niu, Xueying Li, Yuanjun Liao, Jiaming Chen, Yuanyuan Gao, Jiaqi Chen, Minming Chen, Li Zhou, Yuning Zhang, Wei Wang, Xiaoqing Hou, Huaxi Huang, Shixiang Tang, Le Ma, Dingwen Zhang, Xue Yang, Junchi Yan, Yanchi Zhang, Yinqiang Zheng, Xiao Sun, Zhihang Zhong Title: Visionary: The World Model Carrier Built on...
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality 11.12.2025 23:04
🤗 Upvotes: 40 | cs. CV Authors: Zekai Luo, Zongze Du, Zhouhang Zhu, Hao Zhong, Muzhi Zhu, Wen Wang, Yuling Xi, Chenchen Jing, Hao Chen, Chunhua Shen Title: Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality Arxiv: http://arxiv.org/abs/2512.07951v1 Abstract: Video face swapping is crucial in film and entertainment production, where achieving high fidelity and tempor...
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory 11.12.2025 22:37
🤗 Upvotes: 32 | cs. CV Authors: Zhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou, Xiaoke Huang, Zhiheng Liu, Weiming Ren, Kumara Kahatapitiya, Ding Liu, Sen He, Chenyang Zhang, Tao Xiang, Fanny Yang, Serge Belongie, Tian Xie Title: OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory Arxiv: http://arxiv.org/abs/2512.07802v1 Abstract: Storytelling in real-world videos often unfold...
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning 10.12.2025 24:05
🤗 Upvotes: 55 | cs. CL Authors: Tong Wu, Yang Liu, Jun Bai, Zixia Jia, Shuyi Zhang, Ziyong Lin, Yanting Wang, Song-Chun Zhu, Zilong Zheng Title: Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Arxiv: http://arxiv.org/abs/2512.07461v1 Abstract: We introduce Native Parallel Reasoner (NPR), a teacher-free framework that enables Large Language Models (LLMs...
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs 10.12.2025 21:49
🤗 Upvotes: 45 | cs. CL Authors: Xiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Zhaoxiang Liu, Shiguo Lian, Ziwei He, Xipeng Qiu Title: Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs Arxiv: http://arxiv.org/abs/2512.07525v1 Abstract: Rotary Position Embeddings (RoPE) have become a standard for encoding sequence order in Large Language Mode...
Unified Video Editing with Temporal Reasoner 10.12.2025 20:34
🤗 Upvotes: 32 | cs. CV Authors: Xiangpeng Yang, Ji Xie, Yiyuan Yang, Yan Huang, Min Xu, Qiang Wu Title: Unified Video Editing with Temporal Reasoner Arxiv: http://arxiv.org/abs/2512.07469v1 Abstract: Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temporal in-context lea...
Voxify3D: Pixel Art Meets Volumetric Rendering 10.12.2025 22:17
🤗 Upvotes: 30 | cs. CV Authors: Yi-Chuan Huang, Jiewen Chan, Hao-Jen Chien, Yu-Lun Liu Title: Voxify3D: Pixel Art Meets Volumetric Rendering Arxiv: http://arxiv.org/abs/2512.07834v1 Abstract: Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to conflicting requirements of geometric abstraction, semantic p...
Scaling Zero-Shot Reference-to-Video Generation 10.12.2025 23:52
🤗 Upvotes: 26 | cs. CV Authors: Zijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu, Zhaochong An, Weiming Ren, Zhiheng Liu, Xiaoke Huang, Kam Woh Ng, Tian Xie, Xiao Han, Yuren Cong, Hang Li, Chuyan Zhu, Aditya Patel, Tao Xiang, Sen He Title: Scaling Zero-Shot Reference-to-Video Generation Arxiv: http://arxiv.org/abs/2512.06905v1 Abstract: Reference-to-video (R2V) generation aims to synthesize videos...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.