Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems 10.12.2025 24:53
🤗 Upvotes: 24 | cs. AI, cs. SE Authors: Ming Ma, Jue Zhang, Fangkai Yang, Yu Kang, Qingwei Lin, Tianming Yang, Saravan Rajmohan, Dongmei Zhang Title: DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems Arxiv: http://arxiv.org/abs/2512.06749v2 Abstract: Large language model (LLM)-based multi-agent systems are challenging to debug because failures often arise from long, branching...
TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows 09.12.2025 22:53
🤗 Upvotes: 47 | cs. CV Authors: Zhenglin Cheng, Peng Sun, Jianguo Li, Tao Lin Title: TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows Arxiv: http://arxiv.org/abs/2512.05150v1 Abstract: Recent advances in large multi-modal generative models have demonstrated impressive capabilities in multi-modal generation, including image and video generation. These models are...
EditThinker: Unlocking Iterative Reasoning for Any Image Editor 09.12.2025 26:31
🤗 Upvotes: 32 | cs. CV Authors: Hongyu Li, Manyuan Zhang, Dian Zheng, Ziyu Guo, Yimeng Jia, Kaituo Feng, Hao Yu, Yexin Liu, Yan Feng, Peng Pei, Xunliang Cai, Linjiang Huang, Hongsheng Li, Si Liu Title: EditThinker: Unlocking Iterative Reasoning for Any Image Editor Arxiv: http://arxiv.org/abs/2512.05965v1 Abstract: Instruction-based image editing has emerged as a prominent research area, which, b...
From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks 09.12.2025 23:39
🤗 Upvotes: 25 | cs. CL Authors: Changpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang, Yang Li, Qiliang Liang, Hongzhen Wang, Shuai Nie, Jiaming Xu, Runyu Shi, Ying Huang, Guoquan Zhang Title: From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks Arxiv: http://arxiv.org/abs/2512.02580v1 Abstract: Reinforcement learning has emerg...
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture 09.12.2025 24:33
🤗 Upvotes: 22 | cs. CV Authors: Xin He, Longhui Wei, Jianbo Ouyang, Lingxi Xie, Qi Tian Title: EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture Arxiv: http://arxiv.org/abs/2512.04810v2 Abstract: We propose EMMA, an efficient and unified architecture for multimodal understanding, generation and editing. Specifically, EMMA primarily consists of 1) An eff...
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle 06.12.2025 27:25
🤗 Upvotes: 120 | cs. CL, cs. AI Authors: Fangyu Lei, Jinxiang Meng, Yiming Huang, Junjie Zhao, Yitong Zhang, Jianwen Luo, Xin Zou, Ruiyi Yang, Wenbo Shi, Yan Gao, Shizhu He, Zuo Wang, Qian Liu, Yang Wang, Ke Wang, Jun Zhao, Kang Liu Title: DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Arxiv: http://arxiv.org/abs/2512.04324v1 Abstract: Real-world enterprise data inte...
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length 06.12.2025 25:15
🤗 Upvotes: 113 | cs. CV Authors: Yubo Huang, Hailong Guo, Fangtai Wu, Shifeng Zhang, Shijie Huang, Qijun Gan, Lin Liu, Sirui Zhao, Enhong Chen, Jiaming Liu, Steven Hoi Title: Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Arxiv: http://arxiv.org/abs/2512.04677v1 Abstract: Existing diffusion-based video generation methods are fundamentally constrained by seque...
Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction 06.12.2025 24:40
🤗 Upvotes: 58 | cs. CL Authors: Nex-AGI Team, :, Yuxuan Cai, Lu Chen, Qiaoling Chen, Yuyang Ding, Liwen Fan, Wenjie Fu, Yufei Gao, Honglin Guo, Pinxue Guo, Zhenhua Han, Zhengfu He, Hanglei Hu, Kai Hu, Shengjia Hua, Tianyu Huai, Baodai Huang, Li Ji, Zhen Jiang, Zhikai Lei, Bufan Li, Jiahang Lin, Lizhi Lin, Jinxiu Liu, Shichun Liu, Ziming Liu, Yuchen Ni, Pengfang Qian, Yujiong Shen, Qingyun Shi, We...
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning 06.12.2025 23:40
🤗 Upvotes: 35 | cs. CV Authors: Shengyuan Ding, Xinyu Fang, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiangyu Zhao, Haodong Duan, Xiaoyi Dong, Jianze Liang, Bin Wang, Conghui He, Dahua Lin, Jiaqi Wang Title: ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning Arxiv: http://arxiv.org/abs/2512.05111v1 Abstract: Reward models are critical for aligning vis...
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation 06.12.2025 21:09
🤗 Upvotes: 31 | cs. CV Authors: Yunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Jiapeng Zhu, Hengyuan Cao, Zhipeng Zhang, Xing Zhu, Yujun Shen, Min Zhang Title: Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation Arxiv: http://arxiv.org/abs/2512.04678v1 Abstract: Efficient streaming video generation is critical for simu...
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion 06.12.2025 25:40
🤗 Upvotes: 26 | cs. CV Authors: Yueming Pan, Ruoyu Feng, Qi Dai, Yuqi Wang, Wenfeng Lin, Mingyu Guo, Chong Luo, Nanning Zheng Title: Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion Arxiv: http://arxiv.org/abs/2512.04926v1 Abstract: Latent Diffusion Models (LDMs) inherently follow a coarse-to-fine generation process, where high-level semantic st...
PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing 06.12.2025 23:24
🤗 Upvotes: 23 | cs. AI, cs. SE Authors: Junyi Hou, Andre Lin Huikai, Nuo Chen, Yiwei Gong, Bingsheng He Title: PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing Arxiv: http://arxiv.org/abs/2512.02589v1 Abstract: Large language models are increasingly embedded into academic writing workflows, yet existing assistants remain external to the editor,...
Qwen3-VL Technical Report 05.12.2025 27:05
🤗 Upvotes: 91 | cs. CV, cs. AI Authors: Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Li...
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach 05.12.2025 22:30
🤗 Upvotes: 33 | cs. RO, cs. AI Authors: Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li Title: Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach Arxiv: http://arxiv.org/abs/2512.02834v1 Abstract: Vision-Language-Action (VLA) models, trained via flow-matching or diffusion objectives, excel at learning complex behaviors from large...
PretrainZero: Reinforcement Active Pretraining 05.12.2025 22:48
🤗 Upvotes: 31 | cs. CL Authors: Xingrun Xing, Zhiyuan Fan, Jie Lou, Guoqi Li, Jiajun Zhang, Debing Zhang Title: PretrainZero: Reinforcement Active Pretraining Arxiv: http://arxiv.org/abs/2512.03442v1 Abstract: Mimicking human behavior to actively learning from general experience and achieve artificial general intelligence has always been a human dream. Recent reinforcement learning (RL) based lar...
ViDiC: Video Difference Captioning 05.12.2025 23:33
🤗 Upvotes: 24 | cs. CV Authors: Jiangtao Wu, Shihao Li, Zhaozhou Bian, Yuanxing Zhang, Jialu Chen, Runzhe Wen, An Ping, Yiwen He, Jiakai Wang, Jiaheng Liu Title: ViDiC: Video Difference Captioning Arxiv: http://arxiv.org/abs/2512.03405v1 Abstract: Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capabili...
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models 04.12.2025 22:11
🤗 Upvotes: 114 | cs. CL Authors: DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Li, Hao...
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration 04.12.2025 21:23
🤗 Upvotes: 63 | cs. CL, cs. AI, cs. LG, cs. MA Authors: Hongjin Su, Shizhe Diao, Ximing Lu, Mingjie Liu, Jiacheng Xu, Xin Dong, Yonggan Fu, Peter Belcak, Hanrong Ye, Hongxu Yin, Yi Dong, Evelina Bakhturina, Tao Yu, Yejin Choi, Jan Kautz, Pavlo Molchanov Title: ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration Arxiv: http://arxiv.org/abs/2511.21689v1 Abstract: Large...
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 04.12.2025 27:19
🤗 Upvotes: 50 | cs. CV Authors: Qinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian, Quande Liu, Huchuan Lu, Xintao Wang, Pengfei Wan, Kun Gai, Xu Jia Title: MultiShotMaster: A Controllable Multi-Shot Video Generation Framework Arxiv: http://arxiv.org/abs/2512.03041v1 Abstract: Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which re...
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory 04.12.2025 23:42
🤗 Upvotes: 44 | cs. CV, cs. RO Authors: Bo Wang, Jiehong Lin, Chenzhi Liu, Xinting Hu, Yifei Yu, Tianjia Liu, Zhongrui Wang, Xiaojuan Qi Title: MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory Arxiv: http://arxiv.org/abs/2511.22609v1 Abstract: We present MG-Nav (Memory-Guided Navigation), a dual-scale framework for zero-shot visual navigation that unifies global memory-guided planni...
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch 04.12.2025 25:38
🤗 Upvotes: 38 | cs. CV Authors: Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang, Yichen Wei, Shukang Yin, Jiangbo Pei, Wei Shen, Peng Xia, Yi Peng, Tianyidan Xie, Eric Li, Yang Liu, Xuchen Song, Yahui Zhou Title: Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch Arxiv: http://arxiv.org/abs/2512.02395v1 Abstract: Despite recent progress i...
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation 04.12.2025 20:31
🤗 Upvotes: 38 | cs. CV Authors: Hongfei Zhang, Kanghao Chen, Zixin Zhang, Harold Haodong Chen, Yuanhuiyi Lyu, Yuqi Zhang, Shuai Yang, Kun Zhou, Yingcong Chen Title: DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation Arxiv: http://arxiv.org/abs/2511.23127v2 Abstract: This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlle...
Guided Self-Evolving LLMs with Minimal Human Supervision 04.12.2025 25:25
🤗 Upvotes: 37 | cs. AI, cs. CL, cs. LG Authors: Wenhao Yu, Zhenwen Liang, Chengsong Huang, Kishan Panaganti, Tianqing Fang, Haitao Mi, Dong Yu Title: Guided Self-Evolving LLMs with Minimal Human Supervision Arxiv: http://arxiv.org/abs/2512.02472v1 Abstract: AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize kno...
SimScale: Learning to Drive via Real-World Simulation at Scale 04.12.2025 22:38
🤗 Upvotes: 33 | cs. CV, cs. RO Authors: Haochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang, Yihang Qiu, Guang Li, Junli Wang, Yinfeng Gao, Zhang Zhang, Liang Wang, Hangjun Ye, Tieniu Tan, Long Chen, Hongyang Li Title: SimScale: Learning to Drive via Real-World Simulation at Scale Arxiv: http://arxiv.org/abs/2511.23369v1 Abstract: Achieving fully autonomous driving systems requires learning rationa...
InnoGym: Benchmarking the Innovation Potential of AI Agents 04.12.2025 23:37
🤗 Upvotes: 30 | cs. CL, cs. AI, cs. CV, cs. LG, cs. MA Authors: Jintian Zhang, Kewei Xu, Jingsheng Zheng, Zhuoyun Yu, Yuqi Zhu, Yujie Luo, Lanning Wei, Shuofei Qiao, Lun Du, Da Zheng, Shumin Deng, Huajun Chen, Ningyu Zhang Title: InnoGym: Benchmarking the Innovation Potential of AI Agents Arxiv: http://arxiv.org/abs/2512.01822v1 Abstract: LLMs and Agents have achieved impressive progress in code...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.