Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models 05.08.2025 19:37
🤗 Upvotes: 39 | cs. CL Authors: Jinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Dahua Lin Title: Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models Arxiv: http://arxiv.org/abs/2508.00819v1 Abstract: Diffusion Large Language Models (DLLMs) are emerging as a powerful alternative to the dominant Autoregressive Large Language Models, offering efficient parallel...
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training 05.08.2025 20:57
🤗 Upvotes: 38 | cs. AI, cs. CL Authors: Tianqing Fang, Zhisong Zhang, Xiaoyang Wang, Rui Wang, Can Qin, Yuxuan Wan, Jun-Yu Ma, Ce Zhang, Jiaqi Chen, Xiyun Li, Hongming Zhang, Haitao Mi, Dong Yu Title: Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training Arxiv: http://arxiv.org/abs/2508.00414v1 Abstract: General AI Agents are increasingly recognized as fo...
PixNerd: Pixel Neural Field Diffusion 05.08.2025 23:41
🤗 Upvotes: 33 | cs. CV Authors: Shuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang, Limin Wang Title: PixNerd: Pixel Neural Field Diffusion Arxiv: http://arxiv.org/abs/2507.23268v2 Abstract: The current success of diffusion transformers heavily depends on the compressed latent space shaped by the pre-trained variational autoencoder(VAE). However, this two-stage training paradigm inevitably introdu...
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving 02.08.2025 21:15
🤗 Upvotes: 65 | cs. AI, cs. CL Authors: Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Kaijing Ma, Cheng Ren, Jiawei Shen, Wenlei Shi, Tong Sun, He Sun, Jiahui Wang, Siran Wang, Zhihong Wang, Chenrui Wei, Shufa Wei, Yonghui Wu, Yuchen Wu, Yihang Xia, Huajian Xin, Fan Yang, Huaiyuan Ying, Hongyi Yuan, Zheng Yuan, Tianyang Zhan,...
Phi-Ground Tech Report: Advancing Perception in GUI Grounding 02.08.2025 21:26
🤗 Upvotes: 28 | cs. CV, cs. AI, cs. MM Authors: Miaosen Zhang, Ziqiang Xu, Jialiang Zhu, Qi Dai, Kai Qiu, Yifan Yang, Chong Luo, Tianyi Chen, Justin Wagle, Tim Franklin, Baining Guo Title: Phi-Ground Tech Report: Advancing Perception in GUI Grounding Arxiv: http://arxiv.org/abs/2507.23779v1 Abstract: With the development of multimodal reasoning models, Computer Use Agents (CUAs), akin to Jarvis f...
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents 01.08.2025 20:36
🤗 Upvotes: 62 | cs. CV Authors: Yilei Jiang, Yaozhi Zheng, Yuxuan Wan, Jiaming Han, Qunzhong Wang, Michael R. Lyu, Xiangyu Yue Title: ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents Arxiv: http://arxiv.org/abs/2507.22827v1 Abstract: Automating the transformation of user interface (UI) designs into front-end code holds significant promise for...
BANG: Dividing 3D Assets via Generative Exploded Dynamics 01.08.2025 20:32
🤗 Upvotes: 46 | cs. GR Authors: Longwen Zhang, Qixuan Zhang, Haoran Jiang, Yinuo Bai, Wei Yang, Lan Xu, Jingyi Yu Title: BANG: Dividing 3D Assets via Generative Exploded Dynamics Arxiv: http://arxiv.org/abs/2507.21493v1 Abstract: 3D creation has always been a unique human strength, driven by our ability to deconstruct and reassemble objects using our eyes, mind and hand. However, current 3D desig...
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning 01.08.2025 22:33
🤗 Upvotes: 30 | cs. CV, cs. AI, cs. CL Authors: Ruifeng Yuan, Chenghao Xiao, Sicong Leng, Jianyu Wang, Long Li, Weiwen Xu, Hou Pong Chan, Deli Zhao, Tingyang Xu, Zhongyu Wei, Hao Zhang, Yu Rong Title: VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Arxiv: http://arxiv.org/abs/2507.22607v2 Abstract: Reinforcement learning has proven its effectiveness in e...
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels 31.07.2025 24:01
🤗 Upvotes: 67 | cs. CV Authors: HunyuanWorld Team, Zhenwei Wang, Yuhao Liu, Junta Wu, Zixiao Gu, Haoyuan Wang, Xuhui Zuo, Tianyu Huang, Wenhuan Li, Sheng Zhang, Yihang Lian, Yulin Tsai, Lifu Wang, Sicong Liu, Puhua Jiang, Xianghui Yang, Dongyuan Guo, Yixuan Tang, Xinyue Mao, Jiaao Yu, Junlin Yu, Jihong Zhang, Meng Chen, Liang Dong, Yiwen Jia, Chao Zhang, Yonghao Tan, Hao Zhang, Zheng Ye, Peng He,...
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again 31.07.2025 17:59
🤗 Upvotes: 24 | cs. CV Authors: Zigang Geng, Yibing Wang, Yeyao Ma, Chen Li, Yongming Rao, Shuyang Gu, Zhao Zhong, Qinglin Lu, Han Hu, Xiaosong Zhang, Linus, Di Wang, Jie Jiang Title: X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again Arxiv: http://arxiv.org/abs/2507.22058v1 Abstract: Numerous efforts have been made to extend the ``next token predicti...
ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge 31.07.2025 22:39
🤗 Upvotes: 21 | cs. CE, cs. AI Authors: Zihan Zhao, Bo Chen, Ziping Wan, Lu Chen, Xuanze Lin, Shiyang Yu, Situo Zhang, Da Ma, Zichen Zhu, Danyang Zhang, Huayang Wang, Zhongyang Dai, Liyang Wen, Xin Chen, Kai Yu Title: ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge Arxiv: http://arxiv.org/abs/2507.21990v2 Abstract: While large language models (LLMs) have achieved imp...
Agentic Reinforced Policy Optimization 30.07.2025 22:10
🤗 Upvotes: 84 | cs. LG, cs. AI, cs. CL Authors: Guanting Dong, Hangyu Mao, Kai Ma, Licheng Bao, Yifei Chen, Zhongyuan Wang, Zhongxia Chen, Jiazhen Du, Huiyang Wang, Fuzheng Zhang, Guorui Zhou, Yutao Zhu, Ji-Rong Wen, Zhicheng Dou Title: Agentic Reinforced Policy Optimization Arxiv: http://arxiv.org/abs/2507.19849v1 Abstract: Large-scale reinforcement learning with verifiable rewards (RLVR) has de...
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts 30.07.2025 22:18
🤗 Upvotes: 50 | cs. CV Authors: Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan Title: ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Arxiv: http://arxiv.org/abs/2507.20939v1 Abstract: Real-world user-generated short videos, esp...
A Survey of Self-Evolving Agents: On Path to Artificial Super Intelligence 30.07.2025 22:05
🤗 Upvotes: 40 | cs. AI Authors: Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Qihan Ren, Cheng Qian, Zhenghailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, Mengdi Wang Title: A Survey of Self-Evol...
Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning 30.07.2025 20:02
🤗 Upvotes: 36 | cs. LG, cs. CV Authors: Zedong Wang, Siyuan Li, Dan Xu Title: Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning Arxiv: http://arxiv.org/abs/2507.21049v1 Abstract: Despite the promise of Multi-Task Learning in leveraging complementary knowledge across tasks, existing multi-task optimization (MTO) techniques remain fixated on resolving confl...
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment 30.07.2025 23:17
🤗 Upvotes: 32 | cs. LG, cs. AI Authors: Yixin Song, Zhenliang Xue, Dongliang Wei, Feiyang Chen, Jianxiang Gao, Junchen Liu, Hangyu Liang, Guangshuo Qin, Chengrong Tian, Bo Wen, Longyu Zhao, Xinrui Zheng, Zeyu Mi, Haibo Chen Title: SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment Arxiv: http://arxiv.org/abs/2507.20984v1 Abstract: While frontier large...
Reconstructing 4D Spatial Intelligence: A Survey 30.07.2025 21:29
🤗 Upvotes: 27 | cs. CV Authors: Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowei Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, Ziwei Liu Title: Reconstructing 4D Spatial Intelligence: A Survey Arxiv: http://arxiv.org/abs/2507.21045v1 Abstract: Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer...
Deep Researcher with Test-Time Diffusion 29.07.2025 22:59
🤗 Upvotes: 25 | cs. CL Authors: Rujun Han, Yanfei Chen, Zoey CuiZhu, Lesly Miculicich, Guan Sun, Yuanjun Bi, Weiming Wen, Hui Wan, Chunfeng Wen, Solène Maître, George Lee, Vishy Tirumalashetty, Emily Xue, Zizhao Zhang, Salem Haykal, Burak Gokturk, Tomas Pfister, Chen-Yu Lee Title: Deep Researcher with Test-Time Diffusion Arxiv: http://arxiv.org/abs/2507.16075v1 Abstract: Deep research agents, pow...
$\nabla$NABLA: Neighborhood Adaptive Block-Level Attention 26.07.2025 21:11
🤗 Upvotes: 81 | cs. CV Authors: Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva, Vladimir Arkhipkin, Vladimir Korviakov, Vladimir Polovnikov, Viacheslav Vasilev, Evelina Sidorova, Denis Dimitrov Title: $\nabla$NABLA: Neighborhood Adaptive Block-Level Attention Arxiv: http://arxiv.org/abs/2507.13546v1 Abstract: Recent progress in transformer-based architectures has demonstrated remarkable s...
Group Sequence Policy Optimization 26.07.2025 23:18
🤗 Upvotes: 57 | cs. LG, cs. AI, cs. CL Authors: Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, Junyang Lin Title: Group Sequence Policy Optimization Arxiv: http://arxiv.org/abs/2507.18071v1 Abstract: This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement...
MUR: Momentum Uncertainty guided Reasoning for Large Language Models 26.07.2025 22:29
🤗 Upvotes: 31 | cs. CL Authors: Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Luu Anh Tuan, Haiteng Zhao, Qika Lin, Jun Liu Title: MUR: Momentum Uncertainty guided Reasoning for Large Language Models Arxiv: http://arxiv.org/abs/2507.14958v1 Abstract: Large Language Models (LLMs) have achieved impressive performance on reasoning-intensive tasks, yet optimizing the...
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization 26.07.2025 20:14
🤗 Upvotes: 25 | cs. AI, cs. CL Authors: Xingyu Wu, Yuchen Yan, Shangke Lyu, Linjuan Wu, Yiwen Qiu, Yongliang Shen, Weiming Lu, Jian Shao, Jun Xiao, Yueting Zhuang Title: LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Arxiv: http://arxiv.org/abs/2507.15758v1 Abstract: Large reasoning models have achieved remarkable performance through extended chain-of-thought seq...
Pixels, Patterns, but No Poetry: To See The World like Humans 25.07.2025 17:16
🤗 Upvotes: 48 | cs. CV, cs. AI, cs. CL Authors: Hongcheng Gao, Zihao Huang, Lin Xu, Jingyi Tang, Xinhao Li, Yue Liu, Haoyang Li, Taihang Hu, Minhua Lin, Xinlong Yang, Ge Wu, Balong Bi, Hongyu Chen, Wentao Zhang Title: Pixels, Patterns, but No Poetry: To See The World like Humans Arxiv: http://arxiv.org/abs/2507.16863v1 Abstract: Achieving human-like perception and reasoning in Multimodal Large La...
Yume: An Interactive World Generation Model 25.07.2025 26:11
🤗 Upvotes: 45 | cs. CV, cs. AI, cs. HC Authors: Xiaofeng Mao, Shaoheng Lin, Zhen Li, Chuanhao Li, Wenshuo Peng, Tong He, Jiangmiao Pang, Mingmin Chi, Yu Qiao, Kaipeng Zhang Title: Yume: An Interactive World Generation Model Arxiv: http://arxiv.org/abs/2507.17744v1 Abstract: Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration an...
DesignLab: Designing Slides Through Iterative Detection and Correction 25.07.2025 23:08
🤗 Upvotes: 33 | cs. CV, cs. AI Authors: Jooyeol Yun, Heng Wang, Yotaro Shimose, Jaegul Choo, Shingo Takamatsu Title: DesignLab: Designing Slides Through Iterative Detection and Correction Arxiv: http://arxiv.org/abs/2507.17202v1 Abstract: Designing high-quality presentation slides can be challenging for non-experts due to the complexity involved in navigating various design choices. Numerous auto...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.