Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
LLM Safety From Within: Detecting Harmful Content with Internal Representations 28.04.2026 23:23
🤗 Upvotes: 21 | cs. AI Authors: Difan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang, Linfeng Du, Haolun Wu, Ashton Anderson Title: LLM Safety From Within: Detecting Harmful Content with Internal Representations Arxiv: http://arxiv.org/abs/2604.18519v1 Abstract: Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on...
LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics 25.04.2026 26:25
🤗 Upvotes: 77 | cs. AI Authors: Yueyang Ding, HaoPeng Zhang, Rui Dai, Yi Wang, Tianyu Zong, Kaikui Liu, Xiangxiang Chu Title: LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics Arxiv: http://arxiv.org/abs/2604.17295v1 Abstract: Comprehensive understanding of time series remains a significant challenge for Large Language Models (LLMs). Current research...
WorldMark: A Unified Benchmark Suite for Interactive Video World Models 25.04.2026 25:39
🤗 Upvotes: 30 | cs. CV Authors: Xiaojie Xu, Zhengyuan Lin, Kang He, Yukang Feng, Xiaofeng Mao, Yuanyang Yin, Kaipeng Zhang, Yongtao Ge Title: WorldMark: A Unified Benchmark Suite for Interactive Video World Models Arxiv: http://arxiv.org/abs/2604.21686v1 Abstract: Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated...
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling 25.04.2026 26:31
🤗 Upvotes: 25 | cs. RO, cs. AI Authors: Boyu Chen, Yi Chen, Lu Qiu, Jerry Bai, Yuying Ge, Yixiao Ge Title: UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling Arxiv: http://arxiv.org/abs/2604.19734v1 Abstract: Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alter...
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model 24.04.2026 25:15
🤗 Upvotes: 213 | cs. CV Authors: Inclusion AI, Tiwei Bie, Haoxing Chen, Tieyuan Chen, Zhenglin Cheng, Long Cui, Kai Gan, Zhicheng Huang, Zhenzhong Lan, Haoquan Li, Jianguo Li, Tao Lin, Qi Qin, Hongjun Wang, Xiaomei Wang, Haoyuan Wu, Yi Xin, Junbo Zhao Title: LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Arxiv: http://arxiv.org/abs/2604.20796v1...
Near-Future Policy Optimization 24.04.2026 22:20
🤗 Upvotes: 45 | cs. LG Authors: Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang Title: Near-Future Policy Optimization Arxiv: http://arxiv.org/abs/2604.20733v1 Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerate...
DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data 24.04.2026 24:27
🤗 Upvotes: 38 | cs. LG, cs. AI, cs. CL, cs. IR Authors: Venus Team, Sunhao Dai, Yong Deng, Jinzhen Lin, Yusheng Song, Guoqing Wang, Xiaofeng Wu, Yuqi Zhou, Shuo Yang, Zhenzhe Ying, Zhanwei Zhang, Changhua Meng, Weiqiang Wang Title: DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data Arxiv: http://arxiv.org/abs/2604.19859v1 Abstract: Edge-scale deep research agents b...
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis 24.04.2026 26:56
🤗 Upvotes: 26 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Kanzhi Cheng, Zehao Li, Zheng Ma, Nuo Chen, Jialin Cao, Qiushi Sun, Zichen Ding, Fangzhi Xu, Hang Yan, Jiajun Chen, Anh Tuan Luu, Jianbing Zhang, Lewei Lu, Dahua Lin Title: OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Arxiv: http://arxiv.org/abs/2604.15093v1 Abstract: Mobile agents powered by vision-language mod...
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation 24.04.2026 23:40
🤗 Upvotes: 21 | cs. CV Authors: Hyeonwoo Kim, Jeonghwan Kim, Kyungwon Cho, Hanbyul Joo Title: DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation Arxiv: http://arxiv.org/abs/2604.20841v1 Abstract: Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, incl...
Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items 23.04.2026 23:57
🤗 Upvotes: 82 | cs. CV Authors: Mengting Chen, Zhengrui Chen, Yongchao Du, Zuan Gao, Taihang Hu, Jinsong Lan, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Mingzhou Zhang, Jun Zheng, Qinye Zhou, Xiaoyong Zhu, Bo Zheng Title: Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items Arxiv: http://arxiv.org/abs/2604.19748v2 Abstr...
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation 23.04.2026 21:16
🤗 Upvotes: 69 | cs. CV Authors: Xiangyang Luo, Xiaozhe Xin, Tao Feng, Xu Guo, Meiguang Jin, Junfeng Ma Title: CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Arxiv: http://arxiv.org/abs/2604.19636v1 Abstract: Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtua...
AgentSPEX: An Agent SPecification and EXecution Language 23.04.2026 22:39
🤗 Upvotes: 52 | cs. CL Authors: Pengcheng Wang, Jerry Huang, Jiarui Yao, Rui Pan, Peizhi Niu, Yaowenqi Liu, Ruida Wang, Renhao Lu, Yuwei Guo, Tong Zhang Title: AgentSPEX: An Agent SPecification and EXecution Language Arxiv: http://arxiv.org/abs/2604.13346v1 Abstract: Language-model agent systems commonly rely on reactive prompting, in which a single instruction guides the model through an open-en...
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model 23.04.2026 24:31
🤗 Upvotes: 35 | cs. CV Authors: Yutian Chen, Shi Guo, Renbiao Jin, Tianshuo Yang, Xin Cai, Yawen Luo, Mingxin Yang, Mulin Yu, Linning Xu, Tianfan Xue Title: AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model Arxiv: http://arxiv.org/abs/2604.19747v1 Abstract: Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-gener...
TEMPO: Scaling Test-time Training for Large Reasoning Models 23.04.2026 23:30
🤗 Upvotes: 26 | cs. LG Authors: Qingyang Zhang, Xinke Kong, Haitao Wu, Qinghua Hu, Minghao Wu, Baosong Yang, Yu Cheng, Yun Luo, Ganqu Cui, Changqing Zhang Title: TEMPO: Scaling Test-time Training for Large Reasoning Models Arxiv: http://arxiv.org/abs/2604.19295v1 Abstract: Test-time training (TTT) adapts model parameters on unlabeled test instances during inference time, which continuously extend...
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation 22.04.2026 20:55
🤗 Upvotes: 87 | cs. CV Authors: Chenxi Zhao, Chen Zhu, Xiaokun Feng, Aiming Hao, Jiashu Zhu, Jiachen Lei, Jiahong Wu, Xiangxiang Chu, Jufeng Yang Title: Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Arxiv: http://arxiv.org/abs/2604.18168v1 Abstract: Few-step generation has been a long-standing goal, with recent one-step generation methods exe...
OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation 22.04.2026 26:31
🤗 Upvotes: 67 | cs. CV, cs. CL, cs. RO Authors: Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, Lingdong Kong, Yingyan Li, Han Wang, Shaoqing Xu, Yuechen Luo, Fang Li, Chenxu Dang, Junli Wang, Tao Xu, Jing Wu, Jianhua Wu, Xiaoshuai Hao, Wen Zhang, Tianyi Jiang, Lingfeng Zhang, Lei Zhou, Yingbo Tang, Jie Wang, Yinfeng Gao, Xizhou Bu, Haochen Tian, Yihang Qiu, Feiyang Jia, Lin Liu, Yig...
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence 22.04.2026 24:00
🤗 Upvotes: 66 | cs. AI, cs. CL Authors: Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, Jiajie Jin, Yutao Zhu, Hanbin Wang, Fangyu Lei, Qinyu Luo, Mingyang Chen, Zehui Chen, Jiazhan Feng, Ji-Rong Wen, Zhicheng Dou Title: Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence...
OpenGame: Open Agentic Coding for Games 22.04.2026 25:31
🤗 Upvotes: 51 | cs. SE Authors: Yilei Jiang, Jinyuan Hu, Qianyin Xiao, Yaozhi Zheng, Ruize Ma, Kaituo Feng, Jiaming Han, Tianshuo Peng, Kaixuan Fan, Manyuan Zhang, Xiangyu Yue Title: OpenGame: Open Agentic Coding for Games Arxiv: http://arxiv.org/abs/2604.18394v1 Abstract: Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestr...
MultiWorld: Scalable Multi-Agent Multi-View Video World Models 22.04.2026 21:41
🤗 Upvotes: 36 | cs. CV Authors: Haoyu Wu, Jiwen Yu, Yingtian Zou, Xihui Liu Title: MultiWorld: Scalable Multi-Agent Multi-View Video World Models Arxiv: http://arxiv.org/abs/2604.18564v2 Abstract: Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that ta...
EasyVideoR1: Easier RL for Video Understanding 22.04.2026 27:05
🤗 Upvotes: 32 | cs. CV, cs. LG Authors: Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang Title: EasyVideoR1: Easier RL for Video Understanding Arxiv: http://arxiv.org/abs/2604.16893v1 Abstract: Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language...
Elucidating the SNR-t Bias of Diffusion Probabilistic Models 21.04.2026 22:07
🤗 Upvotes: 69 | cs. CV Authors: Meng Yu, Lei Sun, Jianhao Zeng, Xiangxiang Chu, Kun Zhan Title: Elucidating the SNR-t Bias of Diffusion Probabilistic Models Arxiv: http://arxiv.org/abs/2604.16044v1 Abstract: Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Signal-to-Nois...
Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips 21.04.2026 22:28
🤗 Upvotes: 42 | cs. LG, cs. AI, cs. CV Authors: Ido Galil, Moshe Kimhi, Ran El-Yaniv Title: Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips Arxiv: http://arxiv.org/abs/2502.07408v2 Abstract: Deep Neural Networks (DNNs) can be catastrophically disrupted by flipping only a handful of parameter bits. We introduce Deep Neural Lesion (DNL), a data-free...
PersonaVLM: Long-Term Personalized Multimodal LLMs 21.04.2026 24:57
🤗 Upvotes: 39 | cs. CL, cs. CV Authors: Chang Nie, Chaoyou Fu, Yifan Zhang, Haihua Yang, Caifeng Shan Title: PersonaVLM: Long-Term Personalized Multimodal LLMs Arxiv: http://arxiv.org/abs/2604.13074v1 Abstract: Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual preferences remains limited. Prior app...
Qwen3.5-Omni Technical Report 21.04.2026 24:57
🤗 Upvotes: 28 | cs. CL, eess. AS Authors: Qwen Team Title: Qwen3.5-Omni Technical Report Arxiv: http://arxiv.org/abs/2604.15804v1 Abstract: In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging...
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds 18.04.2026 24:06
🤗 Upvotes: 69 | cs. CV Authors: Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, Bo Yuan, Chao Zhang, Coopers Li, Dongyuan Guo, Fan Yang, Haiyu Zhang, Hang Cao, Jianchen Zhu, Jiaxin Lin, Jie Xiao, Jihong Zhang, Junlin Yu, Lei Wang, Lifu Wang, Lilin Wang, Linus, Minghui Chen, Peng He, Penghao Zhao, Qi Chen, Rui Chen, Rui Shao, Sicong L...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.