Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models 17.07.2025 20:04
🤗 Upvotes: 32 | cs. CV Authors: Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen, Alan Yuille, Chen Wei, Junfei Xiao Title: Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models Arxiv: http://arxiv.org/abs/2507.07104v2 Abstract: Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on...
EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes 17.07.2025 19:27
🤗 Upvotes: 24 | cs. CL, cs. AI Authors: LG AI Research, :, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Kyubeen Han, Seokhee Hong, Junwon Hwang, Taewan Hwang, Joonwon Jang, Hyojin Jeon, Kijeong Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Euisoon Kim, Hyosang Kim, Jihoon Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Youchul...
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination 16.07.2025 21:11
🤗 Upvotes: 44 | cs. LG, cs. AI, cs. CL Authors: Mingqi Wu, Zhihao Zhang, Qiaole Dong, Zhiheng Xi, Jun Zhao, Senjie Jin, Xiaoran Fan, Yuhao Zhou, Yanwei Fu, Qin Liu, Songyang Zhang, Qi Zhang Title: Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination Arxiv: http://arxiv.org/abs/2507.10532v1 Abstract: The reasoning capabilities of large language models (...
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation 16.07.2025 20:25
🤗 Upvotes: 43 | cs. CV, eess. AS Authors: Youliang Zhang, Zhaoyang Li, Duomin Wang, Jiahe Zhang, Deyu Zhou, Zixin Yin, Xili Dai, Gang Yu, Xiu Li Title: SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation Arxiv: http://arxiv.org/abs/2507.09862v1 Abstract: The rapid development of large-scale models has catalyzed significant breakthroughs in the di...
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation 16.07.2025 21:55
🤗 Upvotes: 31 | cs. CL, cs. LG Authors: Sangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Aaron Courville, Se-Young Yun Title: Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation Arxiv: http://arxiv.org/abs/2507.10524v1 Abstract: Scaling language models unlocks impressive capabilities, but...
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments 16.07.2025 22:15
🤗 Upvotes: 25 | cs. CV, cs. AI, cs. CL Authors: Mingxian Lin, Wei Huang, Yitang Li, Chengjie Jiang, Kui Wu, Fangwei Zhong, Shengju Qian, Xin Wang, Xiaojuan Qi Title: EmbRACE-3K: Embodied Reasoning and Action in Complex Environments Arxiv: http://arxiv.org/abs/2507.10548v1 Abstract: Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and vide...
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once 16.07.2025 25:24
🤗 Upvotes: 22 | cs. CL Authors: Zhuoshi Pan, Qizhi Pei, Yu Li, Qiyao Sun, Zinan Tang, H. Vicky Zhao, Conghui He, Lijun Wu Title: REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Arxiv: http://arxiv.org/abs/2507.10541v2 Abstract: Recent Large Reasoning Models (LRMs) have achieved remarkable progress on task-specific benchmarks, yet their evaluation methods remain con...
Test-Time Scaling with Reflective Generative Model 15.07.2025 21:33
🤗 Upvotes: 68 | cs. LG, cs. CL Authors: Zixiao Wang, Yuxin Wang, Xiaorui Wang, Mengting Xing, Jie Gao, Jianjun Xu, Guangcan Liu, Chenhui Jin, Zhuo Wang, Shengzhuo Zhang, Hongtao Xie Title: Test-Time Scaling with Reflective Generative Model Arxiv: http://arxiv.org/abs/2507.01951v2 Abstract: We introduce our first reflective generative model MetaStone-S1, which obtains OpenAI o3-mini's performance...
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning 15.07.2025 21:14
🤗 Upvotes: 47 | cs. CV, cs. CL Authors: Yana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin, Jisheng Yin, Jingcheng Hu, Yinmin Zhang, En Yu, Haoran Lv, Zejia Weng, Jia Wang, Chunrui Han, Yuang Peng, Qi Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel Title: Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning Arxiv: http://arxiv.org/abs/2507.05255v1 Abstrac...
NeuralOS: Towards Simulating Operating Systems via Neural Generative Models 15.07.2025 21:04
🤗 Upvotes: 45 | cs. CV, cs. AI, cs. CL, cs. HC, cs. LG Authors: Luke Rivard, Sun Sun, Hongyu Guo, Wenhu Chen, Yuntian Deng Title: NeuralOS: Towards Simulating Operating Systems via Neural Generative Models Arxiv: http://arxiv.org/abs/2507.08800v1 Abstract: We introduce NeuralOS, a neural framework that simulates graphical user interfaces (GUIs) of operating systems by directly predicting screen f...
CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering 15.07.2025 20:45
🤗 Upvotes: 43 | cs. CV Authors: Zhengqing Wang, Yuefan Wu, Jiacheng Chen, Fuyang Zhang, Yasutaka Furukawa Title: CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering Arxiv: http://arxiv.org/abs/2507.08776v2 Abstract: This paper proposes a neural rendering approach that represents a scene as "compressed light-field tokens (CLiFTs)", retaining rich appearance an...
KV Cache Steering for Inducing Reasoning in Small Language Models 15.07.2025 22:47
🤗 Upvotes: 26 | cs. CL, cs. AI Authors: Max Belitsky, Dawid J. Kopiczko, Michael Dorkenwald, M. Jehanzeb Mirza, Cees G. M. Snoek, Yuki M. Asano Title: KV Cache Steering for Inducing Reasoning in Small Language Models Arxiv: http://arxiv.org/abs/2507.08799v1 Abstract: We propose cache steering, a lightweight method for implicit steering of language models via a one-shot intervention applied direct...
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities 15.07.2025 20:41
🤗 Upvotes: 24 | cs. CL, cs. AI Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra...
Neural-Driven Image Editing 15.07.2025 20:36
🤗 Upvotes: 22 | cs. CV Authors: Pengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao, Zilong Ye, Zekai Li, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Ziqiao Wang, Kai Wang, Qian Zheng, Xiaojun Chang, Gang Pan, Shurong Dong, Kaipeng Zhang, Yang You Title: Neural-Driven Image Editing Arxiv: http://arxiv.org/abs/2507.05397v1 Abstract: Traditional image editing typically relies on manual prompting, maki...
Scaling RL to Long Videos 12.07.2025 22:44
🤗 Upvotes: 95 | cs. CV, cs. AI, cs. CL Authors: Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu, Song Han Title: Scaling RL to Long Videos Arxiv: http://arxiv.org/abs/2507.07966v1 Abstract: We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to lon...
T-LoRA: Single Image Diffusion Model Customization Without Overfitting 12.07.2025 23:15
🤗 Upvotes: 83 | cs. CV Authors: Vera Soboleva, Aibek Alanov, Andrey Kuznetsov, Konstantin Sobolev Title: T-LoRA: Single Image Diffusion Model Customization Without Overfitting Arxiv: http://arxiv.org/abs/2507.05964v1 Abstract: While diffusion model fine-tuning offers a powerful approach for customizing pre-trained models to generate specific objects, it frequently suffers from overfitting when tr...
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 12.07.2025 20:07
🤗 Upvotes: 37 | cs. CV, cs. AI, cs. CL Authors: Haochen Wang, Xiangtai Li, Zilong Huang, Anran Wang, Jiacong Wang, Tao Zhang, Jiani Zheng, Sule Bai, Zijian Kang, Jiashi Feng, Zhuochen Wang, Zhaoxiang Zhang Title: Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology Arxiv: http://arxiv.org/abs/2507.07999v1 Abstract: Models like OpenAI-o3 pioneer visual grounded reasoni...
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding 12.07.2025 22:58
🤗 Upvotes: 29 | cs. CV Authors: JingLi Lin, Chenming Zhu, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang, Jiangmiao Pang Title: OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding Arxiv: http://arxiv.org/abs/2507.07984v1 Abstract: Recent advances in multimodal large language models (MLLMs) have shown remarkable capabilities in integrating vision and language...
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 12.07.2025 22:47
🤗 Upvotes: 24 | cs. CV, cs. AI Authors: Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Joon-Young Lee, Seon Joo Kim, Minho Shim Title: Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Arxiv: http://arxiv.org/abs/2507.07990v1 Abstract: Video large language models (LLMs) achieve strong video understanding by leveraging a large...
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling 12.07.2025 20:50
🤗 Upvotes: 23 | cs. CV, cs. AI Authors: Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, Jiang Bian Title: Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Arxiv: http://arxiv.org/abs/2507.07982v1 Abstract: Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models tr...
PyVision: Agentic Vision with Dynamic Tooling 12.07.2025 19:03
🤗 Upvotes: 22 | cs. CL, cs. AI, cs. CV Authors: Shitian Zhao, Haoquan Zhang, Shaoheng Lin, Ming Li, Qilong Wu, Kaipeng Zhang, Chen Wei Title: PyVision: Agentic Vision with Dynamic Tooling Arxiv: http://arxiv.org/abs/2507.07998v1 Abstract: LLMs are increasingly deployed as agents, systems capable of planning, reasoning, and dynamically calling external tools. However, in visual reasoning, prior ap...
4KAgent: Agentic Any Image to 4K Super-Resolution 11.07.2025 26:45
🤗 Upvotes: 56 | cs. CV, eess. IV Authors: Yushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang, Renjie Li, Jian Wang, Yide Zhang, Gengchen Mai, Lihong V. Wang, James Zou, Xiaoyu Wang, Ming-Hsuan Yang, Zhengzhong Tu Title: 4KAgent: Agentic Any Image to 4K Super-Resolution Arxiv: http://arxiv.org/abs/2507.07105v1 Abstract: We present 4KAgent, a unified agentic super-resolution generalist system designed...
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data 11.07.2025 16:47
🤗 Upvotes: 41 | cs. CV Authors: Ke Fan, Shunlin Lu, Minyue Dai, Runyi Yu, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, Jingbo Wang Title: Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data Arxiv: http://arxiv.org/abs/2507.07095v1 Abstract: Generating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging rese...
Perception-Aware Policy Optimization for Multimodal Reasoning 11.07.2025 23:28
🤗 Upvotes: 34 | cs. CL Authors: Zhenhailong Wang, Xuehang Guo, Sofia Stoica, Haiyang Xu, Hongru Wang, Hyeonjeong Ha, Xiusi Chen, Yangyi Chen, Ming Yan, Fei Huang, Heng Ji Title: Perception-Aware Policy Optimization for Multimodal Reasoning Arxiv: http://arxiv.org/abs/2507.06448v1 Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endow...
MIRIX: Multi-Agent Memory System for LLM-Based Agents 11.07.2025 21:31
🤗 Upvotes: 33 | cs. CL, cs. AI Authors: Yu Wang, Xi Chen Title: MIRIX: Multi-Agent Memory System for LLM-Based Agents Arxiv: http://arxiv.org/abs/2507.07957v1 Abstract: Although memory capabilities of AI agents are gaining increasing attention, existing solutions remain fundamentally limited. Most rely on flat, narrowly scoped memory components, constraining their ability to personalize, abstract...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.