Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Soft Adaptive Policy Optimization 27.11.2025 24:08
🤗 Upvotes: 23 | cs. LG, cs. AI, cs. CL Authors: Chang Gao, Chujie Zheng, Xiong-Hui Chen, Kai Dang, Shixuan Liu, Bowen Yu, An Yang, Shuai Bai, Jingren Zhou, Junyang Lin Title: Soft Adaptive Policy Optimization Arxiv: http://arxiv.org/abs/2511.20347v1 Abstract: Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), y...
General Agentic Memory Via Deep Research 26.11.2025 25:36
🤗 Upvotes: 121 | cs. CL, cs. AI, cs. IR, cs. LG Authors: B. Y. Yan, Chaofan Li, Hongjin Qian, Shuqi Lu, Zheng Liu Title: General Agentic Memory Via Deep Research Arxiv: http://arxiv.org/abs/2511.18423v1 Abstract: Memory is critical for AI agents, yet the widely-adopted static memory, aiming to create readily available memory in advance, is inevitably subject to severe information loss. To address...
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning 26.11.2025 23:00
🤗 Upvotes: 79 | cs. AI, cs. CL, cs. LG Authors: Jiayi Zhang, Yiran Peng, Fanqi Kong, Yang Cheng, Yifan Wu, Zhaoyang Yu, Jinyu Xiang, Jianhao Ruan, Jinlin Wang, Maojia Song, HongZhang Liu, Xiangru Tang, Bang Liu, Chenglin Wu, Yuyu Luo Title: AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning Arxiv: http://arxiv.org/abs/2511.19304v1 Abstract: Humans naturally adapt to di...
Computer-Use Agents as Judges for Generative User Interface 26.11.2025 25:50
🤗 Upvotes: 47 | cs. CV, cs. CL, cs. HC Authors: Kevin Qinghong Lin, Siyuan Hu, Linjie Li, Zhengyuan Yang, Lijuan Wang, Philip Torr, Mike Zheng Shou Title: Computer-Use Agents as Judges for Generative User Interface Arxiv: http://arxiv.org/abs/2511.15567v1 Abstract: Computer-Use Agents (CUA) are becoming increasingly capable of autonomously operating digital environments through Graphical User Int...
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation 26.11.2025 25:10
🤗 Upvotes: 44 | cs. CV, cs. AI Authors: Zehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang, Qi Tian Title: DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation Arxiv: http://arxiv.org/abs/2511.19365v1 Abstract: Pixel diffusion aims to generate images directly in pixel space in an end-to-end fashion. This approach avoids the limitations of VAE in the two-stage latent diffusion...
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research 26.11.2025 20:31
🤗 Upvotes: 42 | cs. CL, cs. AI, cs. LG Authors: Rulin Shao, Akari Asai, Shannon Zejiang Shen, Hamish Ivison, Varsha Kishore, Jingming Zhuo, Xinran Zhao, Molly Park, Samuel G. Finlayson, David Sontag, Tyler Murray, Sewon Min, Pradeep Dasigi, Luca Soldaini, Faeze Brahman, Wen-tau Yih, Tongshuang Wu, Luke Zettlemoyer, Yoon Kim, Hannaneh Hajishirzi, Pang Wei Koh Title: DR Tulu: Reinforcement Learning...
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios 26.11.2025 22:20
🤗 Upvotes: 32 | cs. CV Authors: Tian Ye, Song Fei, Lei Zhu Title: UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios Arxiv: http://arxiv.org/abs/2511.18050v1 Abstract: Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect r...
In-Video Instructions: Visual Signals as Generative Control 26.11.2025 22:26
🤗 Upvotes: 26 | cs. CV, cs. AI Authors: Gongfan Fang, Xinyin Ma, Xinchao Wang Title: In-Video Instructions: Visual Signals as Generative Control Arxiv: http://arxiv.org/abs/2511.19401v1 Abstract: Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in the current observation...
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe 25.11.2025 21:24
🤗 Upvotes: 76 | cs. AI, cs. CL Authors: Kaichen Zhang, Keming Wu, Zuhao Yang, Kairui Hu, Bin Wang, Ziwei Liu, Xingxuan Li, Lidong Bing Title: OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe Arxiv: http://arxiv.org/abs/2511.16334v1 Abstract: Recent advancements in large reasoning models have fueled growing interest in extending such capabilities to mu...
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story 25.11.2025 22:13
🤗 Upvotes: 72 | cs. CL, cs. AI, cs. LG Authors: Vladislav Pedashenko, Laida Kushnareva, Yana Khassan Nibal, Eduard Tulchinskii, Kristian Kuznetsov, Vladislav Zharchinskii, Yury Maximov, Irina Piontkovskaya Title: Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story Arxiv: http://arxiv.org/abs/2511.15210v1 Abstract: Intrinsic dimension (ID) is an important tool in moder...
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization 25.11.2025 21:40
🤗 Upvotes: 60 | cs. CV Authors: Yikun Wang, Zuyan Liu, Ziyi Wang, Pengfei Liu, Han Hu, Yongming Rao Title: GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization Arxiv: http://arxiv.org/abs/2511.15705v1 Abstract: Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purp...
SAM 3: Segment Anything with Concepts 25.11.2025 23:53
🤗 Upvotes: 51 | cs. CV, cs. AI Authors: Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Liliane Mo...
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks 21.11.2025 25:59
🤗 Upvotes: 66 | cs. CV, cs. AI Authors: Cheng Yang, Haiyuan Wan, Yiran Peng, Xin Cheng, Zhaoyang Yu, Jiayi Zhang, Junchi Yu, Xinlei Yu, Xiawu Zheng, Dongzhan Zhou, Chenglin Wu Title: Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks Arxiv: http://arxiv.org/abs/2511.15065v1 Abstract: Video Models have achieved remarkable success in high-fidel...
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation 21.11.2025 24:57
🤗 Upvotes: 128 | cs. CV, cs. AI, cs. LG Authors: Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko, Denis Parkhomenko, Viacheslav Vasilev, Alexey Letunovskiy, Maria Kovaleva, Nikolai Vaulin, Ivan Kirillov, Lev Novitskiy, Denis Koposov, Nikita Kiselev, Alexander Varlamov, Dmitrii Mikhailov, Vladimir Polovnikov, Andrey Shutkin, Ilya Vasiliev, Julia Agafonova, Anastasiia Kargapoltseva, Ann...
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity 21.11.2025 22:40
🤗 Upvotes: 47 | cs. AI Authors: Alexis Audran-Reiss, Jordi Armengol Estapé, Karen Hambardzumyan, Amar Budhiraja, Martin Josifoski, Edan Toledo, Rishi Hazra, Despoina Magka, Michael Shvartsman, Parth Pathak, Justine T Kao, Lucia Cipolina-Kun, Bhavul Gauri, Jean-Christophe Gagnon-Audet, Emanuel Tewolde, Jenny Zhang, Taco Cohen, Yossi Adi, Tatiana Shavrina, Yoram Bachrach Title: What Does It Take to...
VisPlay: Self-Evolving Vision-Language Models from Images 21.11.2025 22:28
🤗 Upvotes: 31 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, Yonghui Yang Title: VisPlay: Self-Evolving Vision-Language Models from Images Arxiv: http://arxiv.org/abs/2511.15661v2 Abstract: Reinforcement learning (RL) provides a principled framework for improving Vision-Language Models (VLMs) on complex reasoning tasks. However, existing RL approa...
Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset 21.11.2025 19:22
🤗 Upvotes: 23 | cs. CV Authors: Geon Choi, Hangyul Yoon, Hyunju Shin, Hyunki Park, Sang Hoon Seo, Eunho Yang, Edward Choi Title: Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset Arxiv: http://arxiv.org/abs/2511.15186v1 Abstract: The applicability of current lesion segmentation models for chest X-rays (CXRs) has been limited both by a small n...
VIDEOP2R: Video Understanding from Perception to Reasoning 20.11.2025 25:08
🤗 Upvotes: 70 | cs. CV, cs. AI, cs. LG Authors: Yifan Jiang, Yueying Wang, Rui Zhao, Toufiq Parag, Zhimin Chen, Zhenyu Liao, Jayakrishnan Unnikrishnan Title: VIDEOP2R: Video Understanding from Perception to Reasoning Arxiv: http://arxiv.org/abs/2511.11113v1 Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) h...
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models 20.11.2025 24:58
🤗 Upvotes: 66 | cs. CL, cs. AI, cs. LG, cs. PF Authors: Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang Title: Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models Arxiv: http://arxiv.org/abs/2511.08577v1 Abstract: Improving reasoning capabilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world app...
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models 20.11.2025 23:48
🤗 Upvotes: 58 | cs. CL, cs. AI, cs. LG Authors: Mohammad Zbib, Hasan Abed Al Kader Hammoud, Sina Mukalled, Nadine Rizk, Fatima Karnib, Issam Lakkis, Ammar Mohanna, Bernard Ghanem Title: AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models Arxiv: http://arxiv.org/abs/2511.14295v1 Abstract: We present AraLingBench: a fully human annotated b...
A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space 20.11.2025 23:48
🤗 Upvotes: 41 | cs. CV, cs. AI Authors: Huijie Liu, Shuhao Cui, Haoxiang Cao, Shuai Ma, Kai Wu, Guoliang Kang Title: A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space Arxiv: http://arxiv.org/abs/2511.10555v4 Abstract: Innovative visual stylization is a cornerstone of artistic creation, yet generating novel and consistent visual styles remains a signific...
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark 20.11.2025 22:39
🤗 Upvotes: 32 | cs. CV Authors: Xinxin Liu, Zhaopan Xu, Kai Wang, Yong Jae Lee, Yuzhang Shang Title: Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark Arxiv: http://arxiv.org/abs/2511.13853v1 Abstract: While Chain-of-Thought (CoT) prompting enables sophisticated symbolic reasoning in LLMs, it remains confined to discrete text and cannot simulate the continuous, physic...
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs 20.11.2025 24:27
🤗 Upvotes: 24 | cs. CV Authors: Huiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun, Kaijie Chen, Yan Yan, Xu Yang, Lu Cheng Title: MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Arxiv: http://arxiv.org/abs/2511.14159v1 Abstract: Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and re...
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding 20.11.2025 26:47
🤗 Upvotes: 22 | cs. CV Authors: Jiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen, Boshen Xu, Yuxun Qu, Yijing Chen, Jianzhong Ju, Zhenbo Luo, Jian Luan Title: REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding Arxiv: http://arxiv.org/abs/2511.13026v1 Abstract: Self-reflection mechanisms that rely on purely text-based rethinking processes pe...
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data 19.11.2025 24:24
🤗 Upvotes: 87 | cs. CL, cs. AI, cs. CV Authors: Yunxin Li, Xinyu Chen, Shenyuan Jiang, Haoyuan Shi, Zhenyu Liu, Xuanyu Zhang, Nanhao Deng, Zhenran Xu, Yicheng Ma, Meishan Zhang, Baotian Hu, Min Zhang Title: Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data Arxiv: http://arxiv.org/abs/2511.12609v1 Abstract: We present Uni-MoE 2.0 from the Lychee...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.