Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
P1: Mastering Physics Olympiads with Reinforcement Learning 19.11.2025 22:16
🤗 Upvotes: 107 | cs. LG, cs. AI, cs. CL Authors: Jiacheng Chen, Qianjia Cheng, Fangchen Yu, Haiyuan Wan, Yuchen Zhang, Shenghe Zheng, Junchi Yao, Qingyang Zhang, Haonan He, Yun Luo, Yufeng Zhao, Futing Wang, Li Sheng, Chengxing Xie, Yuxin Zuo, Yizhuo Li, Wenxauan Zeng, Yulun Wu, Rui Huang, Dongzhan Zhou, Kai Chen, Yu Qiao, Lei Bai, Yu Cheng, Ning Ding, Bowen Zhou, Peng Ye, Ganqu Cui Title: P1: Ma...
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling 19.11.2025 27:44
🤗 Upvotes: 104 | cs. CL Authors: MiroMind Team, Song Bai, Lidong Bing, Carson Chen, Guanzheng Chen, Yuntao Chen, Zhe Chen, Ziyi Chen, Jifeng Dai, Xuan Dong, Wenhan Dou, Yue Deng, Yunjie Fu, Junqi Ge, Chenxia Han, Tammy Huang, Zhenhang Huang, Jerry Jiao, Shilei Jiang, Tianyu Jiao, Xiaoqi Jian, Lei Lei, Ruilin Li, Ryan Luo, Tiantong Li, Xiang Lin, Ziyuan Liu, Zhiqi Li, Jie Ni, Qiang Ren, Pax Sun, S...
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance 19.11.2025 23:57
🤗 Upvotes: 75 | cs. CL Authors: Shalini Maiti, Amar Budhiraja, Bhavul Gauri, Gaurav Chaurasia, Anton Protopopov, Alexis Audran-Reiss, Michael Slater, Despoina Magka, Tatiana Shavrina, Roberta Raileanu, Yoram Bachrach Title: Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance Arxiv: http://arxiv.org/abs/2511.13254v1 Abstract: Large Language Models (LLMs) have demonstrated...
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model 19.11.2025 25:57
🤗 Upvotes: 63 | cs. CV Authors: Chunshi Wang, Junliang Ye, Yunhan Yang, Yang Li, Zizhuo Lin, Jun Zhu, Zhuo Chen, Yawei Luo, Chunchao Guo Title: Part-X-MLLM: Part-aware 3D Multimodal Large Language Model Arxiv: http://arxiv.org/abs/2511.13647v1 Abstract: We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structur...
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation 19.11.2025 20:43
🤗 Upvotes: 47 | cs. CV Authors: Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang, Yu Tian, Jiani Zheng, Haochen Wang, Zhiyang Teng, Zhuochen Wang, Yinjie Wang, Yunhai Tong, Mengdi Wang, Xiangtai Li Title: MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation Arxiv: http://arxiv.org/abs/2511.09611v3 Abstract: While thinking-aware generation aims to impro...
GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning 19.11.2025 23:49
🤗 Upvotes: 46 | cs. IR, cs. AI, cs. LG Authors: Duolin Sun, Meixiu Long, Dan Yang, Yihan Jiao, Zhehao Tan, Jie Feng, Junjie Wang, Yue Shen, Peng Wei, Jian Wang, Jinjie Gu Title: GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning Arxiv: http://arxiv.org/abs/2511.11653v1 Abstract: Large Language Models have shown strong potential as rerankers to enhance the overall performan...
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models 19.11.2025 23:11
🤗 Upvotes: 40 | cs. CV Authors: Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu, Zihan Wang, Sirui Chen, Wenkai Cheng, Kanghao Chen, Hongfei Zhang, Zixin Zhang, Rongjin Guo, Yu Cheng, Ying-Cong Chen Title: TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models Arxiv: http://arxiv.org/abs/2511.13704v1 Abstract: The rapid evolution of video generative models has shif...
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image 19.11.2025 25:25
🤗 Upvotes: 31 | cs. CV, cs. RO Authors: Ziang Cao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, Ziwei Liu Title: PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image Arxiv: http://arxiv.org/abs/2511.13648v1 Abstract: 3D modeling is shifting from static visual representations toward physical, articulated assets that can be directly used in simulation and interaction. However, most exist...
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models 18.11.2025 21:58
🤗 Upvotes: 30 | cs. AI Authors: Jingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu, Siyuan Li, Linzhuang Sun, Bihui Yu, Conghui He, Lijun Wu, Cheng Tan Title: GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models Arxiv: http://arxiv.org/abs/2511.11134v1 Abstract: The advent of Unified Multimodal Models (UMMs) signals a paradigm shift in artificial intelligence, moving from...
DoPE: Denoising Rotary Position Embedding 18.11.2025 19:24
🤗 Upvotes: 64 | cs. CL Authors: Jing Xiong, Liyang Fan, Hui Shen, Zunhai Su, Min Yang, Lingpeng Kong, Ngai Wong Title: DoPE: Denoising Rotary Position Embedding Arxiv: http://arxiv.org/abs/2511.09146v1 Abstract: Rotary Position Embedding (RoPE) in Transformer models has inherent limits that weaken length extrapolation. We reinterpret the attention map with positional encoding as a noisy feature m...
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation 18.11.2025 25:25
🤗 Upvotes: 41 | cs. CV Authors: Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou, Xue Song, Liyu Jia, Saining Zhang, Siliang Tang, Juncheng Li, Fengda Zhang, Weijia Wu, Hanwang Zhang, Tat-Seng Chua Title: WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation Arxiv: http://arxiv.org/abs/2511.11434v1 Abstract: Recent advances in unified multimodal models (UMMs...
UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation 18.11.2025 25:46
🤗 Upvotes: 28 | cs. CV Authors: Zhen Yang, Wenyi Hong, Mingde Xu, Xinyue Fan, Weihan Wang, Jiele Cheng, Xiaotao Gu, Jie Tang Title: UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation Arxiv: http://arxiv.org/abs/2511.08195v2 Abstract: User interface (UI) programming is a core yet highly complex part of modern software development. Recent advances in visual...
AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery 18.11.2025 28:52
🤗 Upvotes: 24 | cs. AI, cs. CE, cs. LG Authors: Yuqi Yin, Yibo Fu, Siyuan Wang, Peng Sun, Hongyu Wang, Xiaohui Wang, Lei Zheng, Zhiyong Li, Zhirong Liu, Jianji Wang, Zhaoxi Sun Title: AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery Arxiv: http://arxiv.org/abs/2511.11257v1 Abstract: The discovery of novel Ionic Liquids (ILs) is hindered by critical challenges...
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers 18.11.2025 21:16
🤗 Upvotes: 23 | cs. CV, cs. AI Authors: Dor Shmilovich, Tony Wu, Aviad Dahan, Yuval Domb Title: LiteAttention: A Temporal Sparse Attention for Diffusion Transformers Arxiv: http://arxiv.org/abs/2511.11062v1 Abstract: Diffusion Transformers, particularly for video generation, achieve remarkable quality but suffer from quadratic attention complexity, leading to prohibitive latency. Existing acceler...
Virtual Width Networks 18.11.2025 22:36
🤗 Upvotes: 23 | cs. LG, cs. AI Authors: Seed, Baisheng Li, Banggu Wu, Bole Ma, Bowen Xiao, Chaoyi Zhang, Cheng Li, Chengyi Wang, Chenyin Xu, Chi Zhang, Chong Hu, Daoguang Zan, Defa Zhu, Dongyu Xu, Du Li, Faming Wu, Fan Xia, Ge Zhang, Guang Shi, Haobin Chen, Hongyu Zhu, Hongzhi Huang, Huan Zhou, Huanzhang Dou, Jianhui Duan, Jianqiao Lu, Jianyu Jiang, Jiayi Xu, Jiecao Chen, Jin Chen, Jin Ma, Jing S...
One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models 15.11.2025 22:00
🤗 Upvotes: 55 | cs. CV Authors: Aleksandr Razin, Danil Kazantsev, Ilya Makarov Title: One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models Arxiv: http://arxiv.org/abs/2511.10629v1 Abstract: Diffusion models struggle to scale beyond their training resolutions, as direct high-resolution sampling is slow and costly, while post-hoc image super-res...
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation 15.11.2025 25:56
🤗 Upvotes: 32 | cs. CV, cs. AI, cs. CL, cs. LG Authors: PAN Team, Jiannan Xiang, Yi Gu, Zihan Liu, Zeyu Feng, Qiyue Gao, Yiyan Hu, Benhao Huang, Guangyi Liu, Yichi Yang, Kun Zhou, Davit Abrahamyan, Arif Ahmad, Ganesh Bannur, Junrong Chen, Kimi Chen, Mingkai Deng, Ruobing Han, Xinqi Huang, Haoqiang Kang, Zheqi Li, Enze Ma, Hector Ren, Yashowardhan Shinde, Rohan Shingre, Ramsundar Tanikella, Kaimin...
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist 15.11.2025 27:35
🤗 Upvotes: 26 | cs. CV Authors: Zhengyang Liang, Daoan Zhang, Huichi Zhou, Rui Huang, Bobo Li, Yuechen Zhang, Shengqiong Wu, Xiaohan Wang, Jiebo Luo, Lizi Liao, Hao Fei Title: UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist Arxiv: http://arxiv.org/abs/2511.08521v1 Abstract: While specialized AI models excel at isolated video tasks like generation or understanding...
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains 11.11.2025 25:22
🤗 Upvotes: 34 | cs. CL, cs. AI Authors: Zihao Yi, Qingxuan Jiang, Ruotian Ma, Xingyu Chen, Qu Yang, Mengru Wang, Fanghua Ye, Ying Shen, Zhaopeng Tu, Xiaolong Li, Linus Title: Too Good to be Bad: On the Failure of LLMs to Role-Play Villains Arxiv: http://arxiv.org/abs/2511.04962v1 Abstract: Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of f...
DeepEyesV2: Toward Agentic Multimodal Model 11.11.2025 26:21
🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Jack Hong, Chenxiao Zhao, ChengLin Zhu, Weiheng Lu, Guohai Xu, Xing Yu Title: DeepEyesV2: Toward Agentic Multimodal Model Arxiv: http://arxiv.org/abs/2511.05271v1 Abstract: Agentic multimodal models should not only comprehend text and images, but also actively invoke external tools, such as code execution environments and web search, and integrate these ope...
Visual Spatial Tuning 11.11.2025 24:14
🤗 Upvotes: 30 | cs. CV Authors: Rui Yang, Ziyu Zhu, Yanwei Li, Jingjia Huang, Shen Yan, Siyuan Zhou, Zhe Liu, Xiangtai Li, Shuangye Li, Wenqian Wang, Yi Lin, Hengshuang Zhao Title: Visual Spatial Tuning Arxiv: http://arxiv.org/abs/2511.05491v1 Abstract: Capturing spatial relationships from visual inputs is a cornerstone of human-like general intelligence. Several previous studies have tried to en...
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks 11.11.2025 22:55
🤗 Upvotes: 26 | cs. AI, cs. CL Authors: Yu Feng, Nathaniel Weir, Kaj Bostrom, Sam Bayless, Darion Cassel, Sapana Chaudhary, Benjamin Kiesl-Reiter, Huzefa Rangwala Title: VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks Arxiv: http://arxiv.org/abs/2511.04662v1 Abstract: LLMs can perform multi-step reasoning through Chain-of-Thought (CoT), but they cannot reliably...
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 08.11.2025 24:22
🤗 Upvotes: 127 | cs. CV, cs. CL Authors: Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, Xinchi Chen, Jun Zhao, Xuanjing Huang, Xipeng Qiu Title: Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Arxiv: http://arxiv.org/abs/2511.04570v1 Abstract: "Thinking with Text" and "Thinking...
V-Thinker: Interactive Thinking with Images 08.11.2025 21:01
🤗 Upvotes: 66 | cs. CV Authors: Runqi Qiao, Qiuna Tan, Minghan Yang, Guanting Dong, Peiqing Yang, Shiqiang Lang, Enhui Wan, Xiaowan Wang, Yida Xu, Lan Yang, Chong Sun, Chen Li, Honggang Zhang Title: V-Thinker: Interactive Thinking with Images Arxiv: http://arxiv.org/abs/2511.04460v1 Abstract: Empowering Large Multimodal Models (LMMs) to deeply integrate image interaction with long-horizon reasoni...
Scaling Agent Learning via Experience Synthesis 08.11.2025 23:00
🤗 Upvotes: 51 | cs. AI Authors: Zhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu, Qi Qi, Yifan Wu, Tarun Kalluri, Sara Cao, Yuanhao Xiong, Haibo Tong, Huaxiu Yao, Hengduo Li, Jiacheng Zhu, Xian Li, Dawn Song, Bo Li, Jason Weston, Dat Huynh Title: Scaling Agent Learning via Experience Synthesis Arxiv: http://arxiv.org/abs/2511.03773v1 Abstract: While reinforcement learning (RL) can empower large lang...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.