Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Prompt Orchestration Markup Language 21.08.2025

🤗 Upvotes: 24 | cs. HC, cs. AI, cs. CL, cs. PL Authors: Yuge Zhang, Nan Chen, Jiahang Xu, Yuqing Yang Title: Prompt Orchestration Markup Language Arxiv: http://arxiv.org/abs/2508.13948v1 Abstract: Large Language Models (LLMs) require sophisticated prompting, yet current practices face challenges in structure, data integration, format sensitivity, and tooling. Existing methods lack comprehensive s...

Ovis2.5 Technical Report 20.08.2025

🤗 Upvotes: 79 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Shiyin Lu, Yang Li, Yu Xia, Yuwei Hu, Shanshan Zhao, Yanqing Ma, Zhichao Wei, Yinglun Li, Lunhao Duan, Jianshan Zhao, Yuxuan Han, Haijun Li, Wanying Chen, Junke Tang, Chengkun Hou, Zhixing Du, Tianli Zhou, Wenjie Zhang, Huping Ding, Jiahe Li, Wen Li, Gui Hu, Yiliang Gu, Siran Yang, Jiamang Wang, Hailong Sun, Yibo Wang, Hui Sun, Jinlong Huang...

ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning 20.08.2025

🤗 Upvotes: 47 | cs. CL, cs. AI, cs. LG Authors: Juyuan Wang, Rongchen Zhao, Wei Wei, Yufeng Wang, Mo Yu, Jie Zhou, Jin Xu, Liyan Xu Title: ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning Arxiv: http://arxiv.org/abs/2508.10419v1 Abstract: Narrative comprehension on long stories and novels has been a challenging domain attributed to their intricate plotlines...

4DNeX: Feed-Forward 4D Generative Modeling Made Easy 20.08.2025

🤗 Upvotes: 44 | cs. CV Authors: Zhaoxi Chen, Tianqi Liu, Long Zhuo, Jiawei Ren, Zeng Tao, He Zhu, Fangzhou Hong, Liang Pan, Ziwei Liu Title: 4DNeX: Feed-Forward 4D Generative Modeling Made Easy Arxiv: http://arxiv.org/abs/2508.13154v1 Abstract: We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing...

Next Visual Granularity Generation 20.08.2025

🤗 Upvotes: 37 | cs. CV, cs. AI, cs. LG Authors: Yikai Wang, Zhouxia Wang, Zhonghua Wu, Qingyi Tao, Kang Liao, Chen Change Loy Title: Next Visual Granularity Generation Arxiv: http://arxiv.org/abs/2508.12811v1 Abstract: We propose a novel approach to image generation by decomposing an image into a structured sequence, where each element in the sequence shares the same spatial resolution but differ...

Speed Always Wins: A Survey on Efficient Architectures for Large Language Models 20.08.2025

🤗 Upvotes: 34 | cs. CL, cs. AI, cs. CV Authors: Weigao Sun, Jiaxi Hu, Yucheng Zhou, Jusen Du, Disen Lan, Kexin Wang, Tong Zhu, Xiaoye Qu, Yu Zhang, Xiaoyu Mo, Daizong Liu, Yuxuan Liang, Wenliang Chen, Guoqi Li, Yu Cheng Title: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models Arxiv: http://arxiv.org/abs/2508.09834v1 Abstract: Large Language Models (LLMs) have delive...

When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs 20.08.2025

🤗 Upvotes: 32 | cs. CL, cs. AI Authors: Mikhail Seleznyov, Mikhail Chaichuk, Gleb Ershov, Alexander Panchenko, Elena Tutubalina, Oleg Somov Title: When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs Arxiv: http://arxiv.org/abs/2508.11383v1 Abstract: Large Language Models (LLMs) are highly sensitive to subtle, non-semantic variations in prompt phrasing and form...

Has GPT-5 Achieved Spatial Intelligence? An Empirical Study 20.08.2025

🤗 Upvotes: 22 | cs. CV, cs. CL, cs. LG, cs. MM, cs. RO Authors: Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Jiaqi Li, Xiangyu Fan, Hanming Deng, Lewei Lu, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang Title: Has GPT-5 Achieved Spatial Intelligence? An Empirical Study Arxiv: htt...

HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds 20.08.2025

🤗 Upvotes: 21 | cs. AI Authors: Petr Anokhin, Roman Khalikov, Stefan Rebrikov, Viktor Volkov, Artyom Sorokin, Vincent Bissonnette Title: HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds Arxiv: http://arxiv.org/abs/2508.12782v1 Abstract: Large language models (LLMs) have shown remarkable capabilities in isolated step-by-step reasoning tasks such as mathem...

SSRL: Self-Search Reinforcement Learning 19.08.2025

🤗 Upvotes: 66 | cs. CL Authors: Yuchen Fan, Kaiyan Zhang, Heng Zhou, Yuxin Zuo, Yanxu Chen, Yu Fu, Xinwei Long, Xuekai Zhu, Che Jiang, Yuchen Zhang, Li Kang, Gang Chen, Cheng Huang, Zhizhou He, Bingning Wang, Lei Bai, Ning Ding, Bowen Zhou Title: SSRL: Self-Search Reinforcement Learning Arxiv: http://arxiv.org/abs/2508.10874v1 Abstract: We investigate the potential of large language models (LLMs)...

DINOv3 19.08.2025

🤗 Upvotes: 65 | cs. CV, cs. LG Authors: Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Ma...

Thyme: Think Beyond Images 19.08.2025

🤗 Upvotes: 59 | cs. CV Authors: Yi-Fan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu, Wei Chen, Xiao Hu, Bin Wen, Kaiyu Jiang, Changyi Liu, Tianke Zhang, Haonan Fan, Kaibing Chen, Jiankang Chen, Haojie Ding, Kaiyu Tang, Zhang Zhang, Liang Wang, Fan Yang, Tingting Gao, Guorui Zhou Title: Thyme: Think Beyond Images Arxiv: http://arxiv.org/abs/2508.11630v1 Abstract: Following OpenAI's introduction of th...

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining 19.08.2025

🤗 Upvotes: 40 | cs. LG, cs. CL Authors: Pratyush Maini, Vineeth Dorna, Parth Doshi, Aldo Carranza, Fan Pan, Jack Urbanek, Paul Burstein, Alex Fang, Alvin Deng, Amro Abbas, Brett Larsen, Cody Blakeney, Charvi Bannur, Christina Baek, Darren Teh, David Schwab, Haakon Mongstad, Haoli Yin, Josh Wills, Kaleigh Mentzer, Luke Merrick, Ricardo Monti, Rishabh Adiga, Siddharth Joshi, Spandan Das, Zhengping...

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization 19.08.2025

🤗 Upvotes: 24 | cs. LG Authors: Aditya Tomar, Coleman Hooper, Minjae Lee, Haocheng Xi, Rishabh Tiwari, Wonjun Kang, Luca Manolache, Michael W. Mahoney, Kurt Keutzer, Amir Gholami Title: XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Arxiv: http://arxiv.org/abs/2508.10395v1 Abstract: Although LLM inference has emerged as a critical workload for many downstream a...

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning 16.08.2025

🤗 Upvotes: 117 | cs. AI, cs. CV, cs. LG Authors: Runqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang, Xiaowan Wang, Enhui Wan, Sitong Zhou, Guanting Dong, Yuchen Zeng, Yida Xu, Jie Wang, Chong Sun, Chen Li, Honggang Zhang Title: We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning Arxiv: http://arxiv.org/abs/2508.10433v1 Abstract: Multimodal Large Language Models (...

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale 16.08.2025

🤗 Upvotes: 101 | cs. CV Authors: NextStep Team, Chunrui Han, Guopeng Li, Jingwei Wu, Quan Sun, Yan Cai, Yuang Peng, Zheng Ge, Deyu Zhou, Haomiao Tang, Hongyu Zhou, Kenkun Liu, Ailin Huang, Bin Wang, Changxin Miao, Deshan Sun, En Yu, Fukun Yin, Gang Yu, Hao Nie, Haoran Lv, Hanpeng Hu, Jia Wang, Jian Zhou, Jianjian Sun, Kaijun Tan, Kang An, Kangheng Lin, Liang Zhao, Mei Chen, Peng Xing, Rui Wang, S...

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts 16.08.2025

🤗 Upvotes: 33 | cs. CL, cs. AI Authors: Mo Yu, Tsz Ting Chung, Chulun Zhou, Tong Li, Rui Lu, Jiangnan Li, Liyan Xu, Haoshu Lu, Ning Zhang, Jing Li, Jie Zhou Title: PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts Arxiv: http://arxiv.org/abs/2508.09848v2 Abstract: We introduce PRELUDE, a benchmark for evaluating long-context understanding through the t...

ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing 16.08.2025

🤗 Upvotes: 33 | cs. CV, cs. AI Authors: Lingen Li, Guangzhi Wang, Zhaoyang Zhang, Yaowei Li, Xiaoyu Li, Qi Dou, Jinwei Gu, Tianfan Xue, Ying Shan Title: ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing Arxiv: http://arxiv.org/abs/2508.10881v1 Abstract: Traditional cartoon and anime production involves keyframing, inbetweening, and colorization stages, which require in...

Story2Board: A Training-Free Approach for Expressive Storyboard Generation 15.08.2025

🤗 Upvotes: 42 | cs. CV, cs. GR, cs. LG Authors: David Dinkevich, Matan Levy, Omri Avrahami, Dvir Samuel, Dani Lischinski Title: Story2Board: A Training-Free Approach for Expressive Storyboard Generation Arxiv: http://arxiv.org/abs/2508.09983v1 Abstract: We present Story2Board, a training-free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on...

Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery 15.08.2025

🤗 Upvotes: 29 | cs. CL Authors: Jiatong Li, Weida Wang, Qinggang Zhang, Junxian Li, Di Zhang, Changmeng Zheng, Shufei Zhang, Xiaoyong Wei, Qing Li Title: Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery Arxiv: http://arxiv.org/abs/2508.08401v1 Abstract: Large language models (LLMs), especially Explicit Long Chain-of-Thought (CoT) reasoning models like DeepSeek-R1 and QWQ, have de...

Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation 15.08.2025

🤗 Upvotes: 28 | cs. CV Authors: Bowen Xue, Qixin Yan, Wenjing Wang, Hao Liu, Chen Li Title: Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation Arxiv: http://arxiv.org/abs/2508.07901v2 Abstract: Generating high-fidelity human videos that match user-specified identities is important yet challenging in the field of generative AI. Existing methods often rely on an excessi...

Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing 15.08.2025

🤗 Upvotes: 23 | cs. LG, cs. AI Authors: Xu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin, Hao Zhang, Zhijie Deng Title: Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing Arxiv: http://arxiv.org/abs/2508.09192v1 Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs for text generation, with the potential to deco...

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory 15.08.2025

🤗 Upvotes: 22 | cs. CV Authors: Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, Wei Li Title: Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory Arxiv: http://arxiv.org/abs/2508.09736v1 Abstract: We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visua...

AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving 15.08.2025

🤗 Upvotes: 22 | cs. AI Authors: Zhitian Xie, Qintong Wu, Chengyue Yu, Chenyi Zhuang, Jinjie Gu Title: AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving Arxiv: http://arxiv.org/abs/2508.09889v1 Abstract: The rapid advancement of large language models (LLMs) has empowered intelligent agents to leverage diverse external tools for solving complex real-world pr...

Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL 14.08.2025

🤗 Upvotes: 34 | cs. CL, cs. AI Authors: Jiaxuan Gao, Wei Fu, Minyang Xie, Shusheng Xu, Chuyi He, Zhiyu Mei, Banghua Zhu, Yi Wu Title: Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL Arxiv: http://arxiv.org/abs/2508.07976v2 Abstract: Recent advancements in LLM-based agents have demonstrated remarkable capabilities in handling complex, knowledge-intensive ta...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.