Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents 26.03.2026

🤗 Upvotes: 43 | cs. AI, cs. CL Authors: Ling Yue, Kushal Raj Bhandari, Ching-Yun Ko, Dhaval Patel, Shuxin Lin, Nianjun Zhou, Jianxi Gao, Pin-Yu Chen, Shaowu Pan Title: From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents Arxiv: http://arxiv.org/abs/2603.22386v1 Abstract: Large language model (LLM)-based systems are becoming increasingly popular for sol...

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning 26.03.2026

🤗 Upvotes: 43 | cs. CV, cs. CL Authors: Haoyu Huang, Jinfa Huang, Zhongwei Wan, Xiawu Zheng, Rongrong Ji, Jiebo Luo Title: SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning Arxiv: http://arxiv.org/abs/2603.23483v1 Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities thr...

PEARL: Personalized Streaming Video Understanding Model 26.03.2026

🤗 Upvotes: 36 | cs. CV, cs. AI, cs. IR Authors: Yuanhong Zheng, Ruichuan An, Xiaopeng Lin, Yuxing Liu, Sihan Yang, Huanyu Zhang, Haodong Li, Qintong Zhang, Renrui Zhang, Guopeng Li, Yifan Zhang, Yuheng Li, Wentao Zhang Title: PEARL: Personalized Streaming Video Understanding Model Arxiv: http://arxiv.org/abs/2603.20422v1 Abstract: Human cognition of new concepts is inherently a streaming process:...

DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models 26.03.2026

🤗 Upvotes: 36 | cs. CV Authors: Jaewon Min, Jaeeun Lee, Yeji Choi, Paul Hyunbin Cho, Jin Hyeon Kim, Tae-Young Lee, Jongsik Ahn, Hwayeong Lee, Seonghyun Park, Seungryong Kim Title: DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models Arxiv: http://arxiv.org/abs/2603.23499v1 Abstract: Optical flow models trained on high-quality data often degrade severely when confronted with re...

SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM 26.03.2026

🤗 Upvotes: 33 | cs. CV, cs. GR, cs. RO Authors: Chuanrui Zhang, Minghan Qin, Yuang Wang, Baifeng Xie, Hang Li, Ziwei Wang Title: SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM Arxiv: http://arxiv.org/abs/2603.23386v1 Abstract: High-quality articulated 3D assets are indispensable for embodied AI and physical simulation, yet 3D generation still focuses on static me...

UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation 26.03.2026

🤗 Upvotes: 30 | cs. CV Authors: Jie Liu, Zilyu Ye, Linxiao Yuan, Shenhan Zhu, Yu Gao, Jie Wu, Kunchang Li, Xionghui Wang, Xiaonan Nie, Weilin Huang, Wanli Ouyang Title: UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation Arxiv: http://arxiv.org/abs/2603.23500v1 Abstract: Unified models capable of interleaved generation have emerged as a promising paradigm, with the communi...

RealMaster: Lifting Rendered Scenes into Photorealistic Video 26.03.2026

🤗 Upvotes: 23 | cs. CV Authors: Dana Cohen-Bar, Ido Sobol, Raphael Bensadoun, Shelly Sheynin, Oran Gafni, Or Patashnik, Daniel Cohen-Or, Amit Zohar Title: RealMaster: Lifting Rendered Scenes into Photorealistic Video Arxiv: http://arxiv.org/abs/2603.23462v1 Abstract: State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align gener...

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models 25.03.2026

🤗 Upvotes: 110 | cs. CV Authors: Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng, Rujing Dang, Bingze Song, Ruitian Tian, Jiashu Zhu, Jiachen Lei, Hao Dou, Jing Tang, Lei Sun, Jiahong Wu, Xiangxiang Chu, Zeming Liu, Kaiqi Huang Title: Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models Arxiv: http://arxiv.org/abs/2603.22212v1 Abstract: Video--based world m...

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model 25.03.2026

🤗 Upvotes: 90 | cs. CV Authors: SII-GAIR, Sand. ai, :, Ethan Chern, Hansi Teng, Hanwen Sun, Hao Wang, Hong Pan, Hongyu Jia, Jiadi Su, Jin Li, Junjie Yu, Lijie Liu, Lingzhi Li, Lyumanshan Ye, Min Hu, Qiangang Wang, Quanwei Qi, Steffi Chern, Tao Bu, Taoran Wang, Teren Xu, Tianning Zhang, Tiantian Mi, Weixian Xu, Wenqiang Zhang, Wentai Zhang, Xianping Yi, Xiaojie Cai, Xiaoyang Kang, Yan Ma, Yixiu Li...

LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning 25.03.2026

🤗 Upvotes: 63 | cs. AI, cs. CL Authors: Jianing Wang, Jianfei Zhang, Qi Guo, Linsen Guo, Rumei Li, Chao Zhang, Chong Peng, Cunguang Wang, Dengchang Zhao, Jiarong Shi, Jingang Wang, Liulin Feng, Mengxia Shen, Qi Li, Shengnan An, Shun Wang, Wei Shi, Xiangyu Xi, Xiaoyu Li, Xuezhi Cao, Yi Lu, Yunke Zhao, Zhengyu Chen, Zhimin Lin, Wei Wang, Peng Pei, Xunliang Cai Title: LongCat-Flash-Prover: Advancing...

Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs 25.03.2026

🤗 Upvotes: 60 | cs. CV, cs. AI Authors: Nimrod Shabtay, Moshe Kimhi, Artem Spector, Sivan Haray, Ehud Rivlin, Chaim Baskin, Raja Giryes, Eli Schwartz Title: Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs Arxiv: http://arxiv.org/abs/2603.16932v1 Abstract: Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accur...

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis 25.03.2026

🤗 Upvotes: 55 | cs. IR, cs. AI, cs. CL Authors: Zhuofeng Li, Dongfu Jiang, Xueguang Ma, Haoxiang Zhang, Ping Nie, Yuyu Zhang, Kai Zou, Jianwen Xie, Yu Zhang, Wenhu Chen Title: OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis Arxiv: http://arxiv.org/abs/2603.20278v1 Abstract: Training deep research agents requires long-horizon trajectories that interleave s...

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding 25.03.2026

🤗 Upvotes: 45 | cs. CV Authors: Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu Title: VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding Arxiv: http://arxiv.org/abs/2603.22285v1 Abstract: Long video understanding remains challenging for multimodal large language models (MLLMs) due to limited context windows, which necessitate identify...

SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning 25.03.2026

🤗 Upvotes: 39 | cs. CV Authors: Byungwoo Jeon, Dongyoung Kim, Huiwon Jang, Insoo Kim, Jinwoo Shin Title: SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning Arxiv: http://arxiv.org/abs/2603.22057v1 Abstract: Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly t...

F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting 25.03.2026

🤗 Upvotes: 31 | cs. CV Authors: Injae Kim, Chaehyeon Kim, Minseong Bae, Minseok Joo, Hyunwoo J. Kim Title: F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting Arxiv: http://arxiv.org/abs/2603.21304v1 Abstract: Feed-forward 3D Gaussian Splatting methods enable single-pass reconstruction and real-time rendering. However, they typically adopt rigid pixel-to-Gaussian...

mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT 25.03.2026

🤗 Upvotes: 28 | cs. LG, cs. AI Authors: Woosung Koh, Jeyoung Jeon, Youngjin Song, Yujin Cheon, Soowon Oh, Jaehyeong Choi, Se-Young Yun Title: mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT Arxiv: http://arxiv.org/abs/2603.21606v2 Abstract: Current language model training commonly applies multi-task Supervised Fine-Tuning (SFT) using a homogeneous compute budget ac...

HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning 24.03.2026

🤗 Upvotes: 96 | cs. CV, cs. AI, cs. CL Authors: Shenzhi Wang, Shixuan Liu, Jing Zhou, Chang Gao, Xiong-Hui Chen, Binghai Wang, An Yang, Shiji Song, Bowen Yu, Gao Huang, Junyang Lin Title: HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning Arxiv: http://arxiv.org/abs/2603.17024v2 Abstract: Vision-language models (VLMs) show strong multimodal capabilities but still strug...

Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models 24.03.2026

🤗 Upvotes: 87 | cs. CV Authors: Songchun Zhang, Zeyue Xue, Siming Fu, Jie Huang, Xianghao Kong, Y Ma, Haoyang Huang, Nan Duan, Anyi Rao Title: Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models Arxiv: http://arxiv.org/abs/2603.17051v1 Abstract: Distilled autoregressive (AR) video models enable efficient streaming generation but frequently misalign...

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation 24.03.2026

🤗 Upvotes: 42 | cs. CV Authors: Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu, Begüm Demir, Nicu Sebe, Paolo Rota Title: TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation Arxiv: http://arxiv.org/abs/2603.19039v1 Abstract: Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in pr...

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models 24.03.2026

🤗 Upvotes: 30 | cs. CV Authors: Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, Massimiliano Mancini Title: ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2603.19466v1 Abstract: Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask som...

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow 24.03.2026

🤗 Upvotes: 26 | cs. CV Authors: Zhifei Yang, Guangyao Zhai, Keyang Lu, YuYang Yin, Chao Zhang, Zhen Xiao, Jieyi Long, Nassir Navab, Yikai Wang Title: FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow Arxiv: http://arxiv.org/abs/2603.19598v1 Abstract: Scene generation has extensive industrial applications, demanding both high realism and precise control over...

The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus 24.03.2026

🤗 Upvotes: 25 | cs. LG, cs. AI Authors: Amartya Roy, Rasul Tutunov, Xiaotong Ji, Matthieu Zimmer, Haitham Bou-Ammar Title: The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus Arxiv: http://arxiv.org/abs/2603.20105v1 Abstract: LLMs are increasingly used as general-purpose reasoners, but long inputs remain bottlenecked by a fixed context window. Recursive Language Model...

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation 24.03.2026

🤗 Upvotes: 21 | cs. CV, cs. AI Authors: Jiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu, Hongbin Xu, Hai Ci, Ruigang Niu, Weihua Chen, Fan Wang, Yong Liu Title: LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation Arxiv: http://arxiv.org/abs/2603.20192v1 Abstract: Recent advances in diffusion models have significantly improved text-to-video generation, enabling p...

Hyperagents 24.03.2026

🤗 Upvotes: 21 | cs. AI Authors: Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin, Tatiana Shavrina Title: Hyperagents Arxiv: http://arxiv.org/abs/2603.19461v1 Abstract: Self-improving AI systems aim to reduce reliance on human engineering by learning to improve their own learning and problem-solving processes. Existing approaches to self-improvement rel...

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding 21.03.2026

🤗 Upvotes: 63 | cs. CV, cs. RO Authors: Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia, Yumeng Zhang, Xiaofan Li, Xiao Tan, Xiang Bai Title: Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Arxiv: http://arxiv.org/abs/2603.19235v1 Abstract: While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blin...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.