Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling 03.12.2025

🤗 Upvotes: 140 | cs. CV Authors: Zuhao Yang, Sudong Wang, Kaichen Zhang, Keming Wu, Sicong Leng, Yifan Zhang, Chengwei Qin, Shijian Lu, Xingxuan Li, Lidong Bing Title: LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Arxiv: http://arxiv.org/abs/2511.20785v1 Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Though...

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights 03.12.2025

🤗 Upvotes: 83 | cs. CV, cs. AI Authors: Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu, Cheng Tan Title: Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights Arxiv: http://arxiv.org/abs/2512.01816v1 Abstract: Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using t...

Stabilizing Reinforcement Learning with LLMs: Formulation and Practices 03.12.2025

🤗 Upvotes: 56 | cs. LG, cs. AI, cs. CL Authors: Chujie Zheng, Kai Dang, Bowen Yu, Mingze Li, Huiqiang Jiang, Junrong Lin, Yuqiong Liu, Hao Lin, Chencan Wu, Feng Hu, An Yang, Jingren Zhou, Junyang Lin Title: Stabilizing Reinforcement Learning with LLMs: Formulation and Practices Arxiv: http://arxiv.org/abs/2512.01374v2 Abstract: This paper proposes a novel formulation for reinforcement learning (R...

How Far Are We from Genuinely Useful Deep Research Agents? 03.12.2025

🤗 Upvotes: 44 | cs. CL Authors: Dingling Zhang, He Zhu, Jincheng Ren, Kangqi Song, Xinran Zhou, Boyu Feng, Shudong Liu, Jiabin Luo, Weihao Xie, Zhaohui Wang, Tianrui Qin, King Zhu, Yuqing Wang, Qianben Chen, Yuchen Eleanor Jiang, Wei Wang, Jiaheng Liu, Wangchunshu Zhou Title: How Far Are We from Genuinely Useful Deep Research Agents? Arxiv: http://arxiv.org/abs/2512.01948v1 Abstract: Deep Researc...

What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards 03.12.2025

🤗 Upvotes: 41 | cs. CV Authors: Minh-Quan Le, Yuanzhi Zhu, Vicky Kalogeiton, Dimitris Samaras Title: What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards Arxiv: http://arxiv.org/abs/2512.00425v1 Abstract: Recent video diffusion models can synthesize visually compelling clips, yet often violate basic physical laws-objects float, accelerations drift, and colli...

Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout 03.12.2025

🤗 Upvotes: 38 | cs. CV Authors: Hidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Pinar Yanardag Title: Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout Arxiv: http://arxiv.org/abs/2511.20649v1 Abstract: Current autoregressive video diffusion models are constrained by three core bottlenecks: (i) the finite temporal horizon impo...

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment 03.12.2025

🤗 Upvotes: 36 | cs. CV Authors: Ziheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu, Qibin Hou, Ming-Ming Cheng, Mike Zheng Shou Title: The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment Arxiv: http://arxiv.org/abs/2511.20614v1 Abstract: Previous works have explored various customized generation tasks given a reference image, but they stil...

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models 03.12.2025

🤗 Upvotes: 33 | cs. CV Authors: Zhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou, Shoufa Chen, Haonan Qiu, Xiaoke Huang, Zhaochong An, Fanny Yang, Aditya Patel, Viktar Atliha, Tony Ng, Xiao Han, Chuyan Zhu, Chenyang Zhang, Ding Liu, Juan-Manuel Perez-Rua, Sen He, Jürgen Schmidhuber, Wenhu Chen, Ping Luo, Wei Liu, Tao Xiang, Jonas Schult, Yuren Cong Title: TUNA: Taming Unified Visual Representati...

LFM2 Technical Report 03.12.2025

🤗 Upvotes: 31 | cs. LG, cs. AI Authors: Alexander Amini, Anna Banaszak, Harold Benoit, Arthur Böök, Tarek Dakhran, Song Duong, Alfred Eng, Fernando Fernandes, Marc Härkönen, Anne Harrington, Ramin Hasani, Saniya Karwa, Yuri Khrustalev, Maxime Labonne, Mathias Lechner, Valentine Lechner, Simon Lee, Zetian Li, Noel Loo, Jacob Marks, Edoardo Mosca, Samuel J. Paech, Paul Pak, Rom N. Parnichkun, Alex...

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer 02.12.2025

🤗 Upvotes: 78 | cs. CV Authors: Z-Image Team, Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Shijie Huang, Zhaohui Hou, Dengyang Jiang, Xin Jin, Liangchen Li, Zhen Li, Zhong-Yu Li, David Liu, Dongyang Liu, Junhan Shi, Qilong Wu, Feng Yu, Chi Zhang, Shifeng Zhang, Shilin Zhou Title: Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Arxiv: htt...

REASONEDIT: Towards Reasoning-Enhanced Image Editing Models 02.12.2025

🤗 Upvotes: 40 | cs. CV Authors: Fukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang, Peng Xing, Rui Wang, Wei Cheng, Yingming Wang, Aojie Li, Zixin Yin, Pengtao Chen, Xiangyu Zhang, Daxin Jiang, Xianfang Zeng, Gang Yu Title: REASONEDIT: Towards Reasoning-Enhanced Image Editing Models Arxiv: http://arxiv.org/abs/2511.22625v1 Abstract: Recent advances in image editing models have shown remarkable progres...

Vision Bridge Transformer at Scale 02.12.2025

🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Zhenxiong Tan, Zeqing Wang, Xingyi Yang, Songhua Liu, Xinchao Wang Title: Vision Bridge Transformer at Scale Arxiv: http://arxiv.org/abs/2511.23199v1 Abstract: We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into...

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 02.12.2025

🤗 Upvotes: 25 | cs. AI, cs. CL Authors: Zhihong Shao, Yuxiang Luo, Chengda Lu, Z. Z. Ren, Jiewen Hu, Tian Ye, Zhibin Gou, Shirong Ma, Xiaokang Zhang Title: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning Arxiv: http://arxiv.org/abs/2511.22570v1 Abstract: Large language models have made significant progress in mathematical reasoning, which serves as an important testbed for AI and...

Architecture Decoupling Is Not All You Need For Unified Multimodal Model 02.12.2025

🤗 Upvotes: 23 | cs. CV Authors: Dian Zheng, Manyuan Zhang, Hongyu Li, Kai Zou, Hongbo Liu, Ziyu Guo, Kaituo Feng, Yexin Liu, Ying Luo, Yan Feng, Peng Pei, Xunliang Cai, Hongsheng Li Title: Architecture Decoupling Is Not All You Need For Unified Multimodal Model Arxiv: http://arxiv.org/abs/2511.22663v1 Abstract: Unified multimodal models for image generation and understanding represent a significa...

Multimodal Evaluation of Russian-language Architectures 28.11.2025

🤗 Upvotes: 71 | cs. CL, cs. AI, cs. CV Authors: Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin, Maria Tikhonova, Alexander Kharitonov, Yulia Lyakh, Petr Surovtsev, Denis Shevelev, Vildan Saburov, Vasily Konovalov, Elisei Rykov, Ivan Sviridov, Amina Miftakhova, Ilseyar Alimova, Alexander Panchenko, Alexander Kapitanov, Alena Fenogenova Title: Multimodal Evaluation of Russian-languag...

Latent Collaboration in Multi-Agent Systems 28.11.2025

🤗 Upvotes: 60 | cs. CL, cs. AI, cs. LG Authors: Jiaru Zou, Xiyuan Yang, Ruizhong Qiu, Gaotang Li, Katherine Tieu, Pan Lu, Ke Shen, Hanghang Tong, Yejin Choi, Jingrui He, James Zou, Mengdi Wang, Ling Yang Title: Latent Collaboration in Multi-Agent Systems Arxiv: http://arxiv.org/abs/2511.20639v1 Abstract: Multi-agent systems (MAS) extend large language models (LLMs) from independent single-model r...

Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation 28.11.2025

🤗 Upvotes: 37 | cs. CV, cs. AI Authors: Inferix Team, Tianyu Feng, Yizeng Han, Jiahao He, Yuanyu He, Xi Lin, Teng Liu, Hanfeng Lu, Jiasheng Tang, Wei Wang, Zhiyuan Wang, Jichao Wu, Mingyang Yang, Yinghao Yu, Zeyu Zhang, Bohan Zhuang Title: Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation Arxiv: http://arxiv.org/abs/2511.20714v1 Abstract: World models serve as...

GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms 27.11.2025

🤗 Upvotes: 87 | cs. NE, cs. AI, cs. LG Authors: Valentin Khrulkov, Andrey Galichin, Denis Bashkirov, Dmitry Vinichenko, Oleg Travkin, Roman Alferov, Andrey Kuznetsov, Ivan Oseledets Title: GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms Arxiv: http://arxiv.org/abs/2511.17592v1 Abstract: Recent advances in LLM-guided evolutionary computation, particularly Al...

MedSAM3: Delving into Segment Anything with Medical Concepts 27.11.2025

🤗 Upvotes: 38 | cs. CV, cs. AI Authors: Anglin Liu, Rundong Xue, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen Title: MedSAM3: Delving into Segment Anything with Medical Concepts Arxiv: http://arxiv.org/abs/2511.19046v1 Abstract: Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manu...

Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning 27.11.2025

🤗 Upvotes: 37 | cs. CV, cs. AI Authors: Jiaqi Liu, Kaiwen Xiong, Peng Xia, Yiyang Zhou, Haonian Ji, Lu Feng, Siwei Han, Mingyu Ding, Huaxiu Yao Title: Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning Arxiv: http://arxiv.org/abs/2511.19900v1 Abstract: Vision-language agents have achieved remarkable progress in a variety of multimodal reasoning tasks; however,...

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation 27.11.2025

🤗 Upvotes: 37 | cs. CV Authors: Jiaming Zhang, Shengming Cao, Rui Li, Xiaotong Zhao, Yutao Cui, Xinglin Hou, Gangshan Wu, Haolan Chen, Yu Xu, Limin Wang, Kai Ma Title: SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation Arxiv: http://arxiv.org/abs/2511.19320v1 Abstract: Preserving first-frame identity while ensuring precise motion control is a fundamental cha...

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation 27.11.2025

🤗 Upvotes: 28 | cs. CV Authors: Zhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao, Cheng Chen, Junyi Chen, Jiacheng Wei, Wei Cheng, Shiyu Liu, Yunuo Chen, Gang Yu, Guosheng Lin Title: iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Arxiv: http://arxiv.org/abs/2511.20635v1 Abstract: Pre-trained video models learn powerful priors for generating high-quality, temporally...

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward 27.11.2025

🤗 Upvotes: 26 | cs. CV, cs. CL Authors: Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin, Bin Lin, Zongjian Li, Bin Zhu, Weihao Yu, Li Yuan Title: Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward Arxiv: http://arxiv.org/abs/2511.20561v1 Abstract: Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundament...

GigaWorld-0: World Models as Data Engine to Empower Embodied AI 27.11.2025

🤗 Upvotes: 24 | cs. CV, cs. RO Authors: GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jiagang Zhu, Kerui Li, Mengyuan Xu, Qiuping Deng, Siting Wang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yankai Wang, Yu Cao, Yifan Chang, Yuan Xu, Yun Ye, Yang Wang, Yukun Zhou, Zhengyuan Zhang, Zhehao Dong, Zheng Zhu Title: GigaWorld-0: World Models as Data Engine to Em...

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 27.11.2025

🤗 Upvotes: 23 | cs. CL Authors: Zhenyi Shen, Junru Lu, Lin Gui, Jiazheng Li, Yulan He, Di Yin, Xing Sun Title: SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space Arxiv: http://arxiv.org/abs/2511.20102v1 Abstract: The quadratic complexity of full attention limits efficient long-context processing in large language models (LLMs). Sparse attention mitigates t...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.