Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells 02.04.2026

🤗 Upvotes: 73 | q-bio. QM, cs. AI, q-bio. GN Authors: Han Zhang, Guo-Hua Yuan, Chaohao Yuan, Tingyang Xu, Tian Bian, Hong Cheng, Wenbing Huang, Deli Zhao, Yu Rong Title: Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells Arxiv: http://arxiv.org/abs/2603.25240v1 Abstract: Modeling cellular states and predicting their responses to perturbations are centr...

GEMS: Agent-Native Multimodal Generation with Memory and Skills 02.04.2026

🤗 Upvotes: 63 | cs. CV Authors: Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Yu Cheng, Yang Yang Title: GEMS: Agent-Native Multimodal Generation with Memory and Skills Arxiv: http://arxiv.org/abs/2603.28088v1 Abstract: Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and speciali...

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development 02.04.2026

🤗 Upvotes: 48 | cs. CV, cs. AI Authors: Zhongying Deng, Cheng Tang, Ziyan Huang, Jiashi Lin, Ying Chen, Junzhi Ning, Chenglong Ma, Jiyao Liu, Wei Li, Yinghao Zhu, Shujian Gao, Yanyan Huang, Sibo Ju, Yanzhou Su, Pengcheng Chen, Wenhao Tang, Tianbin Li, Haoyu Wang, Yuanfeng Ji, Hui Sun, Shaobo Min, Liang Peng, Feilong Tang, Haochen Xue, Rulin Zhou, Chaoyang Zhang, Wenjie Li, Shaohao Rui, Weijie Ma,...

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward 02.04.2026

🤗 Upvotes: 44 | cs. CV Authors: Zhaochong An, Orest Kupyn, Théo Uscidda, Andrea Colaco, Karan Ahuja, Serge Belongie, Mar Gonzalez-Franco, Marta Tintore Gazulla Title: VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward Arxiv: http://arxiv.org/abs/2603.26599v1 Abstract: Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric co...

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis 02.04.2026

🤗 Upvotes: 33 | cs. CV, cs. MM Authors: Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou, Kaituo Feng, Wenbo Hu, Yi-Fan Zhang, Yunlong Lin, Wenxuan Huang, Mingyang Song, Dasen Dai, Bolin Jiang, Manyuan Zhang, Shi-Xue Zhang, Zhengkai Jiang, Lucas Wang, Zhao Zhong, Yu Cheng, Nanyun Peng Title: Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis Arxiv: http://arxiv.org/a...

CutClaw: Agentic Hours-Long Video Editing via Music Synchronization 02.04.2026

🤗 Upvotes: 32 | cs. CV Authors: Shifang Zhao, Yihan Hu, Ying Shan, Yunchao Wei, Xiaodong Cun Title: CutClaw: Agentic Hours-Long Video Editing via Music Synchronization Arxiv: http://arxiv.org/abs/2603.29664v1 Abstract: Editing the video content with audio alignment forms a digital human-made art in current social media. However, the time-consuming and repetitive nature of manual video editing has...

daVinci-LLM:Towards the Science of Pretraining 02.04.2026

🤗 Upvotes: 24 | cs. AI, cs. CL Authors: Yiwei Qin, Yixiu Liu, Tiantian Mi, Muhang Xie, Zhen Huang, Weiye Si, Pengrui Lu, Siyuan Feng, Xia Wu, Liming Liu, Ye Luo, Jinlong Hou, Qipeng Guo, Yu Qiao, Pengfei Liu Title: daVinci-LLM:Towards the Science of Pretraining Arxiv: http://arxiv.org/abs/2603.27164v1 Abstract: The foundational pretraining phase determines a model's capability ceiling, as post-tr...

TAPS: Task Aware Proposal Distributions for Speculative Sampling 01.04.2026

🤗 Upvotes: 118 | cs. CL, cs. AI Authors: Mohamad Zbib, Mohamad Bazzi, Ammar Mohanna, Hasan Abed Al Kader Hammoud, Bernard Ghanem Title: TAPS: Task Aware Proposal Distributions for Speculative Sampling Arxiv: http://arxiv.org/abs/2603.27027v1 Abstract: Speculative decoding accelerates autoregressive generation by letting a lightweight draft model propose future tokens that a larger target model th...

Towards a Medical AI Scientist 01.04.2026

🤗 Upvotes: 65 | cs. AI, cs. LG Authors: Hongtao Wu, Boyun Zheng, Dingjie Song, Yu Jiang, Jianfeng Gao, Lei Xing, Lichao Sun, Yixuan Yuan Title: Towards a Medical AI Scientist Arxiv: http://arxiv.org/abs/2603.28589v1 Abstract: Autonomous systems that generate scientific hypotheses, conduct experiments, and draft manuscripts have recently emerged as a promising paradigm for accelerating discovery....

Gen-Searcher: Reinforcing Agentic Search for Image Generation 01.04.2026

🤗 Upvotes: 45 | cs. CV Authors: Kaituo Feng, Manyuan Zhang, Shuang Chen, Yunlong Lin, Kaixuan Fan, Yilei Jiang, Hongyu Li, Dian Zheng, Chenyang Wang, Xiangyu Yue Title: Gen-Searcher: Reinforcing Agentic Search for Image Generation Arxiv: http://arxiv.org/abs/2603.28767v1 Abstract: Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images....

Emergent Social Intelligence Risks in Generative Multi-Agent Systems 01.04.2026

🤗 Upvotes: 41 | cs. MA, cs. CL, cs. CY Authors: Yue Huang, Yu Jiang, Wenjie Wang, Haomin Zhuang, Xiaonan Luo, Yuchen Ma, Zhangchen Xu, Zichen Chen, Nuno Moniz, Zinan Lin, Pin-Yu Chen, Nitesh V Chawla, Nouha Dziri, Huan Sun, Xiangliang Zhang Title: Emergent Social Intelligence Risks in Generative Multi-Agent Systems Arxiv: http://arxiv.org/abs/2603.27771v1 Abstract: Multi-agent systems composed of...

EpochX: Building the Infrastructure for an Emergent Agent Civilization 01.04.2026

🤗 Upvotes: 40 | cs. AI, cs. MA Authors: Huacan Wang, Chaofa Yuan, Xialie Zhuang, Tu Hu, Shuo Zhang, Jun Han, Shi Wei, Daiqiang Li, Jingping Liu, Kunyi Wang, Zihan Yin, Zhenheng Tang, Andy Wang, Henry Peng Zou, Philip S. Yu, Sen Hu, Qizhen Lan, Ronghao Chen Title: EpochX: Building the Infrastructure for an Emergent Agent Civilization Arxiv: http://arxiv.org/abs/2603.27304v1 Abstract: General-purpo...

On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models 01.04.2026

🤗 Upvotes: 30 | cs. LG, cs. AI Authors: Chongyang Zhao, Mingsong Li, Haodong Lu, Dong Gong Title: On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models Arxiv: http://arxiv.org/abs/2603.27481v1 Abstract: Multimodal Continual Instruction Tuning aims to continually enhance Large Vision Language Models (LVLMs) by learning from new dat...

GEditBench v2: A Human-Aligned Benchmark for General Image Editing 01.04.2026

🤗 Upvotes: 27 | cs. CV Authors: Zhangqi Jiang, Zheng Sun, Xianfang Zeng, Yufeng Yang, Xuanyang Zhang, Yongliang Wu, Wei Cheng, Gang Yu, Xu Yang, Bihan Wen Title: GEditBench v2: A Human-Aligned Benchmark for General Image Editing Arxiv: http://arxiv.org/abs/2603.28547v1 Abstract: Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, e...

Make Geometry Matter for Spatial Reasoning 01.04.2026

🤗 Upvotes: 25 | cs. CV, cs. AI Authors: Shihua Zhang, Qiuhong Shen, Shizun Wang, Tianbo Pan, Xinchao Wang Title: Make Geometry Matter for Spatial Reasoning Arxiv: http://arxiv.org/abs/2603.26639v1 Abstract: Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic...

PRBench: End-to-end Paper Reproduction in Physics Research 01.04.2026

🤗 Upvotes: 23 | cs. CL, hep-lat, hep-ph, physics.comp-ph, physics.optics Authors: Shi Qiu, Junyi Deng, Yiwei Deng, Haoran Dong, Jieyu Fu, Mao Li, Zeyu Li, Zhaolong Zhang, Huiwen Zheng, Leidong Bao, Anqi Lv, Zihan Mo, Yadi Niu, Yiyang Peng, Yu Tian, Yili Wang, Ziyu Wang, Zi-Yu Wang, Jiashen Wei, Liuheng Wu, Aoran Xue, Leyi Yang, Guanglu Yuan, Xiarui Zhan, Jingjun Zhang, Zifan Zheng, Pengfei Liu, L...

PixelSmile: Toward Fine-Grained Facial Expression Editing 28.03.2026

🤗 Upvotes: 100 | cs. CV, cs. AI Authors: Jiabin Hua, Hengyuan Xu, Aojie Li, Wei Cheng, Gang Yu, Xingjun Ma, Yu-Gang Jiang Title: PixelSmile: Toward Fine-Grained Facial Expression Editing Arxiv: http://arxiv.org/abs/2603.25728v1 Abstract: Fine-grained facial expression editing has long been limited by intrinsic semantic overlap. To address this, we construct the Flex Facial Expression (FFE) datase...

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale 28.03.2026

🤗 Upvotes: 90 | cs. LG, cs. CL, cs. CV Authors: Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, Xinyu Zhou, Dongzhan Zhou, Zhiwang Zhou, Yuhao Zhou, Bowen Zhou, Zhanping Zhong, Zhijie Zhong, Haiteng Zhao, Penghao Zhao, Xiaomeng Zhao, Zhiyuan Zhao, Yechen Zhang, Jin Zhang, Wenwei Zhang, Hongjie Zhang, Zhuo Zhang, Wenlong Zhang, Bo Zhang, Chao Zhang, Chen Zhang, Yuhang Zan...

Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration 28.03.2026

🤗 Upvotes: 40 | cs. CV Authors: Danil Tokhchukov, Aysel Mirzoeva, Andrey Kuznetsov, Konstantin Sobolev Title: Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration Arxiv: http://arxiv.org/abs/2603.24800v1 Abstract: In this paper, we uncover the hidden potential of Diffusion Transformers (DiTs) to significantly enhance generative tasks. Through an in-depth analysis of the d...

RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models 28.03.2026

🤗 Upvotes: 39 | cs. CV Authors: Yufeng Yang, Xianfang Zeng, Zhangqi Jiang, Fukun Yin, Jianzhuang Liu, Wei Cheng, jinghong lan, Shiyu Liu, Yuqi Peng, Gang YU, Shifeng Chen Title: RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models Arxiv: http://arxiv.org/abs/2603.25502v1 Abstract: Image restoration under real-world degradations is critical for dow...

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data 28.03.2026

🤗 Upvotes: 26 | cs. CV Authors: Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu Title: MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data Arxiv: http://arxiv.org/abs/2603.25319v1 Abstract: Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view...

Voxtral TTS 28.03.2026

🤗 Upvotes: 24 | cs. AI Authors: Alexander H. Liu, Alexis Tacnet, Andy Ehrenberg, Andy Lo, Chen-Yo Sun, Guillaume Lample, Henry Lagarde, Jean-Malo Delignon, Jaeyoung Kim, John Harvill, Khyathi Raghavi Chandu, Lorenzo Signoretti, Margaret Jennings, Patrick von Platen, Pavankumar Reddy Muddireddy, Rohin Arora, Sanchit Gandhi, Samuel Humeau, Soham Ghosh, Srijan Mishra, Van Phung, Abdelaziz Bounhar, A...

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? 27.03.2026

🤗 Upvotes: 27 | cs. CL, cs. LG Authors: Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang Title: Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? Arxiv: http://arxiv.org/abs/2603.24472v1 Abstract: Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while sho...

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding 26.03.2026

🤗 Upvotes: 112 | cs. CV Authors: Hejun Dong, Junbo Niu, Bin Wang, Weijun Zeng, Wentao Zhang, Conghui He Title: MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding Arxiv: http://arxiv.org/abs/2603.22458v1 Abstract: Optical character recognition (OCR) has evolved from line-level transcription to structured document parsing, requiring models to recover long-form seq...

WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG 26.03.2026

🤗 Upvotes: 69 | cs. CV Authors: Zhen Li, Zian Meng, Shuwei Shi, Wenshuo Peng, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang Title: WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG Arxiv: http://arxiv.org/abs/2603.23497v1 Abstract: Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics dri...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.