Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

MinT: Managed Infrastructure for Training and Serving Millions of LLMs 15.05.2026

🤗 Upvotes: 144 | cs. LG, cs. AI, cs. DC Authors: Mind Lab, :, Song Cao, Vic Cao, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Nolan Ho, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Andrew Lei, Kyrie Lei, Alexy Li, Lucian Li, Ray Li, Theo Li,...

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image 15.05.2026

🤗 Upvotes: 122 | cs. LG, cs. CL, cs. CV Authors: Alan Arazi, Eilam Shapira, Shoham Grunblat, Mor Ventura, Elad Hoffer, Gioia Blayer, David Holzmüller, Lennart Purucker, Gaël Varoquaux, Frank Hutter, Roi Reichart Title: MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image Arxiv: http://arxiv.org/abs/2605.10616v1 Abstract: Tabular Foundation Models have recently established the...

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation 15.05.2026

🤗 Upvotes: 79 | cs. CV, cs. AI Authors: Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, Song Han, Han Cai, Mike Zheng Shou Title: AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Arxiv: http://arxiv.org/abs/2605.13724v1 Abstract: Few-step video generation has been significantly advanced by consistency distillation. However, the performance of consistency-distilled mode...

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context 15.05.2026

🤗 Upvotes: 75 | cs. CV Authors: Zhaowei Wang, Lishu Luo, Haodong Duan, Weiwei Liu, Sijin Wu, Ji Luo, Shen Yan, Shuai Peng, Sihang Yuan, Chaoyi Huang, Yi Lin, Yangqiu Song Title: Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Arxiv: http://arxiv.org/abs/2605.13831v1 Abstract: Long-context modeling is becoming a core capability of modern large visio...

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents 15.05.2026

🤗 Upvotes: 57 | cs. SD, cs. AI, cs. CL, cs. LG Authors: Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara Title: EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Arxiv: http://arxiv.org...

Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling 15.05.2026

🤗 Upvotes: 43 | cs. LG, cs. AI, cs. CL, cs. MA Authors: Eilam Shapira, Moshe Tennenholtz, Roi Reichart Title: Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling Arxiv: http://arxiv.org/abs/2605.12411v1 Abstract: AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant n...

Qwen-Image-VAE-2.0 Technical Report 15.05.2026

🤗 Upvotes: 42 | cs. CV Authors: Zekai Zhang, Deqing Li, Kuan Cao, Yujia Wu, Chenfei Wu, Yu Wu, Liang Peng, Hao Meng, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yan Shu, Yanran Zhang, Yilei Chen, Yixian Xu, Yuxiang Chen, Zhendong Wang, Zihao Liu, Zikai Zhou, Yiliang Gu, Yi Wang, Xiaoxiao Xu, Lin Qu Title: Qwen-Image-VAE-...

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking 15.05.2026

🤗 Upvotes: 31 | cs. CV Authors: Jisu Nam, Jahyeok Koo, Soowon Son, Jaewoo Jung, Honggyu An, Junhwa Hur, Seungryong Kim Title: TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Arxiv: http://arxiv.org/abs/2605.12587v1 Abstract: Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-fram...

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling 15.05.2026

🤗 Upvotes: 30 | cs. CV Authors: Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu, Yuran Wang, Yifan Dai, Xinyu Liu, Yiyan Ji, Xiaoling Gu, Yuanxing Zhang Title: Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling Arxiv: http://arxiv.org/abs/2605.13062v1 Abstract: Recent image editing models have achieved remarkable progress in instruction following, mult...

Many-Shot CoT-ICL: Making In-Context Learning Truly Learn 15.05.2026

🤗 Upvotes: 28 | cs. CL, cs. AI Authors: Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung Title: Many-Shot CoT-ICL: Making In-Context Learning Truly Learn Arxiv: http://arxiv.org/abs/2605.13511v1 Abstract: In-context learning (ICL) adapts large language models (LLMs) to new tasks by conditioning on demonstrations in the prompt without parameter updates. With long-context models, many-shot ICL can u...

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents 14.05.2026

🤗 Upvotes: 128 | cs. CR, cs. CL Authors: Yining Chen, Jihao Zhao, Bo Tang, Haofen Wang, Yue Zhang, Fei Huang, Feiyu Xiong, Zhiyu Li Title: MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents Arxiv: http://arxiv.org/abs/2605.09530v2 Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of l...

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 14.05.2026

🤗 Upvotes: 114 | cs. CV Authors: Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, Shihao Bai, Silei Wu, Weichen Fan, Wenjie Ye, Wenwen Tong, Xiangyu Fan, Yan Li, Yubo Wang, Zhijie Cao, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Yuwei Niu, Yue Zhu, Bo Liu, Chengguang Lv, Haojia Yu, Haozhe Xie, Hongli Wang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jingcheng Ni, Junxiang Xu, Kaihuan Liang, Lianqiang Shi, Li...

$δ$-mem: Efficient Online Memory for Large Language Models 14.05.2026

🤗 Upvotes: 90 | cs. AI Authors: Jingdi Lei, Di Zhang, Junxian Li, Weida Wang, Kaixuan Fan, Xiang Liu, Qihan Liu, Xiaoteng Ma, Baian Chen, Soujanya Poria Title: $δ$-mem: Efficient Online Memory for Large Language Models Arxiv: http://arxiv.org/abs/2605.12357v1 Abstract: Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems....

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards 14.05.2026

🤗 Upvotes: 66 | cs. CL, cs. LG Authors: Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister Title: RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Arxiv: http://arxiv.org/abs/2605.10899v1 Abstract: Training deep research agents, namely systems that plan,...

Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics 14.05.2026

🤗 Upvotes: 53 | cs. AI, cs. CL, cs. LG Authors: Jishnu Sethumadhavan Nair, Patrice Bechard, Rishabh Maheshwary, Surajit Dasgupta, Sravan Ramachandran, Aakash Bhagat, Shruthan Radhakrishna, Pulkit Pattnaik, Johan Obando-Ceron, Shiva Krishna Reddy Malay, Sagar Davasam, Seganrasan Subramanian, Vipul Mittal, Sridhar Krishna Nemala, Christopher Pal, Srinivas Sunkara, Sai Rajeswar Title: Do Enterprise...

World Action Models: The Next Frontier in Embodied AI 14.05.2026

🤗 Upvotes: 51 | cs. RO, cs. CL, cs. CV Authors: Siyin Wang, Junhao Shi, Zhaoyang Fu, Xinzhe He, Feihong Liu, Chenchen Yang, Yikang Zhou, Zhaoye Fei, Jingjing Gong, Jinlan Fu, Mike Zheng Shou, Xuanjing Huang, Xipeng Qiu, Yu-Gang Jiang Title: World Action Models: The Next Frontier in Embodied AI Arxiv: http://arxiv.org/abs/2605.12090v1 Abstract: Vision-Language-Action (VLA) models have achieved str...

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization 14.05.2026

🤗 Upvotes: 30 | cs. CV, cs. AI Authors: Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Jing Jin, Yuan Zhou Title: Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Arxiv: http://arxiv.org/abs/2605.10780v2 Abstract: Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generat...

Efficient Pre-Training with Token Superposition 14.05.2026

🤗 Upvotes: 30 | cs. CL Authors: Bowen Peng, Théo Gigant, Jeffrey Quesnelle Title: Efficient Pre-Training with Token Superposition Arxiv: http://arxiv.org/abs/2605.06546v1 Abstract: Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data throughput. In this work, we present Token-Sup...

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward 14.05.2026

🤗 Upvotes: 28 | cs. CV, cs. AI, cs. LG Authors: Runhui Huang, Jie Wu, Rui Yang, Zhe Liu, Hengshuang Zhao Title: AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward Arxiv: http://arxiv.org/abs/2605.12495v1 Abstract: In this paper, we propose AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unifi...

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments 14.05.2026

🤗 Upvotes: 27 | cs. AI, cs. MA Authors: Giridhar Ganapavarapu, Dhaval Patel Title: MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments Arxiv: http://arxiv.org/abs/2605.09131v1 Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how agents conceptualize the...

Qwen-Image-2.0 Technical Report 13.05.2026

🤗 Upvotes: 78 | cs. CV Authors: Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng, Jiahao Li, Jie Zhang, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kuan Cao, Kun Yan, Liang Peng, Lihan Jiang, Niantong Li, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Xihua Wang, Yan Shu, Yanran Zhang, Yi Wang, Yilei Chen, Ying Ba, Yixian Xu, Yujia Wu, Yuxiang Chen, Zecheng Tang, Zekai Zhang, Zhendong Wang...

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs 13.05.2026

🤗 Upvotes: 66 | cs. CL Authors: Guijin Son, Seungone Kim, Catherine Arnett, Hyunwoo Ko, Hyein Lee, Hyeonah Kang, Jiang Longxi, Jin Yun, JungYup Lee, Kyungmin Lee, Sam Yoosuk Kim, Sang Park, Seunghyeok Hong, SeungJae Lee, Seungyeop Yi, Shinae Shin, SunHye Bok, Sunyoung Shin, Yonghoon Ji, Youngtaek Kim, Hanearl Jung, Akari Asai, Graham Neubig, Sean Welleck, Youngjae Yu, Akshelin R, Alexander B. Iva...

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models 13.05.2026

🤗 Upvotes: 54 | cs. CV Authors: Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang Title: CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models Arxiv: http://arxiv.org/abs/2605.08735v1 Abstract: Recent "Thinking with Video" approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reasoning artifacts...

TMAS: Scaling Test-Time Compute via Multi-Agent Synergy 13.05.2026

🤗 Upvotes: 44 | cs. AI Authors: George Wu, Nan Jing, Qing Yi, Chuan Hao, Ming Yang, Feng Chang, Yuan Wei, Jian Yang, Ran Tao, Bryan Dai Title: TMAS: Scaling Test-Time Compute via Multi-Agent Synergy Arxiv: http://arxiv.org/abs/2605.10344v1 Abstract: Test-time scaling has become an effective paradigm for improving the reasoning ability of large language models by allocating additional computation...

PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents 13.05.2026

🤗 Upvotes: 29 | cs. AI, cs. SE Authors: Bihui Yu, Xinglong Xu, Junjie Jiang, Jiabei Cheng, Caijun Jia, Siyuan Li, Conghui He, Jingxuan Wei, Cheng Tan Title: PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents Arxiv: http://arxiv.org/abs/2605.10341v1 Abstract: A LaTeX manuscript that compiles without error is not necessarily publication-ready. The resulting PDFs frequent...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.