Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

AI Can Learn Scientific Taste 18.03.2026

🤗 Upvotes: 215 | cs. CL Authors: Jingqi Tong, Mingzhe Li, Hangcheng Li, Yongzhuo Yang, Yurong Mou, Weijie Ma, Zhiheng Xi, Hongji Chen, Xiaoran Liu, Qinyuan Cheng, Ming Zhang, Qiguang Chen, Weifeng Ge, Qipeng Guo, Tianlei Ying, Tianxiang Sun, Yining Zheng, Xinchi Chen, Jun Zhao, Ning Ding, Xuanjing Huang, Yugang Jiang, Xipeng Qiu Title: AI Can Learn Scientific Taste Arxiv: http://arxiv.org/abs/260...

OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data 18.03.2026

🤗 Upvotes: 127 | cs. AI, cs. CL Authors: Yuwen Du, Rui Ye, Shuo Tang, Xinyu Zhu, Yijun Lu, Yuzhu Cai, Siheng Chen Title: OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data Arxiv: http://arxiv.org/abs/2603.15594v1 Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet the development of high-...

EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings 18.03.2026

🤗 Upvotes: 118 | cs. AI, cs. LG Authors: Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair, Sagar Davasam, Aman Tiwari, Sathwik Tejaswi Madhusudhan, Sridhar Krishna Nemala, Srinivas Sunkara, Sai Rajeswar Title: EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings Arxiv: http://arxiv.org/abs/2603.13594v1 Abstract: Large...

Grounding World Simulation Models in a Real-World Metropolis 18.03.2026

🤗 Upvotes: 103 | cs. CV Authors: Junyoung Seo, Hyunwook Choi, Minkyung Kwon, Jinhyeok Choi, Siyoon Jin, Gayoung Lee, Junho Kim, JoungBin Lee, Geonmo Gu, Dongyoon Han, Sangdoo Yun, Seungryong Kim, Jin-Hwa Kim Title: Grounding World Simulation Models in a Real-World Metropolis Arxiv: http://arxiv.org/abs/2603.15583v1 Abstract: What if a world simulation model could render not an imagined environmen...

HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions 18.03.2026

🤗 Upvotes: 92 | cs. CV, cs. RO Authors: Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, Ziwei Liu Title: HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions Arxiv: http://arxiv.org/abs/2603.15612v1 Abstract: We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual c...

Attention Residuals 18.03.2026

🤗 Upvotes: 60 | cs. CL Authors: Kimi Team, Guangyu Chen, Yu Zhang, Jianlin Su, Weixin Xu, Siyuan Pan, Yaoyu Wang, Yucheng Wang, Guanduo Chen, Bohong Yin, Yutian Chen, Junjie Yan, Ming Wei, Y. Zhang, Fanqing Meng, Chao Hong, Xiaotong Xie, Shaowei Liu, Enzhe Lu, Yunpeng Tai, Yanru Chen, Xin Men, Haiqing Guo, Y. Charles, Haoyu Lu, Lin Sui, Jinguo Zhu, Zaida Zhou, Weiran He, Weixiao Huang, Xinran Xu,...

Mixture-of-Depths Attention 18.03.2026

🤗 Upvotes: 50 | cs. CL, cs. AI Authors: Lianghui Zhu, Yuxin Fang, Bencheng Liao, Shijie Wang, Tianheng Cheng, Zilong Huang, Chen Chen, Lai Wei, Yutao Zeng, Ya Wang, Yi Lin, Yu Li, Xinggang Wang Title: Mixture-of-Depths Attention Arxiv: http://arxiv.org/abs/2603.15619v1 Abstract: Scaling depth is a key driver for large language models (LLMs). Yet, as LLMs become deeper, they often suffer from sign...

Effective Distillation to Hybrid xLSTM Architectures 18.03.2026

🤗 Upvotes: 31 | cs. LG Authors: Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied, Anamaria-Roberta Hartl, David Stap, Pieter-Jan Hoedt, Maximilian Beck, Sebastian Böck, Günter Klambauer, Sepp Hochreiter Title: Effective Distillation to Hybrid xLSTM Architectures Arxiv: http://arxiv.org/abs/2603.15590v1 Abstract: There have been numerous attempts to distill quadratic attention-based large la...

Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models 18.03.2026

🤗 Upvotes: 25 | cs. CV Authors: Lexiang Xiong, Qi Li, Jingwen Ye, Xinchao Wang Title: Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models Arxiv: http://arxiv.org/abs/2603.15557v1 Abstract: Vision-Language Models (VLMs) frequently "hallucinate" - generate plausible yet factually incorrect statements - posing a critical barrier to their trustwor...

ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer 18.03.2026

🤗 Upvotes: 21 | cs. CV Authors: Ruonan Yu, Zhenxiong Tan, Zigeng Chen, Songhua Liu, Xinchao Wang Title: ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer Arxiv: http://arxiv.org/abs/2603.15478v1 Abstract: Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable gene...

LMEB: Long-horizon Memory Embedding Benchmark 17.03.2026

🤗 Upvotes: 53 | cs. CL Authors: Xinping Zhao, Xinshuo Hu, Jiaxin Xu, Danyu Tang, Xin Zhang, Mengjia Zhou, Yan Zhong, Yao Zhou, Zifei Shan, Meishan Zhang, Baotian Hu, Min Zhang Title: LMEB: Long-horizon Memory Embedding Benchmark Arxiv: http://arxiv.org/abs/2603.12572v1 Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in...

Can Vision-Language Models Solve the Shell Game? 17.03.2026

🤗 Upvotes: 30 | cs. CV, cs. CL Authors: Tiedong Liu, Wee Sun Lee Title: Can Vision-Language Models Solve the Shell Game? Arxiv: http://arxiv.org/abs/2603.08436v1 Abstract: Visual entity tracking is an innate cognitive ability in humans, yet it remains a critical bottleneck for Vision-Language Models (VLMs). This deficit is often obscured in existing video benchmarks by visual shortcuts. We introd...

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation 17.03.2026

🤗 Upvotes: 28 | cs. CV, cs. AI Authors: Yichen Zhang, Da Peng, Zonghao Guo, Zijian Zhang, Xuesong Yang, Tong Sun, Shichu Sun, Yidan Zhang, Yanghao Li, Haiyan Zhao, Wang Xu, Qi Shi, Yangang Sun, Chi Chen, Shuo Wang, Yukun Yan, Xu Han, Qiang Ma, Wei Ke, Liang Wang, Zhiyuan Liu, Maosong Sun Title: Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension...

daVinci-Env: Open SWE Environment Synthesis at Scale 17.03.2026

🤗 Upvotes: 21 | cs. SE, cs. AI, cs. CL Authors: Dayuan Fu, Shenyu Wu, Yunze Wu, Zerui Peng, Yaxing Huang, Jie Sun, Ji Zeng, Mohan Jiang, Lin Zhang, Yukun Li, Jiarui Hu, Liming Liu, Jinlong Hou, Pengfei Liu Title: daVinci-Env: Open SWE Environment Synthesis at Scale Arxiv: http://arxiv.org/abs/2603.13023v1 Abstract: Training capable software engineering (SWE) agents demands large-scale, executable...

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections 14.03.2026

🤗 Upvotes: 40 | cs. CL, cs. AI Authors: Łukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha, Ryan Othniel Kearns, Adam Mahdi, Niels Rogge, Clémentine Fourrier, Siwei Han, Huaxiu Yao, Artemis Llabrés, Yiming Xu, Dimosthenis Karatzas, Hao Zhang, Anupam Datta Title: Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections Arxiv: http://arx...

OpenClaw-RL: Train Any Agent Simply by Talking 13.03.2026

🤗 Upvotes: 70 | cs. CL Authors: Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, Ling Yang Title: OpenClaw-RL: Train Any Agent Simply by Talking Arxiv: http://arxiv.org/abs/2603.10165v1 Abstract: Every agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system recovers it as a li...

Flash-KMeans: Fast and Memory-Efficient Exact K-Means 13.03.2026

🤗 Upvotes: 52 | cs. DC Authors: Shuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li, Xiaoze Fan, Jintao Zhang, Han Cai, Yujun Lin, Xiuyu Li, Kurt Keutzer, Song Han, Chenfeng Xu, Ion Stoica Title: Flash-KMeans: Fast and Memory-Efficient Exact K-Means Arxiv: http://arxiv.org/abs/2603.09229v1 Abstract: $k$-means has historically been positioned primarily as an offline processing primitive, typically used...

MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents 13.03.2026

🤗 Upvotes: 25 | cs. CV, cs. AI Authors: Kangsan Kim, Yanlai Yang, Suji Kim, Woongyeong Yeo, Youngwan Lee, Mengye Ren, Sung Ju Hwang Title: MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents Arxiv: http://arxiv.org/abs/2603.09827v2 Abstract: As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the...

LLM2Vec-Gen: Generative Embeddings from Large Language Models 13.03.2026

🤗 Upvotes: 23 | cs. CL Authors: Parishad BehnamGhader, Vaibhav Adlakha, Fabian David Schmidt, Nicolas Chapados, Marius Mosbach, Siva Reddy Title: LLM2Vec-Gen: Generative Embeddings from Large Language Models Arxiv: http://arxiv.org/abs/2603.10913v1 Abstract: LLM-based text embedders typically encode the semantic content of their input. However, embedding tasks require mapping diverse inputs to si...

Urban Socio-Semantic Segmentation with Vision-Language Reasoning 17.01.2026

🤗 Upvotes: 139 | cs. CV, cs. AI, cs. CY Authors: Yu Wang, Yi Wang, Rui Dai, Yujie Wang, Kaikui Liu, Xiangxiang Chu, Yansheng Li Title: Urban Socio-Semantic Segmentation with Vision-Language Reasoning Arxiv: http://arxiv.org/abs/2601.10477v1 Abstract: As hubs of human activity, urban surfaces consist of a wealth of semantic entities. Segmenting these various entities from satellite imagery is cruc...

STEP3-VL-10B Technical Report 17.01.2026

🤗 Upvotes: 130 | cs. CV Authors: Ailin Huang, Chengyuan Yao, Chunrui Han, Fanqi Wan, Hangyu Guo, Haoran Lv, Hongyu Zhou, Jia Wang, Jian Zhou, Jianjian Sun, Jingcheng Hu, Kangheng Lin, Liang Zhao, Mitt Huang, Song Yuan, Wenwen Qu, Xiangfeng Wang, Yanlin Lai, Yingxiu Zhao, Yinmin Zhang, Yukang Shi, Yuyang Chen, Zejia Weng, Ziyang Meng, Ang Li, Aobo Kong, Bo Dong, Changyi Wan, David Wang, Di Qi, Din...

Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 17.01.2026

🤗 Upvotes: 111 | cs. LG, cs. CL Authors: Zhiyuan Hu, Yucheng Wang, Yufei He, Jiaying Wu, Yilun Zhao, See-Kiong Ng, Cynthia Breazeal, Anh Tuan Luu, Hae Won Park, Bryan Hooi Title: Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs Arxiv: http://arxiv.org/abs/2601.08763v2 Abstract: Reinforcement learning (RL) has become a central paradigm for post-training large language m...

Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning 17.01.2026

🤗 Upvotes: 64 | cs. AI, cs. CL Authors: Zhiyuan Hu, Yunhai Hu, Juncheng Liu, Shuyue Stella Li, Yucheng Wang, Zhen Xu, See-Kiong Ng, Anh Tuan Luu, Xinxing Xu, Bryan Hooi, Cynthia Breazeal, Hae Won Park Title: Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning Arxiv: http://arxiv.org/abs/2601.09667v2 Abstract: Multi-agent systems have evolved into practical LLM-driven collabor...

Controlled Self-Evolution for Algorithmic Code Optimization 16.01.2026

🤗 Upvotes: 97 | cs. CL, cs. AI, cs. NE Authors: Tu Hu, Ronghao Chen, Shuo Zhang, Jianghao Yin, Mou Xiao Feng, Jingping Liu, Shaolei Zhang, Wenqi Jiang, Yuqi Fang, Sen Hu, Huacan Wang, Yi Xu Title: Controlled Self-Evolution for Algorithmic Code Optimization Arxiv: http://arxiv.org/abs/2601.07348v4 Abstract: Self-evolution methods enhance code generation through iterative "generate-verify-refine" c...

DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation 16.01.2026

🤗 Upvotes: 92 | cs. CL Authors: Yibo Wang, Lei Wang, Yue Deng, Keming Wu, Yao Xiao, Huanjin Yao, Liwei Kang, Hai Ye, Yongcheng Jing, Lidong Bing Title: DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation Arxiv: http://arxiv.org/abs/2601.09688v1 Abstract: Deep research systems are widely used for multi-step web research, analysis, and cross-source sy...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.