Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts 24.10.2025

🤗 Upvotes: 44 | cs. CL Authors: Siyuan Wang, Gaokai Zhang, Li Lyna Zhang, Ning Shang, Fan Yang, Dongyao Chen, Mao Yang Title: LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts Arxiv: http://arxiv.org/abs/2510.19363v1 Abstract: Reasoning over long contexts is essential for large language models. While reinforcement learning (RL) enhances short-context reasoning by inducing "...

Language Models are Injective and Hence Invertible 24.10.2025

🤗 Upvotes: 42 | cs. LG, cs. AI Authors: Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà Title: Language Models are Injective and Hence Invertible Arxiv: http://arxiv.org/abs/2510.15511v3 Abstract: Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could m...

GigaBrain-0: A World Model-Powered Vision-Language-Action Model 24.10.2025

🤗 Upvotes: 34 | cs. RO, cs. CV Authors: GigaBrain Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jie Li, Jiagang Zhu, Lv Feng, Peng Li, Qiuping Deng, Runqi Ouyang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yang Wang, Yifan Li, Yilong Li, Yiran Ding, Yuan Xu, Yun Ye, Yukun Zhou, Zhehao Dong, Zhenan Wang, Zhichao Liu, Zheng Zhu Title: GigaBrain-0: A World Model-Powered...

LightMem: Lightweight and Efficient Memory-Augmented Generation 23.10.2025

🤗 Upvotes: 86 | cs. CL, cs. AI, cs. CV, cs. LG, cs. MA Authors: Jizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang, Yuqi Tang, Ziwen Xu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Huajun Chen, Ningyu Zhang Title: LightMem: Lightweight and Efficient Memory-Augmented Generation Arxiv: http://arxiv.org/abs/2510.18866v1 Abstract: Despite their remarkable capabilities, Large Language Models (LL...

Efficient Long-context Language Model Training by Core Attention Disaggregation 23.10.2025

🤗 Upvotes: 70 | cs. LG, cs. DC Authors: Yonghao Zhuang, Junda Chen, Bo Pang, Yi Gu, Yibo Zhu, Yimin Jiang, Ion Stoica, Eric Xing, Hao Zhang Title: Efficient Long-context Language Model Training by Core Attention Disaggregation Arxiv: http://arxiv.org/abs/2510.18121v1 Abstract: We present core attention disaggregation (CAD), a technique that improves long-context large language model training by d...

World-in-World: World Models in a Closed-Loop World 23.10.2025

🤗 Upvotes: 68 | cs. CV Authors: Jiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu, Arda Uzunoglu, Shunchi Zhang, Yana Wei, Jiahao Wang, Vishal M. Patel, Paul Pu Liang, Daniel Khashabi, Cheng Peng, Rama Chellappa, Tianmin Shu, Alan Yuille, Yilun Du, Jieneng Chen Title: World-in-World: World Models in a Closed-Loop World Arxiv: http://arxiv.org/abs/2510.18135v1 Abstract: Generative world models (WM...

UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation 23.10.2025

🤗 Upvotes: 59 | cs. CV Authors: Yibin Wang, Zhimin Li, Yuhang Zang, Jiazi Bu, Yujie Zhou, Yi Xin, Junjun He, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang Title: UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation Arxiv: http://arxiv.org/abs/2510.18701v1 Abstract: Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks i...

Chem-R: Learning to Reason as a Chemist 23.10.2025

🤗 Upvotes: 46 | cs. CE Authors: Weida Wang, Benteng Chen, Di Zhang, Wanhao Liu, Shuchen Pu, Ben Gao, Jin Zeng, Xiaoyong Wei, Tianshu Yu, Shuzhou Sun, Tianfan Fu, Wanli Ouyang, Lei Bai, Jiatong Li, Zifu Wang, Yuqiang Li, Shufei Zhang Title: Chem-R: Learning to Reason as a Chemist Arxiv: http://arxiv.org/abs/2510.16880v2 Abstract: Although large language models (LLMs) have significant potential to...

MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation 23.10.2025

🤗 Upvotes: 33 | cs. CV Authors: Weinan Jia, Yuning Lu, Mengqi Huang, Hualiang Wang, Binyuan Huang, Nan Chen, Mu Liu, Jidong Jiang, Zhendong Mao Title: MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation Arxiv: http://arxiv.org/abs/2510.18692v1 Abstract: Long video generation with Diffusion Transformers (DiTs) is bottlenecked by the quadratic scaling of full attention with seque...

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 23.10.2025

🤗 Upvotes: 30 | cs. CV, cs. AI, cs. CL Authors: Haochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou, Yanwei Li, Jiacong Wang, Jiani Zheng, Ye Tian, Jiahao Meng, Zilong Huang, Guangcan Mai, Anran Wang, Yunhai Tong, Zhuochen Wang, Xiangtai Li, Zhaoxiang Zhang Title: Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs Arxiv: http://arxiv.org/abs/2510.18876v2 Abstract:...

Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model 23.10.2025

🤗 Upvotes: 27 | cs. CL, cs. AI Authors: Ling Team, Anqi Shen, Baihui Li, Bin Hu, Bin Jing, Cai Chen, Chao Huang, Chao Zhang, Chaokun Yang, Cheng Lin, Chengyao Wen, Congqi Li, Deng Zhao, Dingbo Yuan, Donghai You, Fagui Mao, Fanzhuang Meng, Feng Xu, Guojie Li, Guowei Wang, Hao Dai, Haonan Zheng, Hong Liu, Jia Guo, Jiaming Liu, Jian Liu, Jianhao Fu, Jiannan Shi, Jianwen Wang, Jianxin Lai, Jin Yang,...

IF-VidCap: Can Video Caption Models Follow Instructions? 23.10.2025

🤗 Upvotes: 24 | cs. CV Authors: Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei, Yiwen He, Runzhe Wen, Chenxi Liao, Chengkang Jiang, An Ping, Shuo Gao, Suhan Wang, Zhaozhou Bian, Zijun Zhou, Jingyi Xie, Jiayi Zhou, Jing Wang, Yifan Yao, Weihao Xie, Yingshui Tan, Yanghai Wang, Qianqian Xie, Zhaoxiang Zhang, Jiaheng Liu Title: IF-VidCap: Can Video Caption Models Follow Instructions? Arxiv: http:/...

DeepAnalyze: Agentic Large Language Models for Autonomous Data Science 22.10.2025

🤗 Upvotes: 56 | cs. AI, cs. CL, cs. DB Authors: Shaolei Zhang, Ju Fan, Meihao Fan, Guoliang Li, Xiaoyong Du Title: DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Arxiv: http://arxiv.org/abs/2510.16872v1 Abstract: Autonomous data science, from raw data sources to analyst-grade deep research reports, has been a long-standing challenge, and is now becoming feasible with the e...

PICABench: How Far Are We from Physically Realistic Image Editing? 22.10.2025

🤗 Upvotes: 55 | cs. CV, cs. AI Authors: Yuandong Pu, Le Zhuo, Songhao Han, Jinbo Xing, Kaiwen Zhu, Shuo Cao, Bin Fu, Si Liu, Hongsheng Li, Yu Qiao, Wenlong Zhang, Xi Chen, Yihao Liu Title: PICABench: How Far Are We from Physically Realistic Image Editing? Arxiv: http://arxiv.org/abs/2510.17681v2 Abstract: Image editing has achieved remarkable progress recently. Modern editing models could already...

Glyph: Scaling Context Windows via Visual-Text Compression 22.10.2025

🤗 Upvotes: 43 | cs. CV, cs. CL, cs. LG Authors: Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang, Hongning Wang, Minlie Huang Title: Glyph: Scaling Context Windows via Visual-Text Compression Arxiv: http://arxiv.org/abs/2510.17800v2 Abstract: Large language models (LLMs) increasingly rely on long-context model...

FineVision: Open Data Is All You Need 22.10.2025

🤗 Upvotes: 32 | cs. CV, cs. AI Authors: Luis Wiedmann, Orr Zohar, Amir Mahla, Xiaohan Wang, Rui Li, Thibaud Frere, Leandro von Werra, Aritra Roy Gosthipaty, Andrés Marafioti Title: FineVision: Open Data Is All You Need Arxiv: http://arxiv.org/abs/2510.17269v1 Abstract: The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public da...

TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model 22.10.2025

🤗 Upvotes: 32 | cs. CL Authors: Bin Yu, Xinming Wang, Shijie Lian, Haotian Li, Changti Wu, Ruina Hu, Bailing Wang, Yuliang Wei, Kai Chen Title: TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model Arxiv: http://arxiv.org/abs/2510.16449v1 Abstract: Large language models (LLMs) have shown remarkable progress in complex reasoning tasks, large...

Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation 22.10.2025

🤗 Upvotes: 28 | cs. CL, cs. AI, cs. IR, cs. LG Authors: Chenghao Zhang, Guanting Dong, Xinyu Yang, Zhicheng Dou Title: Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation Arxiv: http://arxiv.org/abs/2510.17354v1 Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing large language models (LLMs) by retrieving relevant documents from a...

When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling 22.10.2025

🤗 Upvotes: 27 | cs. CL, cs. AI Authors: Heecheol Yun, Kwangmin Ki, Junghyun Lee, Eunho Yang Title: When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling Arxiv: http://arxiv.org/abs/2510.15346v1 Abstract: Ensembling Large Language Models (LLMs) has gained attention as a promising approach to surpass the performance of individual models by leveraging their complementar...

A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning 21.10.2025

🤗 Upvotes: 112 | cs. LG, cs. AI Authors: Zhi Zhou, Yuhao Tan, Zenan Li, Yuan Yao, Lan-Zhe Guo, Yu-Feng Li, Xiaoxing Ma Title: A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning Arxiv: http://arxiv.org/abs/2510.15444v1 Abstract: Test-time scaling seeks to improve the reasoning performance of large language models (LLMs) by adding computational resources. A...

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM 21.10.2025

🤗 Upvotes: 56 | cs. CV, cs. AI, cs. CL Authors: Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Jason Lu, Oluwatobi Olabiyi, Frank Wang, R...

NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks 21.10.2025

🤗 Upvotes: 50 | cs. CV Authors: Junliang Ye, Shenghao Xie, Ruowen Zhao, Zhengyi Wang, Hongyu Yan, Wenqiang Zu, Lei Ma, Jun Zhu Title: NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks Arxiv: http://arxiv.org/abs/2510.15019v1 Abstract: 3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient,...

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs 21.10.2025

🤗 Upvotes: 39 | cs. CL Authors: Nikita Afonin, Nikita Andriyanov, Nikhil Bageshpura, Kyle Liu, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Alexander Panchenko, Oleg Rogov, Elena Tutubalina, Mikhail Seleznyov Title: Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs Arxiv: http://arxiv.org/abs/2510.11288v1 Abstract: Recent work has shown th...

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset 21.10.2025

🤗 Upvotes: 37 | cs. CV Authors: Qingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu, Hanlin Wang, Wen Wang, Ka Leong Cheng, Shuailei Ma, Yanhong Zeng, Zichen Liu, Yinghao Xu, Yujun Shen, Qifeng Chen Title: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset Arxiv: http://arxiv.org/abs/2510.15742v1 Abstract: Instruction-based video editing promises to democratize content creation...

Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery 21.10.2025

🤗 Upvotes: 33 | cs. CV Authors: Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai, Wei-Cheng Chang, Chung-Ho Wu, Jiewen Chan, Zhenjun Zhao, Chieh Hubert Lin, Yu-Lun Liu Title: Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery Arxiv: http://arxiv.org/abs/2510.15869v1 Abstract: Synthesizing large-scale, explorable, and geometrically accurate 3D urban scenes is a challenging yet valua...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.