Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Rethinking Verification for LLM Code Generation: From Generation to Testing 11.07.2025

🤗 Upvotes: 23 | cs. CL Authors: Zihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu, Wenwei Zhang, Minnan Luo, Songyang Zhang, Kai Chen Title: Rethinking Verification for LLM Code Generation: From Generation to Testing Arxiv: http://arxiv.org/abs/2507.06920v2 Abstract: Large language models (LLMs) have recently achieved notable success in code-generation benchmarks such as HumanEval and LiveCodeBench...

SingLoRA: Low Rank Adaptation Using a Single Matrix 10.07.2025

🤗 Upvotes: 68 | cs. AI Authors: David Bensaïd, Noam Rotstein, Roy Velich, Daniel Bensaïd, Ron Kimmel Title: SingLoRA: Low Rank Adaptation Using a Single Matrix Arxiv: http://arxiv.org/abs/2507.05566v1 Abstract: Low-Rank Adaptation (LoRA) has significantly advanced parameter-efficient fine-tuning of large pretrained models. LoRA augments the pre-trained weights of a model by adding the product of...

A Survey on Latent Reasoning 10.07.2025

🤗 Upvotes: 60 | cs. CL Authors: Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu Liu, Jian Yang, Wangchunshu Zhou, Chujie Zheng, Chongxuan Li, Yuyin Zhou, Zhoujun Li, Zhaox...

OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion 10.07.2025

🤗 Upvotes: 45 | cs. CV Authors: Yunhan Yang, Yufan Zhou, Yuan-Chen Guo, Zi-Xin Zou, Yukun Huang, Ying-Tian Liu, Hao Xu, Ding Liang, Yan-Pei Cao, Xihui Liu Title: OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion Arxiv: http://arxiv.org/abs/2507.06165v1 Abstract: The creation of 3D assets with explicit, editable part structures is crucial for advancing interactive...

How to Train Your LLM Web Agent: A Statistical Diagnosis 10.07.2025

🤗 Upvotes: 40 | cs. AI, cs. LG, stat. ML Authors: Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Megh Thakkar, Thibault Le Sellier de Chezelles, Nicolas Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Xue Liu, Alexandre Drouin, Laurent Charlin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia Title: How to Train Your LLM Web Agent: A Statistic...

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling 10.07.2025

🤗 Upvotes: 35 | cs. RO, cs. CV Authors: Meng Wei, Chenyang Wan, Xiqian Yu, Tai Wang, Yuqiang Yang, Xiaohan Mao, Chenming Zhu, Wenzhe Cai, Hanqing Wang, Yilun Chen, Xihui Liu, Jiangmiao Pang Title: StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Arxiv: http://arxiv.org/abs/2507.05240v1 Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires...

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization 10.07.2025

🤗 Upvotes: 35 | cs. CL Authors: Zhongyuan Peng, Yifan Yao, Kaijing Ma, Shuyue Guo, Yizhe Li, Yichi Zhang, Chenchen Zhang, Yifan Zhang, Zhouliang Yu, Luming Li, Minghao Liu, Yihang Xia, Jiawei Shen, Yuchen Wu, Yixin Cao, Zhaoxiang Zhang, Wenhao Huang, Jiaheng Liu, Ge Zhang Title: CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization Arxiv: http://arxiv.org/abs/2507.06181v...

RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents 10.07.2025

🤗 Upvotes: 27 | cs. CL, cs. AI, cs. CY Authors: Peisong Wang, Ruotian Ma, Bang Zhang, Xingyu Chen, Zhiwei He, Kang Luo, Qingsong Lv, Qingxuan Jiang, Zheng Xie, Shanyi Wang, Yuan Li, Fanghua Ye, Jian Li, Yifan Yang, Zhaopeng Tu, Xiaolong Li Title: RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents Arxiv: http://arxiv.org/abs/2507.03112v1 Abstract: Large language mo...

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos 10.07.2025

🤗 Upvotes: 24 | cs. CV, cs. AI Authors: Rongsheng Wang, Junying Chen, Ke Ji, Zhenyang Cai, Shunian Chen, Yunjin Yang, Benyou Wang Title: MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos Arxiv: http://arxiv.org/abs/2507.05675v1 Abstract: Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation...

MemOS: A Memory OS for AI System 09.07.2025

🤗 Upvotes: 83 | cs. CL Authors: Zhiyu Li, Shichao Song, Chenyang Xi, Hanyu Wang, Chen Tang, Simin Niu, Ding Chen, Jiawei Yang, Chunyu Li, Qingchen Yu, Jihao Zhao, Yezhaohui Wang, Peng Liu, Zehao Lin, Pengyuan Wang, Jiahao Huo, Tianyi Chen, Kai Chen, Kehang Li, Zhen Tao, Junpeng Ren, Huayi Lai, Hao Wu, Bo Tang, Zhenren Wang, Zhaoxin Fan, Ningyu Zhang, Linfeng Zhang, Junchi Yan, Mingchuan Yang, Ton...

Should We Still Pretrain Encoders with Masked Language Modeling? 09.07.2025

🤗 Upvotes: 63 | cs. CL Authors: Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse, Duarte M. Alves, Emmanuel Malherbe, André F. T. Martins, Céline Hudelot, Pierre Colombo Title: Should We Still Pretrain Encoders with Masked Language Modeling? Arxiv: http://arxiv.org/abs/2507.00994v2 Abstract: Learning high-quality text representations is fundamental to a wide range of NLP tasks. While e...

Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 09.07.2025

🤗 Upvotes: 36 | cs. CL, cs. AI Authors: Xiangru Tang, Tianrui Qin, Tianhao Peng, Ziyang Zhou, Daniel Shao, Tingting Du, Xinming Wei, Peng Xia, Fang Wu, He Zhu, Ge Zhang, Jiaheng Liu, Xingyao Wang, Sirui Hong, Chenglin Wu, Hao Cheng, Chi Wang, Wangchunshu Zhou Title: Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving Arxiv: http://arxiv.org/abs/2507.06229v1 Abstract: As langu...

4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture 09.07.2025

🤗 Upvotes: 33 | cs. CV Authors: Yutian Chen, Shi Guo, Tianshuo Yang, Lihe Ding, Xiuyuan Yu, Jinwei Gu, Tianfan Xue Title: 4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture Arxiv: http://arxiv.org/abs/2507.05163v1 Abstract: Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majori...

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge 09.07.2025

🤗 Upvotes: 30 | cs. CV, cs. RO Authors: Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, XinQiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, Xin Jin Title: DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Arxiv: http://arxiv.org/abs/2507.04447v1 Abstract: Recent advances in vision-language-action (VLA) models have sho...

Pre-Trained Policy Discriminators are General Reward Models 09.07.2025

🤗 Upvotes: 28 | cs. CL, cs. LG Authors: Shihan Dou, Shichun Liu, Yuming Yang, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Demin Song, Haijun Lv, Songyang Gao, Chengqi Lv, Enyu Zhou, Honglin Guo, Zhiheng Xi, Wenwei Zhang, Qipeng Guo, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Tao Gui, Kai Chen Title: Pre-Trained Policy Discriminators are General Reward Models Arxiv: http://arxiv.or...

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset 09.07.2025

🤗 Upvotes: 22 | cs. CL, cs. AI Authors: Zhiheng Xi, Guanyu Li, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jiaqi Liu, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai, Tao Ji, Tao Gui, Qi Zhang, Philip Torr, Xuanjing Huang Title: BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Arxiv: http://arxiv.org/abs/2507.03483v2 Abstract: In this paper, we introduce BMMR, a...

WebSailor: Navigating Super-human Reasoning for Web Agent 05.07.2025

🤗 Upvotes: 56 | cs. CL, cs. AI Authors: Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Baixuan Li, Zhengwei Tao, Xinyu Wang, Weizhou Shen, Junkai Zhang, Dingchu Zhang, Xixi Wu, Yong Jiang, Ming Yan, Pengjun Xie, Fei Huang, Jingren Zhou Title: WebSailor: Navigating Super-human Reasoning for Web Agent Arxiv: http://arxiv.org/abs/2507.02592v1 Abstract: Transcen...

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion 05.07.2025

🤗 Upvotes: 45 | cs. CV Authors: Fangfu Liu, Hao Li, Jiawei Chi, Hanyang Wang, Minghui Yang, Fudong Wang, Yueqi Duan Title: LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion Arxiv: http://arxiv.org/abs/2507.02813v1 Abstract: Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent develo...

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback 05.07.2025

🤗 Upvotes: 33 | cs. CV Authors: Nina Konovalova, Maxim Nikolaev, Andrey Kuznetsov, Aibek Alanov Title: Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Arxiv: http://arxiv.org/abs/2507.02321v1 Abstract: Despite significant progress in text-to-image diffusion models, achieving precise spatial control over generated outputs remains challenging. ControlNet add...

IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction 05.07.2025

🤗 Upvotes: 32 | q-bio. BM Authors: The IntFold Team, Leon Qiao, Wayne Bai, He Yan, Gary Liu, Nova Xi, Xiang Zhang Title: IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction Arxiv: http://arxiv.org/abs/2507.02025v1 Abstract: We introduce IntFold, a controllable foundation model for both general and specialized biomolecular structure prediction. In...

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy 05.07.2025

🤗 Upvotes: 31 | cs. CL, cs. AI, cs. LG Authors: Chris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He, Jiacai Liu, Chaojie Wang, Rui Yan, Wei Shen, Fuxiang Zhang, Jiacheng Xu, Yang Liu, Yahui Zhou Title: Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy Arxiv: http://arxiv.org/abs/2507.01352v2 Abstract: Despite the critical role of reward models (RMs) in reinforcement learning...

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers 05.07.2025

🤗 Upvotes: 27 | cs. CV Authors: Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu, Yan Ma, Xiaoye Qu, Jiaqi Liu, Yanshu Li, Kaide Zeng, Zhengyuan Yang, Linjie Li, Yu Cheng, Heng Ji, Junxian He, Yi R. Fung Title: Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers Arxiv: http://arxiv.org/abs/2506.23918v3 Abstract: Recent progress in multimodal reasoning has been...

Kwai Keye-VL Technical Report 04.07.2025

🤗 Upvotes: 97 | cs. CV Authors: Kwai Keye Team, Biao Yang, Bin Wen, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Hao Peng, Haojie Ding, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Jin Ouyang, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Shengnan Zhang, Siyang Mao, Sui Huang, Tianke Zhang, Tingting Gao, Wei Chen, Wei Y...

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory 04.07.2025

🤗 Upvotes: 61 | cs. CV Authors: Nan Chen, Mengqi Huang, Yihao Meng, Zhendong Mao Title: LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Arxiv: http://arxiv.org/abs/2507.01945v1 Abstract: Animation colorization is a crucial part of real animation industry production. Long animation colorization has high labor costs. Therefore, automated long animation colorization based o...

Depth Anything at Any Condition 04.07.2025

🤗 Upvotes: 35 | cs. CV, cs. AI Authors: Boyuan Sun, Modi Jin, Bowen Yin, Qibin Hou Title: Depth Anything at Any Condition Arxiv: http://arxiv.org/abs/2507.01634v1 Abstract: We present Depth Anything at Any Condition (DepthAnything-AC), a foundation monocular depth estimation (MDE) model capable of handling diverse environmental conditions. Previous foundation MDE models achieve impressive perform...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.