Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers 02.09.2025

🤗 Upvotes: 40 | cs. CL, cs. AI Authors: Ming Hu, Chenglong Ma, Wei Li, Wanghan Xu, Jiamin Wu, Jucheng Hu, Tianbin Li, Guohang Zhuang, Jiaqi Liu, Yingzhou Lu, Ying Chen, Chaoyang Zhang, Cheng Tan, Jie Ying, Guocheng Wu, Shujian Gao, Pengcheng Chen, Jiashi Lin, Haitao Wu, Lulu Chen, Fengxiang Wang, Yuanyuan Zhang, Xiangyu Zhao, Feilong Tang, Encheng Su, Junzhi Ning, Xinyao Liu, Ye Du, Changkai Ji,...

TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling 28.08.2025

🤗 Upvotes: 59 | cs. LG, cs. CL Authors: Yizhi Li, Qingshui Gu, Zhoufutu Wen, Ziniu Li, Tianshun Xing, Shuyue Guo, Tianyu Zheng, Xin Zhou, Xingwei Qu, Wangchunshu Zhou, Zheng Zhang, Wei Shen, Qian Liu, Chenghua Lin, Jian Yang, Ge Zhang, Wenhao Huang Title: TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling Arxiv: http://arxiv.or...

VibeVoice Technical Report 28.08.2025

🤗 Upvotes: 45 | cs. CL, cs. AI, cs. SD, eess. AS Authors: Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang, Yutao Sun, Li Dong, Yi Zhu, Weijiang Xu, Hangbo Bao, Zehua Wang, Shaohan Huang, Yan Xia, Furu Wei Title: VibeVoice Technical Report Arxiv: http://arxiv.org/abs/2508.19205v1 Abstract: This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple spea...

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics 28.08.2025

🤗 Upvotes: 43 | cs. LG, cs. AI Authors: Weida Wang, Dongchen Huang, Jiatong Li, Tengchao Yang, Ziyang Zheng, Di Zhang, Dong Han, Benteng Chen, Binzhao Luo, Zhiyu Liu, Kunling Liu, Zhiyuan Gao, Shiqi Geng, Wei Ma, Jiaming Su, Xin Li, Shuchen Pu, Yuhan Shui, Qianjia Cheng, Zhihao Dou, Dongfei Cui, Changyong He, Jin Zeng, Zeke Xie, Mao Su, Dongzhan Zhou, Yuqiang Li, Wanli Ouyang, Yunqi Cai, Xi Dai,...

VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space 28.08.2025

🤗 Upvotes: 28 | cs. CV Authors: Lin Li, Zehuan Huang, Haoran Feng, Gengxiong Zhuang, Rui Chen, Chunchao Guo, Lu Sheng Title: VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space Arxiv: http://arxiv.org/abs/2508.19247v1 Abstract: 3D local editing of specified regions is crucial for game industry and robot interaction. Recent methods typically edit rendered multi-view images...

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation 28.08.2025

🤗 Upvotes: 26 | cs. CV Authors: Jianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang, Chao Liang, Wang Liao, Han Liang, Yuan Zhang, Mingyuan Gao Title: OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation Arxiv: http://arxiv.org/abs/2508.19209v1 Abstract: Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likene...

Spacer: Towards Engineered Scientific Inspiration 28.08.2025

🤗 Upvotes: 25 | cs. AI, cs. LG, cs. NE Authors: Minhyeong Lee, Suyoung Hwang, Seunghyun Moon, Geonho Nah, Donghyun Koh, Youngjun Cho, Johyun Park, Hojin Yoo, Jiho Park, Haneul Choi, Sungbin Moon, Taehoon Hwang, Seungwon Kim, Jaeyeong Kim, Seongjun Kim, Juneau Jung Title: Spacer: Towards Engineered Scientific Inspiration Arxiv: http://arxiv.org/abs/2508.17661v1 Abstract: Recent advances in LLMs ha...

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning 28.08.2025

🤗 Upvotes: 23 | cs. LG Authors: Zihao Huang, Yu Bao, Qiyang Min, Siyan Chen, Ran Guo, Hongzhi Huang, Defa Zhu, Yutao Zeng, Banggu Wu, Xun Zhou, Siyuan Qiao Title: UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Arxiv: http://arxiv.org/abs/2508.18756v1 Abstract: While Mixture of Experts (MoE) models achieve remarkable efficiency by activating only subsets...

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 27.08.2025

🤗 Upvotes: 120 | cs. CV Authors: Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, Guanzhou Chen, Zichen Ding, Changyao Tian, Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Yuchen Duan, Xuehui Wang, Songze Li, Xiangy...

Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation 27.08.2025

🤗 Upvotes: 34 | cs. CV Authors: Yaqi Li, Peng Chen, Mingyang Han, Pi Bu, Haoxiang Shi, Runzhou Zhao, Yang Yao, Xuan Zhang, Jun Song, Bo Zheng Title: Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation Arxiv: http://arxiv.org/abs/2508.18032v2 Abstract: Despite the promising progress of recent autoregressive models in text-to-image (T2I) generation, th...

MV-RAG: Retrieval Augmented Multiview Diffusion 27.08.2025

🤗 Upvotes: 31 | cs. CV, cs. AI Authors: Yosef Dayani, Omer Benishu, Sagie Benaim Title: MV-RAG: Retrieval Augmented Multiview Diffusion Arxiv: http://arxiv.org/abs/2508.16577v1 Abstract: Text-to-3D generation approaches have advanced significantly by leveraging pretrained 2D diffusion priors, producing high-quality and 3D-consistent outputs. However, they often fail to produce out-of-domain (OOD)...

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs 26.08.2025

🤗 Upvotes: 58 | cs. LG, cs. CL Authors: Huichi Zhou, Yihang Chen, Siyuan Guo, Xue Yan, Kin Hei Lee, Zihan Wang, Ka Yiu Lee, Guchun Zhang, Kun Shao, Linyi Yang, Jun Wang Title: Memento: Fine-tuning LLM Agents without Fine-tuning LLMs Arxiv: http://arxiv.org/abs/2508.16153v2 Abstract: In this paper, we introduce a novel learning paradigm for Adaptive Large Language Model (LLM) agents that eliminate...

Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR 26.08.2025

🤗 Upvotes: 41 | cs. CL Authors: Xiao Liang, Zhongzhi Li, Yeyun Gong, Yelong Shen, Ying Nian Wu, Zhijiang Guo, Weizhu Chen Title: Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR Arxiv: http://arxiv.org/abs/2508.14029v2 Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), part...

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks 26.08.2025

🤗 Upvotes: 34 | cs. RO, cs. CV Authors: Kaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang, Zeju Li, Bolin Zhang, Wancai Zheng, Xinyi Yu, Hao Chen, Chunhua Shen Title: ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks Arxiv: http://arxiv.org/abs/2508.08240v1 Abstract: Language-guided long-horizon mobile manipulation has long been a grand challenge in embodied semanti...

Intern-S1: A Scientific Multimodal Foundation Model 23.08.2025

🤗 Upvotes: 166 | cs. LG, cs. CL, cs. CV Authors: Lei Bai, Zhongrui Cai, Maosong Cao, Weihan Cao, Chiyu Chen, Haojiong Chen, Kai Chen, Pengcheng Chen, Ying Chen, Yongkang Chen, Yu Cheng, Yu Cheng, Pei Chu, Tao Chu, Erfei Cui, Ganqu Cui, Long Cui, Ziyun Cui, Nianchen Deng, Ning Ding, Nanqin Dong, Peijie Dong, Shihan Dou, Sinan Du, Haodong Duan, Caihua Fan, Ben Gao, Changjiang Gao, Jianfei Gao, Song...

Mobile-Agent-v3: Foundamental Agents for GUI Automation 23.08.2025

🤗 Upvotes: 40 | cs. AI Authors: Jiabo Ye, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Zhaoqing Zhu, Ziwei Zheng, Feiyu Gao, Junjie Cao, Zhengxi Lu, Jitong Liao, Qi Zheng, Fei Huang, Jingren Zhou, Ming Yan Title: Mobile-Agent-v3: Foundamental Agents for GUI Automation Arxiv: http://arxiv.org/abs/2508.15144v1 Abstract: This paper introduces GUI-Owl, a foundational GUI agent model that achieves...

Deep Think with Confidence 23.08.2025

🤗 Upvotes: 26 | cs. LG Authors: Yichao Fu, Xuewei Wang, Yuandong Tian, Jiawei Zhao Title: Deep Think with Confidence Arxiv: http://arxiv.org/abs/2508.15260v1 Abstract: Large Language Models (LLMs) have shown great potential in reasoning tasks through test-time scaling methods like self-consistency with majority voting. However, this approach often leads to diminishing returns in accuracy and high...

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries 23.08.2025

🤗 Upvotes: 26 | cs. CL, cs. AI Authors: Ming Yin, Dinghan Shen, Silei Xu, Jianbing Han, Sixun Dong, Mian Zhang, Yebowen Hu, Shujian Liu, Simin Ma, Song Wang, Sathish Reddy Indurthi, Xun Wang, Yiran Chen, Kaiqiang Song Title: LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries Arxiv: http://arxiv.org/abs/2508.15760v1 Abstract: Tool calling has emerged as a critical...

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization 22.08.2025

🤗 Upvotes: 57 | cs. LG, cs. CL Authors: Shuaijie She, Yu Bao, Yu Lu, Lu Xu, Tao Li, Wenhao Zhu, Shujian Huang, Shanbo Cheng, Lu Lu, Yuxuan Wang Title: DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization Arxiv: http://arxiv.org/abs/2508.14460v1 Abstract: We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via a...

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models 22.08.2025

🤗 Upvotes: 53 | cs. CE Authors: Ziyan Kuang, Feiyu Zhu, Maowei Jiang, Yanzhao Lai, Zelin Wang, Zhitong Wang, Meikang Qiu, Jiajia Huang, Min Peng, Qianqian Xie, Sophia Ananiadou Title: From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models Arxiv: http://arxiv.org/abs/2508.13491v1 Abstract: Large Language Models (LLMs) have shown promise for financial...

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction 22.08.2025

🤗 Upvotes: 47 | cs. AI, cs. LG Authors: Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Jinpeng Wang, Zaiyuan Wang, Yang Yang, Lingyue Yin, Mingren Yin, Zhenwei Zhu, Tianle Cai, Zehui Chen, Jiecao Chen, Yantao Du, Xiang Gao, Jiacheng Guo, Liang Hu, Jianpeng Jiao, Xiangsheng Li, Jingkai Liu, Shuang Ni, Zhoufutu Wen, Ge Zhang, Kaiyuan Zhang, Xin Zhou, Jose Blanchet, Xipeng Qiu, Mengdi...

MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds 22.08.2025

🤗 Upvotes: 29 | cs. GR, cs. CV Authors: Bingquan Dai, Li Ray Luo, Qihong Tang, Jie Wang, Xinyu Lian, Hao Xu, Minghan Qin, Xudong Xu, Bo Dai, Haoqian Wang, Zhaoyang Lyu, Jiangmiao Pang Title: MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds Arxiv: http://arxiv.org/abs/2508.14879v1 Abstract: Reconstructing 3D objects into editable programs is pivotal for applications like re...

Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization 22.08.2025

🤗 Upvotes: 26 | cs. CV Authors: Canyu Zhao, Xiaoman Li, Tianjian Feng, Zhiyue Zhao, Hao Chen, Chunhua Shen Title: Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization Arxiv: http://arxiv.org/abs/2508.14811v1 Abstract: We introduce Tinker, a versatile framework for high-fidelity 3D editing that operates in both one-shot and few-shot regime...

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL 21.08.2025

🤗 Upvotes: 68 | cs. AI, cs. CL Authors: Weizhen Li, Jianbo Lin, Zhuosong Jiang, Jingyi Cao, Xinpeng Liu, Jiayu Zhang, Zhenqiang Huang, Qianben Chen, Weichen Sun, Qiexiang Wang, Hongxuan Lu, Tianrui Qin, Chenghao Zhu, Yi Yao, Shuying Fan, Xiaowan Li, Tiannan Wang, Pai Liu, King Zhu, He Zhu, Dingfeng Shi, Piaohong Wang, Yeyi Guan, Xiangru Tang, Minghao Liu, Yuchen Eleanor Jiang, Jian Yang, Jiaheng...

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos 21.08.2025

🤗 Upvotes: 41 | cs. CV Authors: Chin-Yang Lin, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu Title: LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos Arxiv: http://arxiv.org/abs/2508.14041v1 Abstract: LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.