Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Enabling Scalable Oversight via Self-Evolving Critic 14.01.2025

🤗 Upvotes: 22 | cs. CL, cs. AI, cs. LG Authors: Zhengyang Tang, Ziniu Li, Zhenyang Xiao, Tian Ding, Ruoyu Sun, Benyou Wang, Dayiheng Liu, Fei Huang, Tianyu Liu, Bowen Yu, Junyang Lin Title: Enabling Scalable Oversight via Self-Evolving Critic Arxiv: http://arxiv.org/abs/2501.05727v1 Abstract: Despite their remarkable performance, the development of Large Language Models (LLMs) faces a critical ch...

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models 14.01.2025

🤗 Upvotes: 14 | cs. CL, cs. AI, cs. CV Authors: You Li, Heyu Huang, Chi Chen, Kaiyu Huang, Chao Huang, Zonghao Guo, Zhiyuan Liu, Jinan Xu, Yuhua Li, Ruixuan Li, Maosong Sun Title: Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2501.05767v2 Abstract: The recent advancement of Multimodal Large Language Models (MLLMs)...

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding 14.01.2025

🤗 Upvotes: 10 | cs. CV, cs. CL Authors: Xingyu Fu, Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Jianwei Yang, Dan Roth, Dinei Florencio, Cha Zhang Title: ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Arxiv: http://arxiv.org/abs/2501.05452v1 Abstract: Structured image understanding, such as interpreting tables and charts, requires strategically refocusin...

ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning 14.01.2025

🤗 Upvotes: 10 | cs. CV Authors: Yuzhou Huang, Ziyang Yuan, Quande Liu, Qiulin Wang, Xintao Wang, Ruimao Zhang, Pengfei Wan, Di Zhang, Kun Gai Title: ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning Arxiv: http://arxiv.org/abs/2501.04698v1 Abstract: Text-to-video generation has made remarkable advancements through diffusion models. However,...

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains 14.01.2025

🤗 Upvotes: 8 | cs. CL, cs. AI, cs. LG Authors: Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, Shuang Li, Igor Mordatch Title: Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains Arxiv: http://arxiv.org/abs/2501.05707v1 Abstract: Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlyin...

The GAN is dead; long live the GAN! A Modern GAN Baseline 11.01.2025

🤗 Upvotes: 27 | cs. LG, cs. CV Authors: Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov, James Tompkin Title: The GAN is dead; long live the GAN! A Modern GAN Baseline Arxiv: http://arxiv.org/abs/2501.05441v1 Abstract: There is a widely-spread claim that GANs are difficult to train, and GAN architectures in the literature are littered with empirical tricks. We provide evidence against this claim...

An Empirical Study of Autoregressive Pre-training from Videos 11.01.2025

🤗 Upvotes: 17 | cs. CV, cs. AI Authors: Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman, Christoph Feichtenhofer, Jitendra Malik Title: An Empirical Study of Autoregressive Pre-training from Videos Arxiv: http://arxiv.org/abs/2501.05453v1 Abstract: We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregress...

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives 11.01.2025

🤗 Upvotes: 10 | cs. CV, cs. RO Authors: Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, Liang Pan Title: Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Arxiv: http://arxiv.org/abs/2501.04003v1 Abstract: Recent advancements in Vision-Language Models (VLMs) have sparked interest in their...

Entropy-Guided Attention for Private LLMs 11.01.2025

🤗 Upvotes: 6 | cs. LG, cs. CR Authors: Nandan Kumar Jha, Brandon Reagen Title: Entropy-Guided Attention for Private LLMs Arxiv: http://arxiv.org/abs/2501.03489v2 Abstract: The pervasiveness of proprietary language models has raised critical privacy concerns, necessitating advancements in private inference (PI), where computations are performed directly on encrypted data without revealing users' s...

On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis 11.01.2025

🤗 Upvotes: 5 | cs. LG, cs. AI, cs. CC, cs. CV Authors: Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song Title: On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis Arxiv: http://arxiv.org/abs/2501.04377v1 Abstract: Recently, Visual Autoregressive ($\mathsf{VAR}$) Models introduced a groundbreaking advance...

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model 11.01.2025

🤗 Upvotes: 5 | cs. CL, cs. CV Authors: Gregor Geigle, Florian Schneider, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavaš Title: Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model Arxiv: http://arxiv.org/abs/2501.05122v1 Abstract: Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them strug...

SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution 11.01.2025

🤗 Upvotes: 4 | cs. CL Authors: Chengxing Xie, Bowen Li, Chang Gao, He Du, Wai Lam, Difan Zou, Kai Chen Title: SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution Arxiv: http://arxiv.org/abs/2501.05040v1 Abstract: Large Language Models (LLMs) have demonstrated remarkable proficiency across a variety of complex tasks. One significant application of LLMs is in ta...

Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models 11.01.2025

🤗 Upvotes: 3 | cs. CL Authors: Şaziye Betül Özateş, Tarık Emre Tıraş, Ece Elif Adak, Berat Doğan, Fatih Burak Karagöz, Efe Eren Genç, Esma F. Bilgin Taşdemir Title: Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models Arxiv: http://arxiv.org/abs/2501.04828v1 Abstract: This paper introduces foundational resources and models for natural language processin...

rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking 10.01.2025

🤗 Upvotes: 116 | cs. CL Authors: Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang, Youran Sun, Yi Zhu, Fan Yang, Mao Yang Title: rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking Arxiv: http://arxiv.org/abs/2501.04519v1 Abstract: We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o...

Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought 10.01.2025

🤗 Upvotes: 47 | cs. AI, cs. CL Authors: Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, Louis Castricato, Jan-Philipp Franken, Nick Haber, Chelsea Finn Title: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought Arxiv: http://arxiv.org/abs/2501.04682v1 Abstract: We propo...

URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics 10.01.2025

🤗 Upvotes: 38 | cs. CL, cs. AI, cs. LG Authors: Ruilin Luo, Zhuofan Zheng, Yifan Wang, Yiyao Yu, Xinzhe Ni, Zicheng Lin, Jin Zeng, Yujiu Yang Title: URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics Arxiv: http://arxiv.org/abs/2501.04686v1 Abstract: Chain-of-thought (CoT) reasoning has been widely applied in the mathematical reasoning of Large Language Models...

Agent Laboratory: Using LLM Agents as Research Assistants 10.01.2025

🤗 Upvotes: 38 | cs. HC, cs. AI, cs. CL, cs. LG Authors: Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, Emad Barsoum Title: Agent Laboratory: Using LLM Agents as Research Assistants Arxiv: http://arxiv.org/abs/2501.04227v1 Abstract: Historically, scientific discovery has been a lengthy and costly process, demanding substantial time and resource...

LLM4SR: A Survey on Large Language Models for Scientific Research 10.01.2025

🤗 Upvotes: 21 | cs. CL, cs. DL Authors: Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, Xinya Du Title: LLM4SR: A Survey on Large Language Models for Scientific Research Arxiv: http://arxiv.org/abs/2501.04306v1 Abstract: In recent years, the rapid advancement of Large Language Models (LLMs) has transformed the landscape of scientific research, offering unprecedented support across various stages of...

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection 10.01.2025

🤗 Upvotes: 16 | cs. AI, cs. CL, cs. HC Authors: Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, Fei Wu Title: InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection Arxiv: http://arxiv.org/abs/2501.04575v1 Abstract: Graphical User Interface (GUI) Agents, powered by multimodal large language models (ML...

SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images 10.01.2025

🤗 Upvotes: 12 | cs. CV, cs. GR Authors: Zixuan Huang, Mark Boss, Aaryaman Vasishta, James M. Rehg, Varun Jampani Title: SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images Arxiv: http://arxiv.org/abs/2501.04689v1 Abstract: We study the problem of single-image 3D object reconstruction. Recent works have diverged into two directions: regression-based modeling and generative m...

GeAR: Generation Augmented Retrieval 10.01.2025

🤗 Upvotes: 12 | cs. IR, cs. CL Authors: Haoyu Liu, Shaohan Huang, Jianfeng Liu, Yuefeng Zhan, Hao Sun, Weiwei Deng, Feng Sun, Furu Wei, Qi Zhang Title: GeAR: Generation Augmented Retrieval Arxiv: http://arxiv.org/abs/2501.02772v1 Abstract: Document retrieval techniques form the foundation for the development of large-scale information systems. The prevailing methodology is to construct a bi-encod...

Chirpy3D: Continuous Part Latents for Creative 3D Bird Generation 10.01.2025

🤗 Upvotes: 10 | cs. CV, cs. GR Authors: Kam Woh Ng, Jing Yang, Jia Wei Sii, Jiankang Deng, Chee Seng Chan, Yi-Zhe Song, Tao Xiang, Xiatian Zhu Title: Chirpy3D: Continuous Part Latents for Creative 3D Bird Generation Arxiv: http://arxiv.org/abs/2501.04144v1 Abstract: In this paper, we push the boundaries of fine-grained 3D generation into truly creative territory. Current methods either lack intri...

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization 10.01.2025

🤗 Upvotes: 5 | cs. LG, cs. AI, cs. CL, 68T45 Authors: Amitava Das, Suranjana Trivedy, Danush Khanna, Rajarshi Roy, Gurpreet Singh, Basab Ghosh, Yaswanth Narsupalli, Vinija Jain, Vasu Sharma, Aishwarya Naresh Reganti, Aman Chadha Title: DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Arxiv: http://arxiv.org/abs/2501.03271v2 Abstra...

REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models 09.01.2025

🤗 Upvotes: 51 | cs. CL, cs. LG Authors: Jian Hu Title: REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models Arxiv: http://arxiv.org/abs/2501.03262v1 Abstract: Reinforcement Learning from Human Feedback (RLHF) has emerged as a critical approach for aligning large language models with human preferences, witnessing rapid algorithmic evolution through methods such as Proxim...

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models 09.01.2025

🤗 Upvotes: 32 | cs. CV Authors: Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang, Lefan Wang, Xiaotao Gu, Shiyu Huang, Yuxiao Dong, Jie Tang Title: MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Arxiv: http://arxiv.org/abs/2501.02955v1 Abstract: In recent years, vision language models (VLMs) have made significant advancements in video un...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.