Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Magma: A Foundation Model for Multimodal AI Agents 20.02.2025

🤗 Upvotes: 30 | cs. CV, cs. AI, cs. HC, cs. LG, cs. RO Authors: Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, Lars Liden, Jianfeng Gao Title: Magma: A Foundation Model for Multimodal AI Agents Arxiv: http://arxiv.org/abs/2502.13130v1 Abstract: We present Magma, a foundation model that serves multimodal AI age...

Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation 20.02.2025

🤗 Upvotes: 29 | cs. CV Authors: Bencheng Liao, Hongyuan Tao, Qian Zhang, Tianheng Cheng, Yingyue Li, Haoran Yin, Wenyu Liu, Xinggang Wang Title: Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation Arxiv: http://arxiv.org/abs/2502.13145v1 Abstract: Recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance but face deployment c...

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation 20.02.2025

🤗 Upvotes: 27 | cs. RO, cs. AI, cs. CV Authors: Zekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong, Xinqiang Yu, Jingwen Li, Lingyun Xu, Baoyu Li, Xialin He, Guofan Fan, Jiazhao Zhang, Jiawei He, Jiayuan Gu, Xin Jin, Kaisheng Ma, Zhizheng Zhang, He Wang, Li Yi Title: SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation Arxiv: http://arxiv.org/abs/2502.13143v1 Abstra...

SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models 20.02.2025

🤗 Upvotes: 26 | cs. CL Authors: Seanie Lee, Dong Bok Lee, Dominik Wagner, Minki Kang, Haebin Seong, Tobias Bocklet, Juho Lee, Sung Ju Hwang Title: SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models Arxiv: http://arxiv.org/abs/2502.12464v1 Abstract: Deploying large language models (LLMs) in real-world applications requires robust safety guard...

You Do Not Fully Utilize Transformer's Representation Capacity 20.02.2025

🤗 Upvotes: 25 | cs. LG, cs. CL Authors: Gleb Gerasimov, Yaroslav Aksenov, Nikita Balagansky, Viacheslav Sinii, Daniil Gavrilov Title: You Do Not Fully Utilize Transformer's Representation Capacity Arxiv: http://arxiv.org/abs/2502.09245v1 Abstract: In contrast to RNNs, which compress previous tokens into a single hidden state, Transformers can attend to all previous tokens directly. However, stand...

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 19.02.2025

🤗 Upvotes: 68 | cs. CL, cs. AI, cs. LG Authors: Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng Title: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Arxiv: http://arxiv.org/abs/2502.11089v1 Abstract: Long-context modelin...

Learning Getting-Up Policies for Real-World Humanoid Robots 19.02.2025

🤗 Upvotes: 32 | cs. RO, cs. LG Authors: Xialin He, Runpei Dong, Zixuan Chen, Saurabh Gupta Title: Learning Getting-Up Policies for Real-World Humanoid Robots Arxiv: http://arxiv.org/abs/2502.12152v1 Abstract: Automatic fall recovery is a crucial prerequisite before humanoid robots can be reliably deployed. Hand-designing controllers for getting up is difficult because of the varied configurations...

SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? 19.02.2025

🤗 Upvotes: 27 | cs. LG, cs. SE Authors: Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes Heidecke Title: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? Arxiv: http://arxiv.org/abs/2502.12115v1 Abstract: We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at \$1 million USD total in r...

CRANE: Reasoning with constrained LLM generation 19.02.2025

🤗 Upvotes: 17 | cs. PL, cs. LG Authors: Debangshu Banerjee, Tarun Suresh, Shubham Ugare, Sasa Misailovic, Gagandeep Singh Title: CRANE: Reasoning with constrained LLM generation Arxiv: http://arxiv.org/abs/2502.09061v1 Abstract: Code generation, symbolic math reasoning, and other tasks require LLMs to produce outputs that are both syntactically and semantically correct. Constrained LLM generation...

How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training 19.02.2025

🤗 Upvotes: 16 | cs. LG, cs. AI, cs. CL, cs. CV, cs. HC Authors: Yixin Ou, Yunzhi Yao, Ningyu Zhang, Hui Jin, Jiacheng Sun, Shumin Deng, Zhenguo Li, Huajun Chen Title: How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training Arxiv: http://arxiv.org/abs/2502.11196v1 Abstract: Despite exceptional capabilities in knowledge-intensive tasks, Large Language Models (L...

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation 19.02.2025

🤗 Upvotes: 15 | cs. CV Authors: Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang, Minghao Xu, Wentao Zhang, Bin Cui Title: HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation Arxiv: http://arxiv.org/abs/2502.12148v1 Abstract: The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with power...

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models 19.02.2025

🤗 Upvotes: 14 | cs. LG, cs. AI Authors: Zhenxing Mi, Kuan-Chieh Wang, Guocheng Qian, Hanrong Ye, Runtao Liu, Sergey Tulyakov, Kfir Aberman, Dan Xu Title: I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models Arxiv: http://arxiv.org/abs/2502.10458v1 Abstract: This paper presents ThinkDiff, a novel alignment paradigm that empowers text-to-image diffusion models...

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors 19.02.2025

🤗 Upvotes: 11 | cs. LG, cs. CL Authors: Bohan Lyu, Siqiao Huang, Zichen Liang Title: SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors Arxiv: http://arxiv.org/abs/2502.11167v1 Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, such as code understanding and code generation. However, an equally importan...

Region-Adaptive Sampling for Diffusion Transformers 18.02.2025

🤗 Upvotes: 46 | cs. CV, cs. AI Authors: Ziming Liu, Yifan Yang, Chengruidong Zhang, Yiqi Zhang, Lili Qiu, Yang You, Yuqing Yang Title: Region-Adaptive Sampling for Diffusion Transformers Arxiv: http://arxiv.org/abs/2502.10389v1 Abstract: Diffusion models (DMs) have become the leading choice for generative tasks across diverse domains. However, their reliance on multiple sequential forward passes...

Large Language Diffusion Models 18.02.2025

🤗 Upvotes: 44 | cs. CL, cs. LG Authors: Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li Title: Large Language Diffusion Models Arxiv: http://arxiv.org/abs/2502.09992v1 Abstract: Autoregressive models (ARMs) are widely regarded as the cornerstone of large language models (LLMs). We challenge this notion by introducing LLaDA, a dif...

The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks 18.02.2025

🤗 Upvotes: 41 | cs. AI Authors: Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, Joseph E. Gonzalez Title: The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks Arxiv: http://arxiv.org/abs/2502.08235v1 Ab...

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 18.02.2025

🤗 Upvotes: 38 | cs. CV, cs. CL Authors: Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, Yu Zhou, Deshan Sun, Deyu Zhou, Jian Zhou, Kaijun Tan, Kang An, Mei Chen, Wei Ji, Qiling Wu, Wen Sun, Xin Han, Yanan Wei, Zheng Ge, Aojie Li, Bin Wang, Bizhu Huang, Bo Wang, Brian Li, Changxing Miao, Chen Xu, Chenfei Wu, Chenguang Y...

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models 18.02.2025

🤗 Upvotes: 27 | cs. CV Authors: Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Yan, Wenye Lin, Gyungin Shin, Qiaochu Yang, Anh Totti Nguyen, Kai Han, Samuel Albanie Title: ZeroBench: An Impossib...

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment 18.02.2025

🤗 Upvotes: 22 | cs. CL, cs. CV Authors: Yi-Fan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu, Peiyan Li, Jianshu Zeng, Wulin Xie, Yang Shi, Huanyu Zhang, Junkang Wu, Xue Wang, Yibo Hu, Bin Wen, Fan Yang, Zhang Zhang, Tingting Gao, Di Zhang, Liang Wang, Rong Jin, Tieniu Tan Title: MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Arxiv: http://arxiv.org/abs/2502.10391v1 Abstract: Despite notabl...

ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation 18.02.2025

🤗 Upvotes: 12 | cs. CV, cs. GR Authors: Rotem Shalev-Arkushin, Rinon Gal, Amit H. Bermano, Ohad Fried Title: ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation Arxiv: http://arxiv.org/abs/2502.09411v1 Abstract: Diffusion models enable high-quality and diverse visual content synthesis. However, they struggle to generate rare or unseen concepts. To address this challenge, we ex...

Diverse Inference and Verification for Advanced Reasoning 18.02.2025

🤗 Upvotes: 11 | cs. AI Authors: Iddo Drori, Gaston Longhitano, Mao Mao, Seunghwan Hyun, Yuke Zhang, Sungjun Park, Zachary Meeks, Xin-Yu Zhang, Ben Segev, Howard Yong, Nakul Verma, Avi Shporer, Alon Amit, Madeleine Udell Title: Diverse Inference and Verification for Advanced Reasoning Arxiv: http://arxiv.org/abs/2502.09955v1 Abstract: Reasoning LLMs such as OpenAI o1, o3 and DeepSeek R1 have made...

Precise Parameter Localization for Textual Generation in Diffusion Models 18.02.2025

🤗 Upvotes: 10 | cs. CV Authors: Łukasz Staniszewski, Bartosz Cywiński, Franziska Boenisch, Kamil Deja, Adam Dziedzic Title: Precise Parameter Localization for Textual Generation in Diffusion Models Arxiv: http://arxiv.org/abs/2502.09935v1 Abstract: Novel diffusion models can synthesize photo-realistic images with integrated high-quality text. Surprisingly, we demonstrate through attention activat...

DarwinLM: Evolutionary Structured Pruning of Large Language Models 18.02.2025

🤗 Upvotes: 9 | cs. LG, cs. CL Authors: Shengkun Tang, Oliver Sieberling, Eldar Kurtic, Zhiqiang Shen, Dan Alistarh Title: DarwinLM: Evolutionary Structured Pruning of Large Language Models Arxiv: http://arxiv.org/abs/2502.07780v1 Abstract: Large Language Models (LLMs) have achieved significant success across various NLP tasks. However, their massive computational costs limit their widespread use,...

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU 15.02.2025

🤗 Upvotes: 62 | cs. CL, cs. LG Authors: Heejun Lee, Geon Park, Jaduk Suh, Sung Ju Hwang Title: InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Arxiv: http://arxiv.org/abs/2502.08910v1 Abstract: In modern large language models (LLMs), handling very long context lengths presents significant challenges as it causes slower inference speeds and increased memory cos...

The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding 15.02.2025

🤗 Upvotes: 35 | cs. CL, cs. AI, cs. CV, cs. LG Authors: Mo Yu, Lemao Liu, Junjie Wu, Tsz Ting Chung, Shunchi Zhang, Jiangnan Li, Dit-Yan Yeung, Jie Zhou Title: The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding Arxiv: http://arxiv.org/abs/2502.08946v1 Abstract: In a systematic way, we investigate a widely asked question: Do LLMs really understand wha...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.