Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 23.11.2024

🤗 Paper Upvotes: 42 | cs. CL, cs. CV Authors: Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, Jifeng Dai Title: Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Arxiv: http://arxiv.org/abs/2411.10442v1 Abstract: Existing open-source multimodal large language models (MLLMs) gene...

Multimodal Autoregressive Pre-training of Large Vision Encoders 23.11.2024

🤗 Paper Upvotes: 23 | cs. CV, cs. LG Authors: Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor Guilherme Turrisi da Costa, Louis Béthune, Zhe Gan, Alexander T Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby Title: Multimodal Autoregressive Pre-training of Large Vision Encoders Arxiv: http://arxiv.or...

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions 23.11.2024

🤗 Paper Upvotes: 23 | cs. CL Authors: Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang Title: Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions Arxiv: http://arxiv.org/abs/2411.14405v1 Abstract: Currently OpenAI o1 has sparked a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1...

Hymba: A Hybrid-head Architecture for Small Language Models 23.11.2024

🤗 Paper Upvotes: 20 | cs. CL, cs. AI, cs. LG Authors: Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, Yingyan Lin, Jan Kautz, Pavlo Molchanov Title: Hymba: A Hybrid-head Architecture for Small Language Models Arxiv: http://arxiv.org/abs/2411.13676v1 Abstract: We propose Hymba, a family...

Natural Language Reinforcement Learning 23.11.2024

🤗 Paper Upvotes: 15 | cs. LG, cs. AI, cs. CL Authors: Xidong Feng, Ziyu Wan, Haotian Fu, Bo Liu, Mengyue Yang, Girish A. Koushik, Zhiyuan Hu, Ying Wen, Jun Wang Title: Natural Language Reinforcement Learning Arxiv: http://arxiv.org/abs/2411.14251v1 Abstract: Reinforcement Learning (RL) mathematically formulates decision-making with Markov Decision Process (MDP). With MDPs, researchers have achiev...

OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs 23.11.2024

🤗 Paper Upvotes: 15 | cs. CL, cs. AI, cs. DL, cs. IR, cs. LG Authors: Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D'arcy, David Wadden, Matt Latzke, Minyang Tian, Pan Ji, Shengyan Liu, Hao Tong, Bohao Wu, Yanyu Xiong, Luke Zettlemoyer, Graham Neubig, Dan Weld, Doug Downey, Wen-tau Yih, Pang Wei Koh, Hannaneh H...

Ultra-Sparse Memory Network 23.11.2024

🤗 Paper Upvotes: 14 | cs. LG Authors: Zihao Huang, Qiyang Min, Hongzhi Huang, Defa Zhu, Yutao Zeng, Ran Guo, Xun Zhou Title: Ultra-Sparse Memory Network Arxiv: http://arxiv.org/abs/2411.12364v1 Abstract: It is widely acknowledged that the performance of Transformer models is exponentially related to their number of parameters and computational complexity. While approaches like Mixture of Experts...

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 23.11.2024

🤗 Paper Upvotes: 10 | cs. CV Authors: Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, Ziwei Liu Title: Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2411.14432v1 Abstract: Large Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought...

Stable Flow: Vital Layers for Training-Free Image Editing 23.11.2024

🤗 Paper Upvotes: 7 | cs. CV, cs. GR, cs. LG Authors: Omri Avrahami, Or Patashnik, Ohad Fried, Egor Nemchinov, Kfir Aberman, Dani Lischinski, Daniel Cohen-Or Title: Stable Flow: Vital Layers for Training-Free Image Editing Arxiv: http://arxiv.org/abs/2411.14430v1 Abstract: Diffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional U...

Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models 23.11.2024

🤗 Paper Upvotes: 6 | cs. CL, cs. AI, cs. LG Authors: Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, Neel Nanda Title: Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models Arxiv: http://arxiv.org/abs/2411.14257v1 Abstract: Hallucinations in large language models are a widespread problem, yet the mechanisms behind whether models will hallucinate are poorly under...

SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration 22.11.2024

🤗 Paper Upvotes: 35 | cs. LG, cs. AI, cs. CV, cs. NE, cs. PF Authors: Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen Title: SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration Arxiv: http://arxiv.org/abs/2411.10958v1 Abstract: Although quantization for linear layers has been widely used, its application to accelerate the at...

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models 22.11.2024

🤗 Paper Upvotes: 23 | cs. CV Authors: Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, Ziwei Liu Title: VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models Arxiv: http://arxiv.org/abs/2411.13503v1 Abstract: Video ge...

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation 22.11.2024

🤗 Paper Upvotes: 14 | cs. CV, cs. AI, cs. CL, cs. MM Authors: Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, Junnan Li Title: VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Arxiv: http://arxiv.org/abs/2411.13281v1 Abstract: Large multimodal models (LMMs) with advanced video analysis capabilities have recently gar...

SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory 22.11.2024

🤗 Paper Upvotes: 12 | cs. CV Authors: Cheng-Yen Yang, Hsiang-Wei Huang, Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang Title: SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory Arxiv: http://arxiv.org/abs/2411.11922v1 Abstract: The Segment Anything Model 2 (SAM 2) has demonstrated strong performance in object segmentation tasks but faces challenges in vis...

Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents 22.11.2024

🤗 Paper Upvotes: 9 | cs. AI Authors: Yu Gu, Boyuan Zheng, Boyu Gou, Kai Zhang, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, Huan Sun, Yu Su Title: Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents Arxiv: http://arxiv.org/abs/2411.06559v1 Abstract: Language agents have demonstrated promising capabilities in automating web-based tasks, though their curr...

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training 22.11.2024

🤗 Paper Upvotes: 7 | cs. CL Authors: Haonan Wang, Qian Liu, Chao Du, Tongyao Zhu, Cunxiao Du, Kenji Kawaguchi, Tianyu Pang Title: When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Arxiv: http://arxiv.org/abs/2411.13476v1 Abstract: Extending context window sizes allows large language models (LLMs) to process longer sequences and handle more complex tasks. Rotary Pos...

Stylecodes: Encoding Stylistic Information For Image Generation 22.11.2024

🤗 Paper Upvotes: 6 | cs. CV Authors: Ciara Rowles Title: Stylecodes: Encoding Stylistic Information For Image Generation Arxiv: http://arxiv.org/abs/2411.12811v1 Abstract: Diffusion models excel in image generation, but controlling them remains a challenge. We focus on the problem of style-conditioned image generation. Although example images work, they are cumbersome: srefs (style-reference code...

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models 22.11.2024

🤗 Paper Upvotes: 3 | cs. CV, cs. AI Authors: Vipula Rawte, Sarthak Jain, Aarush Sinha, Garv Kaushik, Aman Bansal, Prathiksha Rumale Vishwanath, Samyak Rajesh Jain, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha, Amit P. Sheth, Amitava Das Title: ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models Arxiv: http://arxiv.org/abs/2411.10867v1 Abstract: Latest dev...

Loss-to-Loss Prediction: Scaling Laws for All Datasets 22.11.2024

🤗 Paper Upvotes: 2 | cs. LG, cs. AI, cs. CL, stat. ML Authors: David Brandfonbrener, Nikhil Anand, Nikhil Vyas, Eran Malach, Sham Kakade Title: Loss-to-Loss Prediction: Scaling Laws for All Datasets Arxiv: http://arxiv.org/abs/2411.12925v1 Abstract: While scaling laws provide a reliable methodology for predicting train loss across compute scales for a single data distribution, less is known about...

ORID: Organ-Regional Information Driven Framework for Radiology Report Generation 22.11.2024

🤗 Paper Upvotes: 2 | cs. CV Authors: Tiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai Title: ORID: Organ-Regional Information Driven Framework for Radiology Report Generation Arxiv: http://arxiv.org/abs/2411.13025v1 Abstract: The objective of Radiology Report Generation (RRG) is to automatically generate coherent textual analyses of diseases based on radiological image...

SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization 21.11.2024

🤗 Paper Upvotes: 13 | cs. CV Authors: Hongrui Jia, Chaoya Jiang, Haiyang Xu, Wei Ye, Mengfan Dong, Ming Yan, Ji Zhang, Fei Huang, Shikun Zhang Title: SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization Arxiv: http://arxiv.org/abs/2411.11909v1 Abstract: As language models continue to scale, Large Language Models (LLMs) have exhib...

Continuous Speculative Decoding for Autoregressive Image Generation 21.11.2024

🤗 Paper Upvotes: 13 | cs. CV Authors: Zili Wang, Robert Zhang, Kun Ding, Qi Yang, Fei Li, Shiming Xiang Title: Continuous Speculative Decoding for Autoregressive Image Generation Arxiv: http://arxiv.org/abs/2411.11925v1 Abstract: Continuous-valued Autoregressive (AR) image generation models have demonstrated notable superiority over their discrete-token counterparts, showcasing considerable recon...

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements 21.11.2024

🤗 Paper Upvotes: 11 | cs. CV Authors: M. Arda Aydın, Efe Mert Çırpar, Elvin Abdinli, Gozde Unal, Yusuf H. Sahin Title: ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Arxiv: http://arxiv.org/abs/2411.12044v1 Abstract: Recent advances in foundational Vision Language Models (VLMs) have reshaped the evaluation paradigm in computer vision tasks....

FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations 21.11.2024

🤗 Paper Upvotes: 10 | cs. GR, cs. CV Authors: Hmrishav Bandyopadhyay, Yi-Zhe Song Title: FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations Arxiv: http://arxiv.org/abs/2411.10818v1 Abstract: Sketch animations offer a powerful medium for visual storytelling, from simple flip-book doodles to professional studio productions. While traditional animation requires teams of skilled ar...

Soft Robotic Dynamic In-Hand Pen Spinning 21.11.2024

🤗 Paper Upvotes: 8 | cs. RO Authors: Yunchao Yao, Uksang Yoo, Jean Oh, Christopher G. Atkeson, Jeffrey Ichnowski Title: Soft Robotic Dynamic In-Hand Pen Spinning Arxiv: http://arxiv.org/abs/2411.12734v1 Abstract: Dynamic in-hand manipulation remains a challenging task for soft robotic systems that have demonstrated advantages in safe compliant interactions but struggle with high-speed dynamic tas...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.