Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning 09.11.2024

🤗 Paper Upvotes: 50 | cs. CV, cs. AI, cs. GR, cs. LG Authors: David Junhao Zhang, Roni Paiss, Shiran Zada, Nikhil Karnad, David E. Jacobs, Yael Pritch, Inbar Mosseri, Mike Zheng Shou, Neal Wadhwa, Nataniel Ruiz Title: ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning Arxiv: http://arxiv.org/abs/2411.05003v1 Abstract: Recently, breakthroughs in vid...

BitNet a4.8: 4-bit Activations for 1-bit LLMs 09.11.2024

🤗 Paper Upvotes: 41 | cs. CL, cs. LG Authors: Hongyu Wang, Shuming Ma, Furu Wei Title: BitNet a4.8: 4-bit Activations for 1-bit LLMs Arxiv: http://arxiv.org/abs/2411.04965v1 Abstract: Recent research on the 1-bit Large Language Models (LLMs), such as BitNet b1.58, presents a promising direction for reducing the inference cost of LLMs while maintaining their performance. In this work, we introduce...

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion 09.11.2024

🤗 Paper Upvotes: 27 | cs. CV, cs. AI, cs. GR Authors: Wenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen, Yueqi Duan, Jun Zhang, Yikai Wang Title: DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion Arxiv: http://arxiv.org/abs/2411.04928v1 Abstract: In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D sc...

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models 09.11.2024

🤗 Paper Upvotes: 25 | cs. CL Authors: Weixin Liang, Lili Yu, Liang Luo, Srinivasan Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Xi Victoria Lin Title: Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Arxiv: http://arxiv.org/abs/2411.04996v1 Abstract: The development of large language models (LLMs) has expanded...

TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation 09.11.2024

🤗 Paper Upvotes: 20 | cs. CV Authors: Wenhao Wang, Yi Yang Title: TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation Arxiv: http://arxiv.org/abs/2411.04709v1 Abstract: Video generation models are revolutionizing content creation, with image-to-video models drawing increasing attention due to their enhanced controllability, visual consistency, and practical a...

Thanos: Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model 09.11.2024

🤗 Paper Upvotes: 15 | cs. CL Authors: Young-Jun Lee, Dokyong Lee, Junyoung Youn, Kyeongjin Oh, Ho-Jin Choi Title: Thanos: Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model Arxiv: http://arxiv.org/abs/2411.04496v1 Abstract: To increase social bonding with interlocutors, humans naturally acquire the ability to respond appropriately in a given situation by considering w...

Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks? 09.11.2024

🤗 Paper Upvotes: 14 | cs. CL Authors: Jonathan Roberts, Kai Han, Samuel Albanie Title: Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks? Arxiv: http://arxiv.org/abs/2411.05000v1 Abstract: As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. In many real-world tasks, decisions depend on d...

DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation 09.11.2024

🤗 Paper Upvotes: 12 | cs. RO, cs. LG Authors: Peiqi Liu, Zhanqiu Guo, Mohit Warke, Soumith Chintala, Chris Paxton, Nur Muhammad Mahi Shafiullah, Lerrel Pinto Title: DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation Arxiv: http://arxiv.org/abs/2411.04999v1 Abstract: Significant progress has been made in open-vocabulary mobile manipulation, where the goal is for a ro...

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos 09.11.2024

🤗 Paper Upvotes: 12 | cs. CV Authors: Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao, Eric Xing, Fahad Shahbaz Khan, Salman Khan Title: VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos Arxiv: http://arxiv.org/abs/2411.04923v1 Abstract: Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing...

Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination 08.11.2024

🤗 Paper Upvotes: 33 | cs. CV, cs. AI, cs. CL, cs. MM Authors: Dingjie Song, Sicheng Lai, Shunian Chen, Lichao Sun, Benyou Wang Title: Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination Arxiv: http://arxiv.org/abs/2411.03823v1 Abstract: The rapid progression of multimodal large language models (MLLMs) has demonstrated superior performance on various multimodal...

Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level 08.11.2024

🤗 Paper Upvotes: 26 | cs. LG, cs. AI Authors: Antoine Grosnit, Alexandre Maraval, James Doran, Giuseppe Paolo, Albert Thomas, Refinath Shahul Hameed Nabeezath Beevi, Jonas Gonzalez, Khyati Khandelwal, Ignacio Iacobacci, Abdelhakim Benechehab, Hamza Cherkaoui, Youssef Attia El-Hili, Kun Shao, Jianye Hao, Jun Yao, Balazs Kegl, Haitham Bou-Ammar, Jun Wang Title: Large Language Models Orchestrating S...

Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models 08.11.2024

🤗 Paper Upvotes: 10 | cs. CL, cs. AI, cs. LG Authors: Zhijian Zhuo, Ya Wang, Yutao Zeng, Xiaoqing Li, Xun Zhou, Jinwen Ma Title: Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models Arxiv: http://arxiv.org/abs/2411.03884v1 Abstract: Transformers have found extensive applications across various domains due to the powerful fitting capabilities. This success can be pa...

Self-Consistency Preference Optimization 08.11.2024

🤗 Paper Upvotes: 5 | cs. CL, cs. AI, cs. LG Authors: Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Yu Title: Self-Consistency Preference Optimization Arxiv: http://arxiv.org/abs/2411.04109v1 Abstract: Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing r...

From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond 08.11.2024

🤗 Paper Upvotes: 3 | cs. CL Authors: Harsha Nori, Naoto Usuyama, Nicholas King, Scott Mayer McKinney, Xavier Fernandes, Sheng Zhang, Eric Horvitz Title: From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond Arxiv: http://arxiv.org/abs/2411.03590v1 Abstract: Run-time steering strategies like Medprompt are valuable for guiding large language models (LLMs...

HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems 07.11.2024

🤗 Paper Upvotes: 34 | cs. IR Authors: Jiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang, Weipeng Chen, Ji-Rong Wen Title: HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems Arxiv: http://arxiv.org/abs/2411.02959v1 Abstract: Retrieval-Augmented Generation (RAG) has been shown to improve knowledge capabilities and alleviate the hallucination problem of LLMs. The Web...

LLaMo: Large Language Model-based Molecular Graph Assistant 07.11.2024

🤗 Paper Upvotes: 13 | cs. LG, cs. AI, q-bio. MN Authors: Jinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. Kim Title: LLaMo: Large Language Model-based Molecular Graph Assistant Arxiv: http://arxiv.org/abs/2411.00871v1 Abstract: Large Language Models (LLMs) have demonstrated remarkable generalization and instruction-following capabilities with instruction tuning. The advancements in LLMs and ins...

DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution 07.11.2024

🤗 Paper Upvotes: 10 | cs. RO, cs. AI, cs. LG Authors: Yang Yue, Yulin Wang, Bingyi Kang, Yizeng Han, Shenzhi Wang, Shiji Song, Jiashi Feng, Gao Huang Title: DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution Arxiv: http://arxiv.org/abs/2411.02359v1 Abstract: MLLMs have demonstrated remarkable comprehension and reasoning capabilities with complex language...

Controlling Language and Diffusion Models by Transporting Activations 07.11.2024

🤗 Paper Upvotes: 8 | cs. LG, cs. AI, cs. CL, cs. CV, 68T07, 49Q22, I.2.6; I.2.7; I.4.8 Authors: Pau Rodriguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Marco Cuturi, Xavier Suau Title: Controlling Language and Diffusion Models by Transporting Activations Arxiv: http://arxiv.org/abs/2410.23054v1 Abstract: The increasing capabilities of large generative models and their ever mo...

Sample-Efficient Alignment for LLMs 07.11.2024

🤗 Paper Upvotes: 8 | cs. LG, cs. AI, cs. CL Authors: Zichen Liu, Changyu Chen, Chao Du, Wee Sun Lee, Min Lin Title: Sample-Efficient Alignment for LLMs Arxiv: http://arxiv.org/abs/2411.01493v1 Abstract: We study methods for efficiently aligning large language models (LLMs) with human preferences given budgeted online feedback. We first formulate the LLM alignment problem in the frame of contextua...

DreamPolish: Domain Score Distillation With Progressive Geometry Generation 07.11.2024

🤗 Paper Upvotes: 6 | cs. CV, cs. AI Authors: Yean Cheng, Ziqi Cai, Ming Ding, Wendi Zheng, Shiyu Huang, Yuxiao Dong, Jie Tang, Boxin Shi Title: DreamPolish: Domain Score Distillation With Progressive Geometry Generation Arxiv: http://arxiv.org/abs/2411.01602v1 Abstract: We introduce DreamPolish, a text-to-3D generation model that excels in producing refined geometry and high-quality textures. In...

Adaptive Length Image Tokenization via Recurrent Allocation 07.11.2024

🤗 Paper Upvotes: 4 | cs. CV, cs. AI, cs. LG, cs. RO Authors: Shivam Duggal, Phillip Isola, Antonio Torralba, William T. Freeman Title: Adaptive Length Image Tokenization via Recurrent Allocation Arxiv: http://arxiv.org/abs/2411.02393v1 Abstract: Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intellig...

GarVerseLOD: High-Fidelity 3D Garment Reconstruction from a Single In-the-Wild Image using a Dataset with Levels of Details 07.11.2024

🤗 Paper Upvotes: 3 | cs. CV, cs. GR Authors: Zhongjin Luo, Haolin Liu, Chenghong Li, Wanghao Du, Zirong Jin, Wanhu Sun, Yinyu Nie, Weikai Chen, Xiaoguang Han Title: GarVerseLOD: High-Fidelity 3D Garment Reconstruction from a Single In-the-Wild Image using a Dataset with Levels of Details Arxiv: http://arxiv.org/abs/2411.03047v1 Abstract: Neural implicit functions have brought impressive advances...

Zebra-Llama: A Context-Aware Large Language Model for Democratizing Rare Disease Knowledge 07.11.2024

🤗 Paper Upvotes: 3 | cs. CL Authors: Karthik Soman, Andrew Langdon, Catalina Villouta, Chinmay Agrawal, Lashaw Salta, Braian Peetoom, Gianmarco Bellucci, Orion J Buske Title: Zebra-Llama: A Context-Aware Large Language Model for Democratizing Rare Disease Knowledge Arxiv: http://arxiv.org/abs/2411.02657v1 Abstract: Rare diseases present unique challenges in healthcare, often suffering from delaye...

Inference Optimal VLMs Need Only One Visual Token but Larger Models 07.11.2024

🤗 Paper Upvotes: 2 | cs. CV, cs. AI, cs. LG Authors: Kevin Y. Li, Sachin Goyal, Joao D. Semedo, J. Zico Kolter Title: Inference Optimal VLMs Need Only One Visual Token but Larger Models Arxiv: http://arxiv.org/abs/2411.03312v1 Abstract: Vision Language Models (VLMs) have demonstrated strong capabilities across various visual understanding and reasoning tasks. However, their real-world deployment...

AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents 06.11.2024

🤗 Paper Upvotes: 40 | cs. AI Authors: Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, Yuxiao Dong Title: AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents Arxiv: http://arxiv.org/abs/2410.24024v2 Abstract: Autonomous agents have become increasingly important for interacting with the real world. Android agents, in parti...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.