Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis 21.12.2024 21:08
🤗 Upvotes: 12 | cs. CV Authors: Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Qifeng Chen, Yujun Shen, Limin Wang Title: LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis Arxiv: http://arxiv.org/abs/2412.15214v1 Abstract: The intuitive nature of drag-based interaction has led to its growing adoption for controlling object trajectories in image-to-video synthesis. Still, ex...
DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation 21.12.2024 23:08
🤗 Upvotes: 8 | cs. CV, cs. AI, cs. GR Authors: Wang Zhao, Yan-Pei Cao, Jiale Xu, Yuejiang Dong, Ying Shan Title: DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation Arxiv: http://arxiv.org/abs/2412.15200v1 Abstract: Procedural Content Generation (PCG) is powerful in creating high-quality 3D contents, yet controlling it to produce desired shap...
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling 21.12.2024 24:09
🤗 Upvotes: 7 | cs. CL, cs. AI, cs. LG Authors: Zihan Liu, Yang Chen, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping Title: AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling Arxiv: http://arxiv.org/abs/2412.15084v1 Abstract: In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective rewa...
No More Adam: Learning Rate Scaling at Initialization is All You Need 20.12.2024 21:59
🤗 Upvotes: 177 | cs. LG, cs. AI Authors: Minghao Xu, Lichuan Xiang, Xu Cai, Hongkai Wen Title: No More Adam: Learning Rate Scaling at Initialization is All You Need Arxiv: http://arxiv.org/abs/2412.11768v2 Abstract: In this work, we question the necessity of adaptive gradient methods for training deep neural networks. SGD-SaI is a simple yet effective enhancement to stochastic gradient descent wi...
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference 20.12.2024 21:56
🤗 Upvotes: 36 | cs. CL, cs. AI Authors: Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, Iacopo Poli Title: Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Arx...
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks 20.12.2024 24:45
🤗 Upvotes: 30 | cs. CL Authors: Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, Graham Neubig Title: TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks Arxiv: http://...
AniDoc: Animation Creation Made Easier 20.12.2024 22:20
🤗 Upvotes: 29 | cs. CV Authors: Yihao Meng, Hao Ouyang, Hanlin Wang, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Zhiheng Liu, Yujun Shen, Huamin Qu Title: AniDoc: Animation Creation Made Easier Arxiv: http://arxiv.org/abs/2412.14173v1 Abstract: The production of 2D animation follows an industry-standard workflow, encompassing four essential stages: character design, keyframe animation, in-betweening, a...
FashionComposer: Compositional Fashion Image Generation 20.12.2024 19:47
🤗 Upvotes: 13 | cs. CV Authors: Sihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu, Hao Luo, Hengshuang Zhao Title: FashionComposer: Compositional Fashion Image Generation Arxiv: http://arxiv.org/abs/2412.14168v2 Abstract: We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, para...
GUI Agents: A Survey 20.12.2024 21:01
🤗 Upvotes: 11 | cs. AI, cs. HC Authors: Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A. Ross...
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning 20.12.2024 22:42
🤗 Upvotes: 10 | cs. LG, cs. RO Authors: Moritz Reuss, Jyothish Pari, Pulkit Agrawal, Rudolf Lioutikov Title: Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning Arxiv: http://arxiv.org/abs/2412.12953v1 Abstract: Diffusion Policies have become widely used in Imitation Learning, offering several appealing properties, such as generating multimodal and dis...
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation 20.12.2024 20:41
🤗 Upvotes: 10 | cs. CV Authors: Haotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng, Jiaming Sun, Minghuan Liu, Hujun Bao, Jiashi Feng, Xiaowei Zhou, Bingyi Kang Title: Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation Arxiv: http://arxiv.org/abs/2412.14015v1 Abstract: Prompts play a critical role in unleashing the power of language and vision foundation models for speci...
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces 20.12.2024 20:52
🤗 Upvotes: 9 | cs. CV Authors: Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, Saining Xie Title: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Arxiv: http://arxiv.org/abs/2412.14171v1 Abstract: Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language...
Are Your LLMs Capable of Stable Reasoning? 19.12.2024 24:11
🤗 Upvotes: 61 | cs. AI, cs. CL Authors: Junnan Liu, Hongwei Liu, Linchen Xiao, Ziyi Wang, Kuikun Liu, Songyang Gao, Wenwei Zhang, Songyang Zhang, Kai Chen Title: Are Your LLMs Capable of Stable Reasoning? Arxiv: http://arxiv.org/abs/2412.13147v2 Abstract: The rapid advancement of Large Language Models (LLMs) has demonstrated remarkable progress in complex reasoning tasks. However, a significant d...
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models 19.12.2024 22:34
🤗 Upvotes: 29 | cs. AI, cs. CL, cs. CV Authors: YiFan Zhang, Shanglin Lei, Runqi Qiao, Zhuoma GongQue, Xiaoshuai Song, Guanting Dong, Qiuna Tan, Zhe Wei, Peiqing Yang, Ye Tian, Yadong Xue, Xiaofei Wang, Honggang Zhang Title: Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models Arxiv: http://arxiv.org/abs/2412.12606v1 Abstract: The rapidly developing field...
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain 19.12.2024 23:15
🤗 Upvotes: 29 | cs. CL Authors: Shuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong Wen Title: OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain Arxiv: http://arxiv.org/abs/2412.13018v1 Abstract: As a typical and practical application of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) techniques have gained extensive attention, particularly in...
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations 19.12.2024 23:05
🤗 Upvotes: 21 | cs. CL Authors: Jeffrey Cheng, Benjamin Van Durme Title: Compressed Chain of Thought: Efficient Reasoning Through Dense Representations Arxiv: http://arxiv.org/abs/2412.13171v1 Abstract: Chain-of-thought (CoT) decoding enables language models to improve reasoning performance at the cost of high generation latency in decoding. Recent proposals have explored variants of contemplatio...
Emergence of Abstractions: Concept Encoding and Decoding Mechanism for In-Context Learning in Transformers 19.12.2024 22:52
🤗 Upvotes: 9 | cs. CL, cs. AI, cs. LG Authors: Seungwook Han, Jinyeop Song, Jeff Gore, Pulkit Agrawal Title: Emergence of Abstractions: Concept Encoding and Decoding Mechanism for In-Context Learning in Transformers Arxiv: http://arxiv.org/abs/2412.12276v2 Abstract: Humans distill complex experiences into fundamental abstractions that enable rapid learning and adaptation. Similarly, autoregressiv...
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration 19.12.2024 20:44
🤗 Upvotes: 7 | cs. CV Authors: Mark Endo, Xiaohan Wang, Serena Yeung-Levy Title: Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration Arxiv: http://arxiv.org/abs/2412.13180v1 Abstract: Recent works on accelerating Vision-Language Models show that strong performance can be maintained across a variety of vision-language tasks despite highly compressing visual...
Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents 19.12.2024 23:53
🤗 Upvotes: 5 | cs. LG, cs. AI, cs. CV Authors: Yifei Zhou, Qianlan Yang, Kaixiang Lin, Min Bai, Xiong Zhou, Yu-Xiong Wang, Sergey Levine, Erran Li Title: Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents Arxiv: http://arxiv.org/abs/2412.13194v1 Abstract: The vision of a broadly capable and goal-directed agent, such as an Internet-browsing agent in the...
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation 19.12.2024 23:12
🤗 Upvotes: 4 | cs. CL Authors: Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami, Ryan A. Rossi, Dinesh Manocha Title: VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation Arxiv: http://arxiv.org/abs/2412.10704v1 Abstract: Understanding information from a collection of multiple documents, particularly those with visually rich elements,...
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner 19.12.2024 20:27
🤗 Upvotes: 2 | cs. CV Authors: Yufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Nanxuan Zhao, Jing Shi, Tong Sun Title: SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Arxiv: http://arxiv.org/abs/2412.10533v1 Abstract: We present SUGAR, a zero-shot method for subject-driven video customization. Given an input image, SUGAR is capable of generating videos for the subject contained in the image...
Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion 19.12.2024 20:33
🤗 Upvotes: 2 | cs. CV, cs. LG Authors: Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, Anton Obukhov Title: Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion Arxiv: http://arxiv.org/abs/2412.13389v1 Abstract: Depth completion upgrades sparse depth measurements into dense depth maps guided by a conventional image. Existing methods...
Byte Latent Transformer: Patches Scale Better Than Tokens 18.12.2024 25:08
🤗 Upvotes: 39 | cs. CL Authors: Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer Title: Byte Latent Transformer: Patches Scale Better Than Tokens Arxiv: http://arxiv.org/abs/2412.09871v1 Abstract: We introduce the Byte Latent Transformer (BLT),...
RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation 18.12.2024 21:46
🤗 Upvotes: 25 | cs. CL, cs. AI, cs. IR Authors: Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu, Zhonghua Li, Qi Ye, Zhicheng Dou Title: RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation Arxiv: http://arxiv.org/abs/2412.11919v1 Abstract: Large language models (LLMs) exhibit remarkable generative capabilities but often suffer from hallucinations. Retriev...
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models 18.12.2024 21:10
🤗 Upvotes: 25 | cs. CV, cs. AI, cs. CL Authors: Fan Zhang, Shulin Tian, Ziqi Huang, Yu Qiao, Ziwei Liu Title: Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Arxiv: http://arxiv.org/abs/2412.09645v2 Abstract: Recent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, eval...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.