Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views 11.12.2024 22:08
🤗 Upvotes: 4 | cs. CV, cs. GR Authors: Antoine Guédon, Tomoki Ichikawa, Kohei Yamashita, Ko Nishino Title: MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views Arxiv: http://arxiv.org/abs/2412.06767v1 Abstract: We present a novel appearance model that simultaneously realizes explicit high-quality 3D surface mesh recovery and photorealistic novel view synt...
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment 10.12.2024 20:24
🤗 Upvotes: 33 | cs. CV Authors: Yibin Wang, Zhiyu Tan, Junyan Wang, Xiaomeng Yang, Cheng Jin, Hao Li Title: LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment Arxiv: http://arxiv.org/abs/2412.04814v1 Abstract: Recent advancements in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos wit...
EXAONE 3.5: Series of Large Language Models for Real-world Use Cases 10.12.2024 22:00
🤗 Upvotes: 31 | cs. CL Authors: LG AI Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Yountae Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Youchul Kim, Edward Hwayoung Lee, Haeju Lee, Honglak Lee, Jinsik Lee, Kyungmin Lee,...
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 10.12.2024 21:58
🤗 Upvotes: 30 | cs. CL, cs. CV Authors: Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, Xiang Yue Title: MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Arxiv: http://arxiv.org/abs/2412.05237v1 Abstract: Open-source multimodal large language models (MLLMs) have shown significant potential in a broad range of multimo...
APOLLO: SGD-like Memory, AdamW-level Performance 10.12.2024 19:44
🤗 Upvotes: 27 | cs. LG, cs. AI, cs. PF Authors: Hanqing Zhu, Zhenyu Zhang, Wenyan Cong, Xi Liu, Sem Park, Vikas Chandra, Bo Long, David Z. Pan, Zhangyang Wang, Jinwon Lee Title: APOLLO: SGD-like Memory, AdamW-level Performance Arxiv: http://arxiv.org/abs/2412.05270v2 Abstract: Large language models (LLMs) are notoriously memory-intensive during training, particularly with the popular AdamW optimi...
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion 10.12.2024 20:09
🤗 Upvotes: 19 | cs. CV Authors: Trong-Tung Nguyen, Quang Nguyen, Khoi Nguyen, Anh Tran, Cuong Pham Title: SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion Arxiv: http://arxiv.org/abs/2412.04301v2 Abstract: Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-ba...
Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 10.12.2024 20:18
🤗 Upvotes: 18 | cs. RO, cs. AI, cs. CL, cs. CV, cs. LG Authors: Yi Chen, Yuying Ge, Yizhuo Li, Yixiao Ge, Mingyu Ding, Ying Shan, Xihui Liu Title: Moto: Latent Motion Token as the Bridging Language for Robot Manipulation Arxiv: http://arxiv.org/abs/2412.04445v1 Abstract: Recent developments in Large Language Models pre-trained on extensive corpora have shown significant success in various natural...
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration 10.12.2024 22:51
🤗 Upvotes: 13 | cs. CV Authors: Kaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin, Yu Wang, Xihui Liu Title: GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Arxiv: http://arxiv.org/abs/2412.04440v1 Abstract: Text-to-video generation models have shown significant progress in the recent years. However, they still struggle with generating complex dynamic scenes based on...
Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction 10.12.2024 21:05
🤗 Upvotes: 12 | cs. CV Authors: Jixuan Fan, Wanhua Li, Yifei Han, Yansong Tang Title: Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction Arxiv: http://arxiv.org/abs/2412.04887v1 Abstract: 3D Gaussian Splatting has demonstrated notable success in large-scale scene reconstruction, but challenges persist due to high training memory consumption and storage ov...
CompCap: Improving Multimodal Large Language Models with Composite Captions 10.12.2024 21:55
🤗 Upvotes: 11 | cs. CV, cs. AI, cs. LG Authors: Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh, Qifan Wang, David Yang, ShengYun Peng, Hanchao Yu, Shen Yan, Xuewen Zhang, Baosheng He Title: CompCap: Improving Multimodal Large Language Models with Composite Captions Arxiv: http://arxiv.org/abs/2412.05243v1 Abstract: How well can Multimodal Large Language Models (MLLMs) understand co...
VisionZip: Longer is Better but Not Necessary in Vision Language Models 08.12.2024 21:48
🤗 Upvotes: 83 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, Jiaya Jia Title: VisionZip: Longer is Better but Not Necessary in Vision Language Models Arxiv: http://arxiv.org/abs/2412.04467v1 Abstract: Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them...
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 08.12.2024 19:23
🤗 Upvotes: 46 | cs. CV, cs. AI Authors: Jiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li, Jianfeng Gao, Tianyi Zhou, Bin Xiao Title: Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Arxiv: http://arxiv.org/abs/2412.04424v1 Abstract: We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representat...
NVILA: Efficient Frontier Visual Language Models 08.12.2024 19:31
🤗 Upvotes: 36 | cs. CV Authors: Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Vishwesh Nath, Jinyi Hu, Sifei Liu, Ranjay Krishna, Daguang Xu, Xiaolong Wang, Pavlo Molchanov, Jan Kautz, Hongxu Yin, Song Han, Yao Lu Title: NVILA: Efficient F...
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction 08.12.2024 20:40
🤗 Upvotes: 32 | cs. CL Authors: Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, Caiming Xiong Title: Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Arxiv: http://arxiv.org/abs/2412.04454v1 Abstract: Graphical User Interfaces (GUIs) are critical to human-computer interaction, yet automating GUI tasks remains challenging due to the com...
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection 08.12.2024 22:57
🤗 Upvotes: 32 | cs. RO, cs. AI, cs. CV, cs. LG Authors: Enshen Zhou, Qi Su, Cheng Chi, Zhizheng Zhang, Zhongyuan Wang, Tiejun Huang, Lu Sheng, He Wang Title: Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection Arxiv: http://arxiv.org/abs/2412.04455v1 Abstract: Automatic detection and prevention of open-set failures are crucial in closed-loop r...
Evaluating Language Models as Synthetic Data Generators 08.12.2024 21:02
🤗 Upvotes: 30 | cs. CL Authors: Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig Title: Evaluating Language Models as Synthetic Data Generators Arxiv: http://arxiv.org/abs/2412.03679v1 Abstract: Given the increasing use of synthetic data in language model (LM) post-training, an LM's ability to gen...
A Noise is Worth Diffusion Guidance 08.12.2024 21:20
🤗 Upvotes: 25 | cs. CV, cs. AI, cs. LG Authors: Donghoon Ahn, Jiwon Kang, Sanghyun Lee, Jaewon Min, Minjae Kim, Wooseok Jang, Hyoungwon Cho, Sayak Paul, SeonHwa Kim, Eunju Cha, Kyong Hwan Jin, Seungryong Kim Title: A Noise is Worth Diffusion Guidance Arxiv: http://arxiv.org/abs/2412.03895v1 Abstract: Diffusion models excel in generating high-quality images. However, current diffusion models strug...
Structured 3D Latents for Scalable and Versatile 3D Generation 08.12.2024 23:37
🤗 Upvotes: 22 | cs. CV Authors: Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, Jiaolong Yang Title: Structured 3D Latents for Scalable and Versatile 3D Generation Arxiv: http://arxiv.org/abs/2412.01506v1 Abstract: We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LAT...
Negative Token Merging: Image-based Adversarial Feature Guidance 08.12.2024 19:33
🤗 Upvotes: 21 | cs. CV, cs. AI, cs. GR, cs. LG, stat. ML Authors: Jaskirat Singh, Lindsey Li, Weijia Shi, Ranjay Krishna, Yejin Choi, Pang Wei Koh, Michael F. Cohen, Stephen Gould, Liang Zheng, Luke Zettlemoyer Title: Negative Token Merging: Image-based Adversarial Feature Guidance Arxiv: http://arxiv.org/abs/2412.01339v2 Abstract: Text-based adversarial guidance using a negative prompt has emerg...
MV-Adapter: Multi-view Consistent Image Generation Made Easy 08.12.2024 21:23
🤗 Upvotes: 17 | cs. CV Authors: Zehuan Huang, Yuan-Chen Guo, Haoran Wang, Ran Yi, Lizhuang Ma, Yan-Pei Cao, Lu Sheng Title: MV-Adapter: Multi-view Consistent Image Generation Made Easy Arxiv: http://arxiv.org/abs/2412.03632v1 Abstract: Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to (...
ShowUI: One Vision-Language-Action Model for GUI Visual Agent 28.11.2024 24:33
🤗 Paper Upvotes: 48 | cs. CV, cs. AI, cs. CL, cs. HC Authors: Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Weixian Lei, Lijuan Wang, Mike Zheng Shou Title: ShowUI: One Vision-Language-Action Model for GUI Visual Agent Arxiv: http://arxiv.org/abs/2411.17465v1 Abstract: Building Graphical User Interface (GUI) assistants holds significant promise for enhancing hum...
Star Attention: Efficient LLM Inference over Long Sequences 28.11.2024 20:34
🤗 Paper Upvotes: 32 | cs. CL, cs. AI, cs. LG Authors: Shantanu Acharya, Fei Jia, Boris Ginsburg Title: Star Attention: Efficient LLM Inference over Long Sequences Arxiv: http://arxiv.org/abs/2411.17116v1 Abstract: Inference with Transformer-based Large Language Models (LLMs) on long sequences is both costly and slow due to the quadratic complexity of the self-attention mechanism. We introduce Sta...
Pathways on the Image Manifold: Image Editing via Video Generation 28.11.2024 25:04
🤗 Paper Upvotes: 23 | cs. CV, cs. AI, cs. LG Authors: Noam Rotstein, Gal Yona, Daniel Silver, Roy Velich, David Bensaïd, Ron Kimmel Title: Pathways on the Image Manifold: Image Editing via Video Generation Arxiv: http://arxiv.org/abs/2411.16819v1 Abstract: Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as...
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 28.11.2024 26:29
🤗 Paper Upvotes: 15 | cs. CV, cs. AI, cs. CL Authors: Chaoyou Fu, Yi-Fan Zhang, Shukang Yin, Bo Li, Xinyu Fang, Sirui Zhao, Haodong Duan, Xing Sun, Ziwei Liu, Liang Wang, Caifeng Shan, Ran He Title: MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Arxiv: http://arxiv.org/abs/2411.15296v1 Abstract: As a prominent direction of Artificial General Intelligence (AGI), Multimodal Lar...
Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration 28.11.2024 22:13
🤗 Paper Upvotes: 14 | cs. CV Authors: Yuhang Han, Xuyang Liu, Pengxiang Ding, Donglin Wang, Honggang Chen, Qingsen Yan, Siteng Huang Title: Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration Arxiv: http://arxiv.org/abs/2411.17686v1 Abstract: To accelerate the inference of heavy Multimodal Large Language Models (MLLMs), this study rethinks the current la...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.