Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Can sparse autoencoders be used to decompose and interpret steering vectors? 15.11.2024 21:54
🤗 Paper Upvotes: 6 | cs. LG, cs. AI, cs. CL Authors: Harry Mayne, Yushi Yang, Adam Mahdi Title: Can sparse autoencoders be used to decompose and interpret steering vectors? Arxiv: http://arxiv.org/abs/2411.08790v1 Abstract: Steering vectors are a promising approach to control the behaviour of large language models. However, their underlying mechanisms remain poorly understood. While sparse autoen...
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation 15.11.2024 18:59
🤗 Paper Upvotes: 5 | cs. AI, cs. MM, cs. SD, eess. AS Authors: Yungang Yi, Weihua Li, Matthew Kuo, Quan Bai Title: PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation Arxiv: http://arxiv.org/abs/2411.08307v1 Abstract: Music generation has progressed significantly, especially in the domain of audio generation. However, generating symbo...
SAMPart3D: Segment Any Part in 3D Objects 14.11.2024 20:51
🤗 Paper Upvotes: 18 | cs. CV Authors: Yunhan Yang, Yukun Huang, Yuan-Chen Guo, Liangjun Lu, Xiaoyang Wu, Edmund Y. Lam, Yan-Pei Cao, Xihui Liu Title: SAMPart3D: Segment Any Part in 3D Objects Arxiv: http://arxiv.org/abs/2411.07184v1 Abstract: 3D part segmentation is a crucial and challenging task in 3D perception, playing a vital role in applications such as robotics, 3D generation, and 3D editin...
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation 14.11.2024 22:30
🤗 Paper Upvotes: 14 | cs. CV, cs. AI, cs. CL Authors: Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai yu, Liang Zhao, Yisong Wang, Jiaying Liu, Chong Ruan Title: JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation Arxiv: http://arxiv.org/abs/2411.07975v1 Abstract: We pre...
Stronger Models are NOT Stronger Teachers for Instruction Tuning 14.11.2024 27:47
🤗 Paper Upvotes: 13 | cs. AI, cs. CL Authors: Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Radha Poovendran Title: Stronger Models are NOT Stronger Teachers for Instruction Tuning Arxiv: http://arxiv.org/abs/2411.07133v2 Abstract: Instruction tuning has been widely adopted to ensure large language models (LLMs) follow user instructions effectively. The resulting instruction-following...
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions 14.11.2024 20:39
🤗 Paper Upvotes: 11 | cs. CV, cs. AI Authors: Anas Awadalla, Le Xue, Manli Shu, An Yan, Jun Wang, Senthil Purushwalkam, Sheng Shen, Hannah Lee, Oscar Lo, Jae Sung Park, Etash Guha, Silvio Savarese, Ludwig Schmidt, Yejin Choi, Caiming Xiong, Ran Xu Title: BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions Arxiv: http://arxiv.org/abs/2411.07461v1 Abstract: We introduce BLIP3-KALE, a dataset...
Scaling Properties of Diffusion Models for Perceptual Tasks 14.11.2024 25:09
🤗 Paper Upvotes: 7 | cs. CV, cs. AI Authors: Rahul Ravishankar, Zeeshan Patel, Jathushan Rajasegaran, Jitendra Malik Title: Scaling Properties of Diffusion Models for Perceptual Tasks Arxiv: http://arxiv.org/abs/2411.08034v2 Abstract: In this paper, we argue that iterative computation with diffusion models offers a powerful paradigm for not only generation but also visual perception tasks. We uni...
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings 14.11.2024 22:22
🤗 Paper Upvotes: 5 | cs. CV, cs. AI, cs. LG Authors: Aditya Sanghi, Aliasghar Khani, Pradyumna Reddy, Arianna Rampini, Derek Cheung, Kamal Rahimi Malekshan, Kanika Madan, Hooman Shayani Title: Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings Arxiv: http://arxiv.org/abs/2411.08017v1 Abstract: Large-scale 3D generative models require substantial...
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models 13.11.2024 23:36
🤗 Paper Upvotes: 44 | cs. CV, cs. AI, cs. GR, cs. LG Authors: Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, Gal Chechik Title: Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models Arxiv: http://arxiv.org/abs/2411.07232v2 Abstract: Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance be...
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision 13.11.2024 19:47
🤗 Paper Upvotes: 39 | cs. CV, cs. AI Authors: Cong Wei, Zheyang Xiong, Weiming Ren, Xinrun Du, Ge Zhang, Wenhu Chen Title: OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision Arxiv: http://arxiv.org/abs/2411.07199v1 Abstract: Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or m...
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models 13.11.2024 21:15
🤗 Paper Upvotes: 30 | cs. CL Authors: Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Hui Huang, Weixun Wang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Xuepeng Liu, Dekai Sun, Wenbo Su, Bo Zheng Title: Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models Arxiv: http://arxiv.org/abs/2411.07140v1 Abstract: New LLM evaluation benchmarks are important to align with...
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework 13.11.2024 20:43
🤗 Paper Upvotes: 28 | cs. CL Authors: Yew Ken Chia, Liying Cheng, Hou Pong Chan, Chaoqun Liu, Maojia Song, Sharifah Mahani Aljunied, Soujanya Poria, Lidong Bing Title: M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework Arxiv: http://arxiv.org/abs/2411.06176v1 Abstract: The ability to understand and answer questions over documents can be...
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models 13.11.2024 24:47
🤗 Paper Upvotes: 21 | cs. CV, cs. LG Authors: NVIDIA, :, Yuval Atzmon, Maciej Bala, Yogesh Balaji, Tiffany Cai, Yin Cui, Jiaojiao Fan, Yunhao Ge, Siddharth Gururani, Jacob Huffman, Ronald Isaac, Pooya Jannaty, Tero Karras, Grace Lam, J. P. Lewis, Aaron Licata, Yen-Chen Lin, Ming-Yu Liu, Qianli Ma, Arun Mallya, Ashlee Martino-Tarr, Doug Mendez, Seungjun Nah, Chris Pruett, Fitsum Reda, Jiaming Song...
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models 13.11.2024 24:32
🤗 Paper Upvotes: 18 | cs. SE, cs. LG Authors: Nizar Islah, Justine Gehring, Diganta Misra, Eilif Muller, Irina Rish, Terry Yue Zhuo, Massimo Caccia Title: GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models Arxiv: http://arxiv.org/abs/2411.05830v1 Abstract: The rapid evolution of software libraries presents a significant challenge for code generation models, which...
Watermark Anything with Localized Messages 13.11.2024 23:25
🤗 Paper Upvotes: 11 | cs. CV, cs. CR Authors: Tom Sander, Pierre Fernandez, Alain Durmus, Teddy Furon, Matthijs Douze Title: Watermark Anything with Localized Messages Arxiv: http://arxiv.org/abs/2411.07231v1 Abstract: Image watermarking methods are not tailored to handle small watermarked areas. This restricts applications in real-world scenarios where parts of the image may come from different...
Autoregressive Models in Vision: A Survey 13.11.2024 22:52
🤗 Paper Upvotes: 3 | cs. CV, cs. CL Authors: Jing Xiong, Gongye Liu, Lun Huang, Chengyue Wu, Taiqiang Wu, Yao Mu, Yuan Yao, Hui Shen, Zhongwei Wan, Jinfa Huang, Chaofan Tao, Shen Yan, Huaxiu Yao, Lingpeng Kong, Hongxia Yang, Mi Zhang, Guillermo Sapiro, Jiebo Luo, Ping Luo, Ngai Wong Title: Autoregressive Models in Vision: A Survey Arxiv: http://arxiv.org/abs/2411.05902v1 Abstract: Autoregressive...
LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation 12.11.2024 25:29
🤗 Paper Upvotes: 15 | cs. CV, cs. CL Authors: Weiquan Huang, Aoqi Wu, Yifan Yang, Xufang Luo, Yuqing Yang, Liang Hu, Qi Dai, Xiyang Dai, Dongdong Chen, Chong Luo, Lili Qiu Title: LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation Arxiv: http://arxiv.org/abs/2411.04997v1 Abstract: CLIP is one of the most important multimodal foundational models today. What powers CLIP's capabili...
Balancing Pipeline Parallelism with Vocabulary Parallelism 12.11.2024 23:34
🤗 Paper Upvotes: 10 | cs. DC Authors: Man Tsung Yeung, Penghui Qi, Min Lin, Xinyi Wan Title: Balancing Pipeline Parallelism with Vocabulary Parallelism Arxiv: http://arxiv.org/abs/2411.05288v1 Abstract: Pipeline parallelism is widely used to scale the training of transformer-based large language models, various works have been done to improve its throughput and memory footprint. In this paper, we...
StdGEN: Semantic-Decomposed 3D Character Generation from Single Images 12.11.2024 21:47
🤗 Paper Upvotes: 10 | cs. CV Authors: Yuze He, Yanning Zhou, Wang Zhao, Zhongkai Wu, Kaiwen Xiao, Wei Yang, Yong-Jin Liu, Xiao Han Title: StdGEN: Semantic-Decomposed 3D Character Generation from Single Images Arxiv: http://arxiv.org/abs/2411.05738v1 Abstract: We present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling br...
DELIFT: Data Efficient Language model Instruction Fine Tuning 12.11.2024 21:17
🤗 Paper Upvotes: 5 | cs. CL Authors: Ishika Agarwal, Krishnateja Killamsetty, Lucian Popa, Marina Danilevksy Title: DELIFT: Data Efficient Language model Instruction Fine Tuning Arxiv: http://arxiv.org/abs/2411.04425v2 Abstract: Fine-tuning large language models (LLMs) is essential for enhancing their performance on specific tasks but is often resource-intensive due to redundant or uninformative...
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study 12.11.2024 25:06
🤗 Paper Upvotes: 4 | cs. SE, cs. AI, cs. LG Authors: André Storhaug, Jingyue Li Title: Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study Arxiv: http://arxiv.org/abs/2411.02462v1 Abstract: The advent of large language models (LLMs) like GitHub Copilot has significantly enhanced programmers' productivity, particularly in code generation. However,...
RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models 12.11.2024 22:22
🤗 Paper Upvotes: 3 | cs. CV, cs. AI Authors: Maya Varma, Jean-Benoit Delbrouck, Zhihong Chen, Akshay Chaudhari, Curtis Langlotz Title: RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models Arxiv: http://arxiv.org/abs/2411.04097v1 Abstract: Fine-tuned vision-language models (VLMs) often capture spurious correlations between image features and textual attribute...
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities 12.11.2024 24:01
🤗 Paper Upvotes: 3 | cs. CL Authors: Zhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu, Yoon Kim Title: The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities Arxiv: http://arxiv.org/abs/2411.04986v1 Abstract: Modern language models can process inputs across diverse languages and modalities. We hypothesize that models acquire this capab...
Improving the detection of technical debt in Java source code with an enriched dataset 12.11.2024 26:17
🤗 Paper Upvotes: 2 | cs. SE Authors: Nam Le Hai, Anh M. T. Bui, Phuong T. Nguyen, Davide Di Ruscio, Rick Kazman Title: Improving the detection of technical debt in Java source code with an enriched dataset Arxiv: http://arxiv.org/abs/2411.05457v1 Abstract: Technical debt (TD) is a term used to describe the additional work and costs that emerge when developers have opted for a quick and easy solut...
OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models 09.11.2024 22:46
🤗 Paper Upvotes: 69 | cs. CL, cs. PL Authors: Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, Wei Chu Title: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models Arxiv: http://arxiv.org/abs/2411.04905v1 Abst...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.