Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Slow Perception: Let's Perceive Geometric Figures Step-by-step 01.01.2025 23:19
🤗 Upvotes: 5 | cs. CV Authors: Haoran Wei, Youyang Yin, Yumeng Li, Jia Wang, Liang Zhao, Jianjian Sun, Zheng Ge, Xiangyu Zhang Title: Slow Perception: Let's Perceive Geometric Figures Step-by-step Arxiv: http://arxiv.org/abs/2412.20631v1 Abstract: Recently, "visual o1" began to enter people's vision, with expectations that this slow-thinking design can solve visual reasoning tasks, especially geo...
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs 31.12.2024 23:19
🤗 Upvotes: 53 | cs. CL, cs. AI, cs. LG Authors: Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, Benyou Wang Title: HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs Arxiv: http://arxiv.org/abs/2412.18925v1 Abstract: The breakthrough of OpenAI o1 highlights the potential of enhancing reasoning to improve LLM. Yet, most research in reasoning has focu...
1.58-bit FLUX 31.12.2024 22:59
🤗 Upvotes: 24 | cs. CV, cs. AI, cs. LG Authors: Chenglin Yang, Celong Liu, Xueqing Deng, Dongwon Kim, Xing Mei, Xiaohui Shen, Liang-Chieh Chen Title: 1.58-bit FLUX Arxiv: http://arxiv.org/abs/2412.18653v1 Abstract: We present 1.58-bit FLUX, the first successful approach to quantizing the state-of-the-art text-to-image generation model, FLUX.1-dev, using 1.58-bit weights (i.e., values in {-1, 0, +...
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 31.12.2024 17:30
🤗 Upvotes: 17 | cs. CL, cs. AI, cs. CV, cs. LG, cs. MM, eess. AS Authors: Liang Chen, Zekun Wang, Shuhuai Ren, Lei Li, Haozhe Zhao, Yunshui Li, Zefan Cai, Hongcheng Guo, Lei Zhang, Yizhe Xiong, Yichi Zhang, Ruoyu Wu, Qingxiu Dong, Ge Zhang, Jian Yang, Lingwei Meng, Shujie Hu, Yulong Chen, Junyang Lin, Shuai Bai, Andreas Vlachos, Xu Tan, Minjia Zhang, Wen Xiao, Aaron Yee, Tianyu Liu, Baobao Chang...
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models 31.12.2024 23:16
🤗 Upvotes: 11 | cs. CV Authors: Zehan Wang, Ziang Zhang, Tianyu Pang, Chao Du, Hengshuang Zhao, Zhou Zhao Title: Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models Arxiv: http://arxiv.org/abs/2412.18605v1 Abstract: Orientation is a key attribute of objects, crucial for understanding their spatial pose and arrangement in images. However, practical solutions for...
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment 31.12.2024 25:03
🤗 Upvotes: 11 | cs. CV Authors: Ziang Yan, Zhilin Li, Yinan He, Chenting Wang, Kunchang Li, Xinhao Li, Xiangyu Zeng, Zilei Wang, Yali Wang, Yu Qiao, Limin Wang, Yi Wang Title: Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment Arxiv: http://arxiv.org/abs/2412.19326v1 Abstract: Current multimodal large language models (MLLMs) struggle with fine-grai...
From Elements to Design: A Layered Approach for Automatic Graphic Design Composition 31.12.2024 22:38
🤗 Upvotes: 11 | cs. CV Authors: Jiawei Lin, Shizhao Sun, Danqing Huang, Ting Liu, Ji Li, Jiang Bian Title: From Elements to Design: A Layered Approach for Automatic Graphic Design Composition Arxiv: http://arxiv.org/abs/2412.19712v1 Abstract: In this work, we investigate automatic design composition from multimodal graphic elements. Although recent studies have developed various generative models...
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models 31.12.2024 23:53
🤗 Upvotes: 8 | cs. CV Authors: Tao Wu, Yong Zhang, Xiaodong Cun, Zhongang Qi, Junfu Pu, Huanzhang Dou, Guangcong Zheng, Ying Shan, Xi Li Title: VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Arxiv: http://arxiv.org/abs/2412.19645v2 Abstract: Zero-shot customized video generation has gained significant attention due to its substantial applicatio...
The Superposition of Diffusion Models Using the Itô Density Estimator 31.12.2024 23:22
🤗 Upvotes: 8 | cs. LG Authors: Marta Skreta, Lazar Atanackovic, Avishek Joey Bose, Alexander Tong, Kirill Neklyudov Title: The Superposition of Diffusion Models Using the Itô Density Estimator Arxiv: http://arxiv.org/abs/2412.17762v1 Abstract: The Cambrian explosion of easily accessible pre-trained diffusion models suggests a demand for methods that combine multiple different pre-trained diffusio...
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging 31.12.2024 19:03
🤗 Upvotes: 6 | cs. CL Authors: Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee Title: Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging Arxiv: http://arxiv.org/abs/2412.19512v1 Abstract: Fine-tuning large language models (LLMs) for downstream tasks is a widely adopted approach, but it often leads to safety degradation in safety-aligned LLMs. Curren...
CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era 31.12.2024 25:00
🤗 Upvotes: 3 | cs. CL, cs. AI, cs. DB Authors: Yanlin Feng, Simone Papicchio, Sajjadur Rahman Title: CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era Arxiv: http://arxiv.org/abs/2412.18702v1 Abstract: Retrieval from graph data is crucial for augmenting large language models (LLM) with both open-domain knowledge and private enterprise data, and it is al...
YuLan-Mini: An Open Data-efficient Language Model 28.12.2024 19:39
🤗 Upvotes: 27 | cs. CL Authors: Yiwen Hu, Huatong Song, Jia Deng, Jiapeng Wang, Jie Chen, Kun Zhou, Yutao Zhu, Jinhao Jiang, Zican Dong, Wayne Xin Zhao, Ji-Rong Wen Title: YuLan-Mini: An Open Data-efficient Language Model Arxiv: http://arxiv.org/abs/2412.17743v2 Abstract: Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the comple...
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression 28.12.2024 21:47
🤗 Upvotes: 17 | cs. CL Authors: Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou Title: A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression Arxiv: http://arxiv.org/abs/2412.17483v1 Abstract: In this work, we provide a thorough investigation of gist-based context compression methods to improve l...
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 28.12.2024 21:10
🤗 Upvotes: 4 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Wan-Cyuan Fan, Tanzila Rahman, Leonid Sigal Title: MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Arxiv: http://arxiv.org/abs/2412.18072v1 Abstract: With advances in foundational and vision-language models, and effective fine-tuning techniques, a large number of both general and special-purpose models have been develo...
Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation 28.12.2024 22:18
🤗 Upvotes: 2 | cs. IR, cs. AI Authors: Yucong Luo, Qitao Qin, Hao Zhang, Mingyue Cheng, Ruiran Yan, Kefan Wang, Jie Ouyang Title: Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation Arxiv: http://arxiv.org/abs/2412.18176v1 Abstract: Sequential recommendation (SR) systems have evolved significantly over the past decade, transitioning from traditiona...
DepthLab: From Partial to Complete 26.12.2024 22:07
🤗 Upvotes: 21 | cs. CV Authors: Zhiheng Liu, Ka Leong Cheng, Qiuyu Wang, Shuzhe Wang, Hao Ouyang, Bin Tan, Kai Zhu, Yujun Shen, Qifeng Chen, Ping Luo Title: DepthLab: From Partial to Complete Arxiv: http://arxiv.org/abs/2412.18153v1 Abstract: Missing values remain a common challenge for depth data across its wide range of applications, stemming from various causes like incomplete data acquisition...
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization 26.12.2024 21:34
🤗 Upvotes: 20 | cs. AI, cs. CL Authors: Ermo Hua, Che Jiang, Xingtai Lv, Kaiyan Zhang, Ning Ding, Youbang Sun, Biqing Qi, Yuchen Fan, Xue Kai Zhu, Bowen Zhou Title: Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization Arxiv: http://arxiv.org/abs/2412.17739v1 Abstract: Extending the context length of Language Models (LMs) by improving Rotary Position Embed...
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation 26.12.2024 22:13
🤗 Upvotes: 10 | cs. CV, cs. AI, cs. MM Authors: Minghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu, Zhaoyang Zhang, Yong Zhang, Ying Shan, Xiangyu Yue Title: DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation Arxiv: http://arxiv.org/abs/2412.18597v1 Abstract: Sora-like video generation models have achieved remarkable progre...
In Case You Missed It: ARC 'Challenge' Is Not That Challenging 26.12.2024 24:19
🤗 Upvotes: 8 | cs. CL, cs. AI Authors: Łukasz Borchmann Title: In Case You Missed It: ARC 'Challenge' Is Not That Challenging Arxiv: http://arxiv.org/abs/2412.17758v1 Abstract: ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. Although some researchers have quie...
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing 26.12.2024 20:56
🤗 Upvotes: 8 | cs. LG Authors: Ziteng Wang, Jianfei Chen, Jun Zhu Title: ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing Arxiv: http://arxiv.org/abs/2412.14711v1 Abstract: Sparsely activated Mixture-of-Experts (MoE) models are widely adopted to scale up model capacity without increasing the computation budget. However, vanilla TopK routers are trained in a discontinuous, non-diff...
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval 26.12.2024 22:17
🤗 Upvotes: 6 | cs. CL Authors: Aakash Mahalingam, Vinesh Kumar Gande, Aman Chadha, Vinija Jain, Divya Chaudhary Title: SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval Arxiv: http://arxiv.org/abs/2412.15443v1 Abstract: Retrieval-Augmented Generation (RAG) systems have become pivotal in leveraging vast corpora to generate informed and contextually relevant responses,...
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models 26.12.2024 26:24
🤗 Upvotes: 5 | cs. CV Authors: Minghao Chen, Roman Shapovalov, Iro Laina, Tom Monnier, Jianyuan Wang, David Novotny, Andrea Vedaldi Title: PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models Arxiv: http://arxiv.org/abs/2412.18608v1 Abstract: Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures. These asse...
MotiF: Making Text Count in Image Animation with Motion Focal Loss 26.12.2024 22:35
🤗 Upvotes: 3 | cs. CV, cs. AI Authors: Shijie Wang, Samaneh Azadi, Rohit Girdhar, Saketh Rambhatla, Chen Sun, Xi Yin Title: MotiF: Making Text Count in Image Animation with Motion Focal Loss Arxiv: http://arxiv.org/abs/2412.16153v1 Abstract: Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image ani...
Bridging the Data Provenance Gap Across Text, Speech and Video 26.12.2024 25:29
🤗 Upvotes: 3 | cs. AI, cs. CL, cs. CY, cs. LG, cs. MM Authors: Shayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary, Joanna Materzynska, William Brannon, Robert Mahari, Manan Dey, Mohammed Hamdy, Nayan Saxena, Ahmad Mustafa Anis, Emad A. Alghamdi, Vu Minh Chien, Naana Obeng-Marnu, Da Yin, Kun Qian, Yizhi Li, Minnie Liang, An Dinh, Shrestha Mohanty, Deividas Mataciunas, Tobin South, Jiangu...
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response 25.12.2024 21:39
🤗 Upvotes: 64 | cs. CL, cs. AI Authors: Junyu Luo, Xiao Luo, Kaize Ding, Jingyang Yuan, Zhiping Xiao, Ming Zhang Title: RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response Arxiv: http://arxiv.org/abs/2412.14922v1 Abstract: Supervised fine-tuning (SFT) plays a crucial role in adapting large language models (LLMs) to specific domains or tasks. However, as demonstr...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.