Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners 25.12.2024

🤗 Upvotes: 29 | cs. AI, cs. CL, cs. LG Authors: Weihao Zeng, Yuzhen Huang, Lulu Zhao, Yijun Wang, Zifei Shan, Junxian He Title: B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Arxiv: http://arxiv.org/abs/2412.17256v1 Abstract: In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their o...

Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching 25.12.2024

🤗 Upvotes: 26 | cs. CV, cs. LG Authors: Enshu Liu, Xuefei Ning, Yu Wang, Zinan Lin Title: Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching Arxiv: http://arxiv.org/abs/2412.17153v2 Abstract: Autoregressive (AR) models have achieved state-of-the-art performance in text and image generation but suffer from slow generation due to the token-by-token process. W...

Diving into Self-Evolving Training for Multimodal Reasoning 25.12.2024

🤗 Upvotes: 23 | cs. CL, cs. AI, cs. CV, cs. LG Authors: Wei Liu, Junlong Li, Xiwen Zhang, Fan Zhou, Yu Cheng, Junxian He Title: Diving into Self-Evolving Training for Multimodal Reasoning Arxiv: http://arxiv.org/abs/2412.17451v1 Abstract: Reasoning ability is essential for Large Multimodal Models (LMMs). In the absence of multimodal chain-of-thought annotated data, self-evolving training, where t...

Deliberation in Latent Space via Differentiable Cache Augmentation 25.12.2024

🤗 Upvotes: 16 | cs. CL, cs. AI, cs. LG Authors: Luyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie, Arthur Szlam Title: Deliberation in Latent Space via Differentiable Cache Augmentation Arxiv: http://arxiv.org/abs/2412.17747v1 Abstract: Techniques enabling large language models (LLMs) to "think more" by generating and attending to intermediate reasoning steps have shown promise in solving complex p...

Large Motion Video Autoencoding with Cross-modal Video VAE 25.12.2024

🤗 Upvotes: 15 | cs. CV Authors: Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen Title: Large Motion Video Autoencoding with Cross-modal Video VAE Arxiv: http://arxiv.org/abs/2412.17805v1 Abstract: Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying imag...

OpenAI o1 System Card 25.12.2024

🤗 Upvotes: 12 | cs. AI Authors: OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennett, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrew Duberstein, Andrew Kondrich, Andrey...

Revisiting In-Context Learning with Long Context Language Models 25.12.2024

🤗 Upvotes: 12 | cs. CL, cs. AI, cs. LG Authors: Jinheon Baek, Sun Jae Lee, Prakhar Gupta, Geunseob, Oh, Siddharth Dalmia, Prateek Kolhar Title: Revisiting In-Context Learning with Long Context Language Models Arxiv: http://arxiv.org/abs/2412.16926v1 Abstract: In-Context Learning (ICL) is a technique by which language models make predictions based on examples provided in their input context. Previ...

Outcome-Refining Process Supervision for Code Generation 25.12.2024

🤗 Upvotes: 11 | cs. CL, cs. AI, cs. LG, cs. SE Authors: Zhuohao Yu, Weizheng Gu, Yidong Wang, Zhengran Zeng, Jindong Wang, Wei Ye, Shikun Zhang Title: Outcome-Refining Process Supervision for Code Generation Arxiv: http://arxiv.org/abs/2412.15118v1 Abstract: Large Language Models have demonstrated remarkable capabilities in code generation, yet they often struggle with complex programming tasks t...

LearnLM: Improving Gemini for Learning 25.12.2024

🤗 Upvotes: 9 | cs. CY, cs. AI, cs. LG Authors: LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla, Aliya Rysbek, Andrea Huber, Brett Wiltshire, Brian Veprek, Daniel Gillick, Daniel Kasenberg, Derek Ahmed, Irina Jurenka, James Cohan, Jennifer She, Julia Wilkowski, Kaiz Alarakyia, Kevin McKee, Lisa Wang, Markus Kunesch, Mike Schaekermann, Miruna Pîslar, Nikhil Joshi, Parsa Mahmoudieh, Paul Jhu...

Parallelized Autoregressive Visual Generation 24.12.2024

🤗 Upvotes: 34 | cs. CV Authors: Yuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han, Haoyuan Guo, Zhenheng Yang, Difan Zou, Jiashi Feng, Xihui Liu Title: Parallelized Autoregressive Visual Generation Arxiv: http://arxiv.org/abs/2412.15119v1 Abstract: Autoregressive models have emerged as a powerful approach for visual generation but suffer from slow inference speed due to their sequential token-by-to...

Offline Reinforcement Learning for LLM Multi-Step Reasoning 24.12.2024

🤗 Upvotes: 19 | cs. LG, cs. AI, cs. CL Authors: Huaijie Wang, Shibo Hao, Hanze Dong, Shenao Zhang, Yilin Bao, Ziran Yang, Yi Wu Title: Offline Reinforcement Learning for LLM Multi-Step Reasoning Arxiv: http://arxiv.org/abs/2412.16145v1 Abstract: Improving the multi-step reasoning ability of large language models (LLMs) with offline reinforcement learning (RL) is essential for quickly adapting the...

SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation 24.12.2024

🤗 Upvotes: 17 | cs. CL Authors: Jialong Wu, Zhenglin Wang, Linhai Zhang, Yilong Lai, Yulan He, Deyu Zhou Title: SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation Arxiv: http://arxiv.org/abs/2412.13649v1 Abstract: Key-Value (KV) cache has become a bottleneck of LLMs for long-context generation. Despite the numerous efforts in this area, the optimization for the decoding phas...

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up 24.12.2024

🤗 Upvotes: 13 | cs. CV Authors: Songhua Liu, Zhenxiong Tan, Xinchao Wang Title: CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up Arxiv: http://arxiv.org/abs/2412.16112v1 Abstract: Diffusion Transformers (DiT) have become a leading architecture in image generation. However, the quadratic complexity of attention mechanisms, which are responsible for modeling token-wise rela...

Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis 24.12.2024

🤗 Upvotes: 12 | cs. CV, cs. LG, cs. SD, eess. AS Authors: Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, Yuki Mitsufuji Title: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis Arxiv: http://arxiv.org/abs/2412.15322v1 Abstract: We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a no...

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage 24.12.2024

🤗 Upvotes: 9 | cs. CV Authors: Saehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi, Sungroh Yoon Title: Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Arxiv: http://arxiv.org/abs/2412.15484v1 Abstract: Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. O...

Sequence Matters: Harnessing Video Models in 3D Super-Resolution 24.12.2024

🤗 Upvotes: 6 | cs. CV, 68U10, 68T10, I.4.5; I.2.10 Authors: Hyun-kyu Ko, Dongheok Park, Youngin Park, Byeonghyeon Lee, Juhee Han, Eunbyung Park Title: Sequence Matters: Harnessing Video Models in 3D Super-Resolution Arxiv: http://arxiv.org/abs/2412.11525v3 Abstract: 3D super-resolution aims to reconstruct high-fidelity 3D models from low-resolution (LR) multi-view images. Early studies primarily...

TRecViT: A Recurrent Video Transformer 24.12.2024

🤗 Upvotes: 5 | cs. CV, cs. LG Authors: Viorica Pătrăucean, Xu Owen He, Joseph Heyward, Chuhan Zhang, Mehdi S. M. Sajjadi, George-Cristian Muraru, Artem Zholus, Mahdi Karami, Ross Goroshin, Yutian Chen, Simon Osindero, João Carreira, Razvan Pascanu Title: TRecViT: A Recurrent Video Transformer Arxiv: http://arxiv.org/abs/2412.14294v1 Abstract: We propose a novel block for video modelling. It relie...

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design 24.12.2024

🤗 Upvotes: 4 | cs. LG Authors: Zhen Zheng, Xiaonan Song, Chuanjie Liu Title: MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Arxiv: http://arxiv.org/abs/2412.14590v1 Abstract: Quantization has become one of the most effective methodologies to compress LLMs into smaller size. However, the existing quantization solutions still show lim...

Multi-LLM Text Summarization 24.12.2024

🤗 Upvotes: 3 | cs. CL Authors: Jiangnan Fang, Cheng-Tse Liu, Jieun Kim, Yash Bhedaru, Ethan Liu, Nikhil Singh, Nedim Lipka, Puneet Mathur, Nesreen K. Ahmed, Franck Dernoncourt, Ryan A. Rossi, Hanieh Deilamsalehy Title: Multi-LLM Text Summarization Arxiv: http://arxiv.org/abs/2412.15487v1 Abstract: In this work, we propose a Multi-LLM summarization framework, and investigate two different multi-LL...

Qwen2.5 Technical Report 21.12.2024

🤗 Upvotes: 236 | cs. CL Authors: Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu X...

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval 21.12.2024

🤗 Upvotes: 44 | cs. CV, cs. CL Authors: Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian, Yongping Xiong Title: MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval Arxiv: http://arxiv.org/abs/2412.14475v1 Abstract: Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack...

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 21.12.2024

🤗 Upvotes: 23 | cs. CL, cs. AI Authors: Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li Title: LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks Arxiv: http://arxiv.org/abs/2412.15204v1 Abstract: This paper introduces LongBench v2, a benchmark designed to assess th...

How to Synthesize Text Data without Model Collapse? 21.12.2024

🤗 Upvotes: 19 | cs. CL, cs. AI, cs. LG Authors: Xuekai Zhu, Daixuan Cheng, Hengli Li, Kaiyan Zhang, Ermo Hua, Xingtai Lv, Ning Ding, Zhouhan Lin, Zilong Zheng, Bowen Zhou Title: How to Synthesize Text Data without Model Collapse? Arxiv: http://arxiv.org/abs/2412.14689v1 Abstract: Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline i...

Flowing from Words to Pixels: A Framework for Cross-Modality Evolution 21.12.2024

🤗 Upvotes: 17 | cs. CV Authors: Qihao Liu, Xi Yin, Alan Yuille, Andrew Brown, Mannat Singh Title: Flowing from Words to Pixels: A Framework for Cross-Modality Evolution Arxiv: http://arxiv.org/abs/2412.15213v1 Abstract: Diffusion models, and their generalization, flow matching, have had a remarkable impact on the field of media generation. Here, the conventional approach is to learn the complex m...

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion 21.12.2024

🤗 Upvotes: 13 | cs. CV Authors: Jixuan He, Wanhua Li, Ye Liu, Junsik Kim, Donglai Wei, Hanspeter Pfister Title: Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion Arxiv: http://arxiv.org/abs/2412.14462v1 Abstract: As a common image editing operation, image composition involves integrating foreground objects into background scenes. In this paper, we expand the application of the conce...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.