Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters 16.04.2025

🤗 Upvotes: 95 | cs. DC, cs. AI, 68T50, I.2.7; I.2.11 Authors: Zonghang Li, Tao Li, Wenjiao Feng, Mohsen Guizani, Hongfang Yu Title: PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters Arxiv: http://arxiv.org/abs/2504.08791v1 Abstract: Emergency of DeepSeek R1 and QwQ 32B have broken through performance barriers for running frontier large language models (LLMs) on...

VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning 16.04.2025

🤗 Upvotes: 36 | cs. LG, cs. AI Authors: Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, Wenhu Chen Title: VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning Arxiv: http://arxiv.org/abs/2504.08837v1 Abstract: Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through...

FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding 16.04.2025

🤗 Upvotes: 35 | cs. CV Authors: Zheng Liu, Mengjie Liu, Jingzhou Chen, Jingwei Xu, Bin Cui, Conghui He, Wentao Zhang Title: FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Arxiv: http://arxiv.org/abs/2504.09925v1 Abstract: We introduce FUSION, a family of multimodal large language models (MLLMs) with a fully vision-language alignment and integration...

Iterative Self-Training for Code Generation via Reinforced Re-Ranking 16.04.2025

🤗 Upvotes: 29 | cs. CL, cs. IR, cs. SE Authors: Nikita Sorokin, Ivan Sedykh, Valentin Malykh Title: Iterative Self-Training for Code Generation via Reinforced Re-Ranking Arxiv: http://arxiv.org/abs/2504.09643v1 Abstract: Generating high-quality code that solves complex programming tasks is challenging, especially with current decoder-based models that produce highly stochastic outputs. In code ge...

Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model 15.04.2025

🤗 Upvotes: 83 | cs. CV, cs. AI Authors: Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, Feng Cheng, Feilong Zuo Xuejiao Zeng, Ziyan Yang, Fangyuan Kong, Zhiwu Qing, Fei Xiao, Meng Wei, Tuyen Hoang, Siyu Zhang, Peihao Zhu, Qi Zhao, Jiangqiao Yan, Liangke Gui, Sheng Bi, Jiashi Li, Yuxi Ren, Rui Wang, Huixia Li, Xuefeng Xiao, Shu Li...

GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 15.04.2025

🤗 Upvotes: 32 | cs. CV Authors: Tianwei Xiong, Jun Hao Liew, Zilong Huang, Jiashi Feng, Xihui Liu Title: GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation Arxiv: http://arxiv.org/abs/2504.08736v1 Abstract: In autoregressive (AR) image generation, visual tokenizers compress images into compact discrete latent tokens, enabling efficient training of downs...

MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft 15.04.2025

🤗 Upvotes: 25 | cs. CV, cs. AI Authors: Junliang Guo, Yang Ye, Tianyu He, Haoyu Wu, Yushu Jiang, Tim Pearce, Jiang Bian Title: MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft Arxiv: http://arxiv.org/abs/2504.08388v1 Abstract: World modeling is a crucial task for enabling intelligent agents to effectively interact with humans and operate in dynamic environments. In this...

Kimi-VL Technical Report 12.04.2025

🤗 Upvotes: 71 | cs. CV Authors: Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, Dehao Zhang, Dikang Du, Dongliang Wang, Enming Yuan, Enzhe Lu, Fang Li, Flood Sung, Guangda Wei, Guokun Lai, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haoning Wu, Haotian Yao, Haoyu Lu, Heng Wang, Hongcheng Gao, Huabin Zheng, J...

C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing 12.04.2025

🤗 Upvotes: 37 | cs. LG Authors: Zhongyang Li, Ziyue Li, Tianyi Zhou Title: C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing Arxiv: http://arxiv.org/abs/2504.07964v1 Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) suffer from severely sub-optimal expert pathways-our study reveals that naive expert selection learned from pretrainin...

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 12.04.2025

🤗 Upvotes: 34 | cs. CV, cs. AI, cs. CL Authors: Yukun Qi, Yiming Zhao, Yu Zeng, Xikun Bao, Wenxuan Huang, Lin Chen, Zehui Chen, Jie Zhao, Zhongang Qi, Feng Zhao Title: VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning Arxiv: http://arxiv.org/abs/2504.07956v1 Abstract: The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities...

DeepSeek-R1 Thoughtology: Let's <think> about LLM Reasoning 12.04.2025

🤗 Upvotes: 33 | cs. CL Authors: Sara Vera Marjanović, Arkil Patel, Vaibhav Adlakha, Milad Aghajohari, Parishad BehnamGhader, Mehar Bhatia, Aditi Khandelwal, Austin Kraft, Benno Krojer, Xing Han Lù, Nicholas Meade, Dongchan Shin, Amirhossein Kazemnejad, Gaurav Kamath, Marius Mosbach, Karolina Stańczak, Siva Reddy Title: DeepSeek-R1 Thoughtology: Let's about LLM Reasoning Arxiv: http://arxiv.org/ab...

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning 12.04.2025

🤗 Upvotes: 33 | cs. CV Authors: Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Zhen Li, Peng Gao, Zhanyu Ma, Ming-Ming Cheng Title: VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning Arxiv: http://arxiv.org/abs/2504.07960v1 Abstract: Recent progress in diffusion models significantly advances various image generation tasks. However, the current mainstream approach re...

MM-IFEngine: Towards Multimodal Instruction Following 12.04.2025

🤗 Upvotes: 26 | cs. CV Authors: Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, Jiaqi Wang Title: MM-IFEngine: Towards Multimodal Instruction Following Arxiv: http://arxiv.org/abs/2504.07957v1 Abstract: The Instruction Following (IF) ability measures how well Multi-modal Large Language Models (MLLMs) understand exactly what users...

HoloPart: Generative 3D Part Amodal Segmentation 12.04.2025

🤗 Upvotes: 23 | cs. CV Authors: Yunhan Yang, Yuan-Chen Guo, Yukun Huang, Zi-Xin Zou, Zhipeng Yu, Yangguang Li, Yan-Pei Cao, Xihui Liu Title: HoloPart: Generative 3D Part Amodal Segmentation Arxiv: http://arxiv.org/abs/2504.07943v1 Abstract: 3D part amodal segmentation--decomposing a 3D shape into complete, semantically meaningful parts, even when occluded--is a challenging but crucial task for 3D...

DDT: Decoupled Diffusion Transformer 11.04.2025

🤗 Upvotes: 51 | cs. CV, cs. AI Authors: Shuai Wang, Zhi Tian, Weilin Huang, Limin Wang Title: DDT: Decoupled Diffusion Transformer Arxiv: http://arxiv.org/abs/2504.05741v2 Abstract: Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs...

OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens 11.04.2025

🤗 Upvotes: 43 | cs. CL Authors: Jiacheng Liu, Taylor Blanton, Yanai Elazar, Sewon Min, YenSung Chen, Arnavi Chheda-Kothary, Huy Tran, Byron Bischoff, Eric Marsh, Michael Schmitz, Cassidy Trier, Aaron Sarnat, Jenna James, Jon Borchardt, Bailey Kuehl, Evie Cheng, Karen Farley, Sruthi Sreeram, Taira Anderson, David Albright, Carissa Schoenick, Luca Soldaini, Dirk Groeneveld, Rock Yuren Pang, Pang We...

A Unified Agentic Framework for Evaluating Conditional Image Generation 11.04.2025

🤗 Upvotes: 25 | cs. CV, cs. CL Authors: Jifang Wang, Xue Yang, Longyue Wang, Zhenran Xu, Yiyu Wang, Yaowei Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang Title: A Unified Agentic Framework for Evaluating Conditional Image Generation Arxiv: http://arxiv.org/abs/2504.07046v1 Abstract: Conditional image generation has gained significant attention for its ability to personalize content. However...

Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? 11.04.2025

🤗 Upvotes: 24 | cs. AI, cs. CL, cs. LG Authors: Chenrui Fan, Ming Li, Lichao Sun, Tianyi Zhou Title: Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? Arxiv: http://arxiv.org/abs/2504.06514v1 Abstract: We find that the response length of reasoning LLMs, whether trained by reinforcement learning or supervised learning, drastically increases for ill-pose...

OmniSVG: A Unified Scalable Vector Graphics Generation Model 10.04.2025

🤗 Upvotes: 91 | cs. CV Authors: Yiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng, Jiaxu Zhang, Liao Wang, Gang Yu, Xingjun Ma, Yu-Gang Jiang Title: OmniSVG: A Unified Scalable Vector Graphics Generation Model Arxiv: http://arxiv.org/abs/2504.06263v1 Abstract: Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and...

Hogwild! Inference: Parallel LLM Generation via Concurrent Attention 10.04.2025

🤗 Upvotes: 73 | cs. LG, cs. CL Authors: Gleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev, Vage Egiazarian, Anton Sinitsin, Denis Kuznedelev, Dan Alistarh Title: Hogwild! Inference: Parallel LLM Generation via Concurrent Attention Arxiv: http://arxiv.org/abs/2504.06261v2 Abstract: Large Language Models (LLMs) have demonstrated the ability to tackle increasingly complex tasks through ad...

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought 10.04.2025

🤗 Upvotes: 62 | cs. CV, cs. CL Authors: Yi Peng, Chris, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, Rongxian Zhuang, Xuchen Song, Yang Liu, Yahui Zhou Title: Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought Arxiv: http://arxiv.org/abs/2504.05599v1 Abstract: We introduce Skywork R1V, a multimodal reasoning model exte...

An Empirical Study of GPT-4o Image Generation Capabilities 10.04.2025

🤗 Upvotes: 50 | cs. CV Authors: Sixiang Chen, Jinbin Bai, Zhuoran Zhao, Tian Ye, Qingyu Shi, Donghao Zhou, Wenhao Chai, Xin Lin, Jianzong Wu, Chao Tang, Shilin Xu, Tao Zhang, Haobo Yuan, Yikang Zhou, Wei Chow, Linfeng Li, Xiangtai Li, Lei Zhu, Lu Qi Title: An Empirical Study of GPT-4o Image Generation Capabilities Arxiv: http://arxiv.org/abs/2504.05979v1 Abstract: The landscape of image generatio...

COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values 10.04.2025

🤗 Upvotes: 36 | cs. CL Authors: M-A-P Team, Siwei Wu, Jincheng Ren, Xinrun Du, Shuyue Guo, Xingwei Qu, Yiming Liang, Jie Liu, Yunwen Li, Tianyu Zheng, Boyu Feng, Huaqing Yuan, Zenith Wang, Jiaheng Liu, Wenhao Huang, Chenglin Cai, Haoran Que, Jian Yang, Yuelin Bai, Zekun Moore Wang, Zhouliang Yu, Qunshu Lin, Ding Pan, Yuchen Jiang, Tiannan Wang, Wangchunshu Zhou, Shenzhi Wang, Xingyuan Bu, Minghao...

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation 10.04.2025

🤗 Upvotes: 27 | cs. CV, cs. LG Authors: Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng, Fei Ding, Qian He Title: Less-to-More Generalization: Unlocking More Controllability by In-Context Generation Arxiv: http://arxiv.org/abs/2504.02160v1 Abstract: Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data s...

SmolVLM: Redefining small and efficient multimodal models 09.04.2025

🤗 Upvotes: 96 | cs. AI, cs. CV Authors: Andrés Marafioti, Orr Zohar, Miquel Farré, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tunstall, Leandro von Werra, Thomas Wolf Title: SmolVLM: Redefining small and efficient multimodal models Arxiv: http://arxiv.org/abs/2504.0529...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.