Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

RL + Transformer = A General-Purpose Problem Solver 28.01.2025

🤗 Upvotes: 7 | cs. LG, cs. AI Authors: Micah Rentschler, Jesse Roberts Title: RL + Transformer = A General-Purpose Problem Solver Arxiv: http://arxiv.org/abs/2501.14176v1 Abstract: What if artificial intelligence could not only solve problems for which it was trained but also learn to teach itself to solve new problems (i.e., meta-learn)? In this study, we demonstrate that a pre-trained transform...

Relightable Full-Body Gaussian Codec Avatars 28.01.2025

🤗 Upvotes: 5 | cs. CV, cs. GR Authors: Shaofei Wang, Tomas Simon, Igor Santesteban, Timur Bagautdinov, Junxuan Li, Vasu Agrawal, Fabian Prada, Shoou-I Yu, Pace Nalbone, Matt Gramlich, Roman Lubachersky, Chenglei Wu, Javier Romero, Jason Saragih, Michael Zollhoefer, Andreas Geiger, Siyu Tang, Shunsuke Saito Title: Relightable Full-Body Gaussian Codec Avatars Arxiv: http://arxiv.org/abs/2501.14726v...

Question Answering on Patient Medical Records with Private Fine-Tuned LLMs 28.01.2025

🤗 Upvotes: 4 | cs. CL, cs. AI Authors: Sara Kothari, Ayush Gupta Title: Question Answering on Patient Medical Records with Private Fine-Tuned LLMs Arxiv: http://arxiv.org/abs/2501.13687v1 Abstract: Healthcare systems continuously generate vast amounts of electronic health records (EHRs), commonly stored in the Fast Healthcare Interoperability Resources (FHIR) standard. Despite the wealth of infor...

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing 28.01.2025

🤗 Upvotes: 3 | cs. CV Authors: Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, Salman Khan Title: GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing Arxiv: http://arxiv.org/abs/2501.13925v1 Abstract: Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the...

AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation 28.01.2025

🤗 Upvotes: 2 | cs. CV Authors: Yuning Cui, Syed Waqas Zamir, Salman Khan, Alois Knoll, Mubarak Shah, Fahad Shahbaz Khan Title: AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation Arxiv: http://arxiv.org/abs/2403.14614v1 Abstract: In the image acquisition process, various forms of degradation, including noise, haze, and rain, are frequently introduced. These degradatio...

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning 28.01.2025

🤗 Upvotes: 2 | cs. CV Authors: Yang You, Yixin Li, Congyue Deng, Yue Wang, Leonidas Guibas Title: Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning Arxiv: http://arxiv.org/abs/2411.19458v1 Abstract: Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their...

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding 25.01.2025

🤗 Upvotes: 46 | cs. LG, cs. AI, cs. MA, I.2.11 Authors: Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev Title: SRMT: Shared Memory for Multi-agent Lifelong Pathfinding Arxiv: http://arxiv.org/abs/2501.13200v1 Abstract: Multi-agent reinforcement learning (MARL) demonstrates significant progress in solving cooperative and competitive multi-agent problems in various environments. One of the principal c...

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models 25.01.2025

🤗 Upvotes: 33 | cs. CL Authors: Zhenghao Lin, Zihao Tang, Xiao Liu, Yeyun Gong, Yi Cheng, Qi Chen, Hang Li, Ying Xin, Ziyue Yang, Kailai Yang, Yu Yan, Xiao Liang, Shuai Lu, Yiming Huang, Zheheng Luo, Lei Qu, Xuan Feng, Yaoxiang Wang, Yuqing Xia, Feiyang Chen, Yuting Jiang, Yasen Hu, Hao Ni, Binyang Li, Guoshuai Zhao, Jui-Hao Chiang, Zhongxin Guo, Chen Lin, Kun Kuang, Wenjie Li, Yelong Shen, Jian...

Improving Video Generation with Human Feedback 25.01.2025

🤗 Upvotes: 30 | cs. CV, cs. AI, cs. GR, cs. LG Authors: Jie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Wenyu Qin, Menghan Xia, Xintao Wang, Xiaohong Liu, Fei Yang, Pengfei Wan, Di Zhang, Kun Gai, Yujiu Yang, Wanli Ouyang Title: Improving Video Generation with Human Feedback Arxiv: http://arxiv.org/abs/2501.13918v1 Abstract: Video generation has a...

Temporal Preference Optimization for Long-Form Video Understanding 25.01.2025

🤗 Upvotes: 15 | cs. CV, cs. AI, cs. CL, cs. LG, cs. RO Authors: Rui Li, Xiaohan Wang, Yuhui Zhang, Zeyu Wang, Serena Yeung-Levy Title: Temporal Preference Optimization for Long-Form Video Understanding Arxiv: http://arxiv.org/abs/2501.13919v1 Abstract: Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains...

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step 25.01.2025

🤗 Upvotes: 14 | cs. CV, cs. AI, cs. CL Authors: Ziyu Guo, Renrui Zhang, Chengzhuo Tong, Zhizheng Zhao, Peng Gao, Hongsheng Li, Pheng-Ann Heng Title: Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Arxiv: http://arxiv.org/abs/2501.13926v1 Abstract: Chain-of-Thought (CoT) reasoning has been extensively explored in large models to tackle complex understandin...

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 25.01.2025

🤗 Upvotes: 10 | cs. CV, cs. CL Authors: Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, Ziwei Liu Title: Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Arxiv: http://arxiv.org/abs/2501.13826v1 Abstract: Humans acquire knowledge through three cognitive stages: perceiving information, comprehending knowledge, and adapting knowledg...

DiffuEraser: A Diffusion Model for Video Inpainting 25.01.2025

🤗 Upvotes: 8 | cs. CV Authors: Xiaowen Li, Haolan Xue, Peiran Ren, Liefeng Bo Title: DiffuEraser: A Diffusion Model for Video Inpainting Arxiv: http://arxiv.org/abs/2501.10018v1 Abstract: Recent video inpainting algorithms integrate flow-based pixel propagation with transformer-based generation to leverage optical flow for restoring textures and objects using information from neighboring frames,...

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models 25.01.2025

🤗 Upvotes: 8 | cs. CV, cs. CL, cs. LG Authors: Jiayi Lei, Renrui Zhang, Xiangfei Hu, Weifeng Lin, Zhen Li, Wenjian Sun, Ruoyi Du, Le Zhuo, Zhongyu Li, Xinyue Li, Shitian Zhao, Ziyu Guo, Yiting Lu, Peng Gao, Hongsheng Li Title: IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Arxiv: http://arxiv.org/abs/2501.13920v1 Abstract: With the rapid development o...

Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback 25.01.2025

🤗 Upvotes: 7 | cs. LG, cs. AI Authors: Yen-Ting Lin, Di Jin, Tengyu Xu, Tianhao Wu, Sainbayar Sukhbaatar, Chen Zhu, Yun He, Yun-Nung Chen, Jason Weston, Yuandong Tian, Arash Rahnama, Sinong Wang, Hao Ma, Han Fang Title: Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback Arxiv: http://arxiv.org/abs/2501.10799v1 Abstract: Large language models (LLMs) have recently demonstr...

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt 25.01.2025

🤗 Upvotes: 5 | cs. CV, cs. AI, cs. LG Authors: Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, Ming-Ming Cheng Title: One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt Arxiv: http://arxiv.org/abs/2501.13554v1 Abstract: Text-to-image generation models can create high-quality images from input prompt...

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning 24.01.2025

🤗 Upvotes: 109 | cs. CL, cs. AI, cs. LG Authors: DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan,...

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 24.01.2025

🤗 Upvotes: 44 | cs. CV Authors: Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, Peng Jin, Wenqi Zhang, Fan Wang, Lidong Bing, Deli Zhao Title: VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Arxiv: http://arxiv.org/abs/2501.13106v2 Abstract: In this paper, we propose VideoLLaMA3, a...

FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces 24.01.2025

🤗 Upvotes: 43 | cs. CL, cs. GR, cs. MA Authors: Zhenran Xu, Longyue Wang, Jifang Wang, Zhouyi Li, Senbao Shi, Xue Yang, Yiyu Wang, Baotian Hu, Jun Yu, Min Zhang Title: FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces Arxiv: http://arxiv.org/abs/2501.12909v1 Abstract: Virtual film production requires intricate decision-making processes, including scriptwriting...

Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback 24.01.2025

🤗 Upvotes: 42 | cs. CL Authors: Yafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li, Yu Cheng Title: Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback Arxiv: http://arxiv.org/abs/2501.12895v1 Abstract: Large language models (LLMs) demonstrate impressive performance but lack the flexibility to adapt to human preferences quickly without retraining. In this work, we introdu...

Kimi k1.5: Scaling Reinforcement Learning with LLMs 24.01.2025

🤗 Upvotes: 39 | cs. AI, cs. LG Authors: Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, Chuning Tang, Congcong Wang, Dehao Zhang, Enming Yuan, Enzhe Lu, Fengxiang Tang, Flood Sung, Guangda Wei, Guokun Lai, Haiqing Guo, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haotian Yao, Haotian Zhao, Haoyu Lu, Haoze Li, Haozhe...

Autonomy-of-Experts Models 24.01.2025

🤗 Upvotes: 31 | cs. CL, cs. AI, cs. LG Authors: Ang Lv, Ruobing Xie, Yining Qian, Songhao Wu, Xingwu Sun, Zhanhui Kang, Di Wang, Rui Yan Title: Autonomy-of-Experts Models Arxiv: http://arxiv.org/abs/2501.13074v1 Abstract: Mixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argu...

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning 24.01.2025

🤗 Upvotes: 13 | cs. CL Authors: Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, Dacheng Tao Title: O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning Arxiv: http://arxiv.org/abs/2501.12570v1 Abstract: Recently, long-thought reasoning LLMs, such as OpenAI's O1, adopt extended reasoning processes similar to how humans ponder over com...

Pairwise RM: Perform Best-of-N Sampling with Knockout Tournament 24.01.2025

🤗 Upvotes: 13 | cs. CL Authors: Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, Juanzi Li Title: Pairwise RM: Perform Best-of-N Sampling with Knockout Tournament Arxiv: http://arxiv.org/abs/2501.13007v1 Abstract: Best-of-N (BoN) sampling, a common strategy for test-time scaling of Large Language Models (LLMs), relies on reward models to select the best candidate solution from multiple generat...

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems 24.01.2025

🤗 Upvotes: 7 | cs. CL, cs. AI, cs. LG Authors: Elad Levi, Ilan Kadar Title: IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Arxiv: http://arxiv.org/abs/2501.11067v1 Abstract: Large Language Models (LLMs) are transforming artificial intelligence, evolving into task-oriented systems capable of autonomous planning and execution. One of the primary applications of LLMs i...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.