Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Cosmos World Foundation Model Platform for Physical AI 09.01.2025

🤗 Upvotes: 31 | cs. CV, cs. AI, cs. LG, cs. RO Authors: NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huffman, Poo...

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token 09.01.2025

🤗 Upvotes: 22 | cs. CV, cs. AI, cs. CL Authors: Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng Title: LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token Arxiv: http://arxiv.org/abs/2501.03895v1 Abstract: The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual i...

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 09.01.2025

🤗 Upvotes: 18 | cs. CV Authors: Haobo Yuan, Xiangtai Li, Tao Zhang, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang Title: Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Arxiv: http://arxiv.org/abs/2501.04001v1 Abstract: This work presents Sa2VA, the first unified model for dense grounded understanding of both images an...

Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control 09.01.2025

🤗 Upvotes: 13 | cs. CV, cs. AI, cs. GR Authors: Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, Yuan Liu Title: Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control Arxiv: http://arxiv.org/abs/2501.03847v1 Abstract: Diffusion models have demonstrated impressive performance in generating hig...

OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis 09.01.2025

🤗 Upvotes: 10 | cs. CL, cs. CV Authors: Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu, Xiong Liu, Min Yang, Yongbin Li, Longze Chen, Jiaming Li, Lei Zhang, Yangyi Chen, Hamid Alinejad-Rokny, Fei Huang Title: OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis Arxiv: http://arxiv.org/abs/2501.04561v1 Abstract: Re...

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides 09.01.2025

🤗 Upvotes: 10 | cs. AI, cs. CL Authors: Hao Zheng, Xinyan Guan, Hao Kong, Jia Zheng, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, Le Sun Title: PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides Arxiv: http://arxiv.org/abs/2501.03936v1 Abstract: Automatically generating presentations from documents is a challenging task that requires balancing content quality, visual design, a...

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model 09.01.2025

🤗 Upvotes: 6 | cs. CL, cs. AI Authors: Yueqin Yin, Shentao Yang, Yujia Xie, Ziyi Yang, Yuting Sun, Hany Awadalla, Weizhu Chen, Mingyuan Zhou Title: Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Arxiv: http://arxiv.org/abs/2501.02790v1 Abstract: Reinforcement learning from human feedback (RLHF) has been widely adopted to align language models (LMs) with human prefe...

MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting 09.01.2025

🤗 Upvotes: 6 | cs. CV Authors: Sangwoon Kwak, Joonsoo Kim, Jun Young Jeong, Won-Sik Cheong, Jihyong Oh, Munchurl Kim Title: MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting Arxiv: http://arxiv.org/abs/2501.03714v1 Abstract: 3D Gaussian Splatting (3DGS) has made significant strides in scene representation and neural rendering...

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution 08.01.2025

🤗 Upvotes: 38 | cs. CV Authors: Rui Xie, Yinhong Liu, Penghao Zhou, Chen Zhao, Jun Zhou, Kai Zhang, Zhenyu Zhang, Jian Yang, Zhenheng Yang, Ying Tai Title: STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution Arxiv: http://arxiv.org/abs/2501.02976v1 Abstract: Image diffusion models have been adapted for real-world video super-resolution to tackle ove...

Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 08.01.2025

🤗 Upvotes: 23 | cs. CV Authors: Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang Title: Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction Arxiv: http://arxiv.org/abs/2501.03218v1 Abstract: Active Real-time interaction with video LLMs introduces a new paradigm for human-computer intera...

BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 08.01.2025

🤗 Upvotes: 22 | cs. CL, cs. AI, cs. LG Authors: Beichen Zhang, Yuhong Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Haodong Duan, Yuhang Cao, Dahua Lin, Jiaqi Wang Title: BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning Arxiv: http://arxiv.org/abs/2501.03226v1 Abstract: Cutting-edge large language models (LLMs) demonstrate promising performance i...

Personalized Graph-Based Retrieval for Large Language Models 08.01.2025

🤗 Upvotes: 19 | cs. CL Authors: Steven Au, Cameron J. Dimacali, Ojasmitha Pedirappagari, Namyong Park, Franck Dernoncourt, Yu Wang, Nikos Kanakaris, Hanieh Deilamsalehy, Ryan A. Rossi, Nesreen K. Ahmed Title: Personalized Graph-Based Retrieval for Large Language Models Arxiv: http://arxiv.org/abs/2501.02157v1 Abstract: As large language models (LLMs) evolve, their ability to deliver personalized...

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring 08.01.2025

🤗 Upvotes: 13 | q-bio. GN, cs. AI, cs. CL, cs. LG Authors: Ollie Liu, Sami Jaghouar, Johannes Hagemann, Shangshang Wang, Jason Wiemels, Jeff Kaufman, Willie Neiswanger Title: METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring Arxiv: http://arxiv.org/abs/2501.02045v1 Abstract: We pretrain METAGENE-1, a 7-billion-parameter autoregressive transformer model, which we refer to as a metag...

GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking 08.01.2025

🤗 Upvotes: 12 | cs. CV Authors: Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yijin Li, Fu-Yun Wang, Hongsheng Li Title: GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking Arxiv: http://arxiv.org/abs/2501.02690v1 Abstract: 4D video control is essential in video generation as it enables the use of sophisticated lens techniques, such as multi-cam...

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 08.01.2025

🤗 Upvotes: 12 | cs. CV, cs. AI, cs. LG Authors: Guy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin, Yaniv Taigman, Yossi Adi, Sagie Benaim, Adam Polyak Title: Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Arxiv: http://arxiv.org/abs/2501.03059v1 Abstract: We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into rea...

TransPixar: Advancing Text-to-Video Generation with Transparency 08.01.2025

🤗 Upvotes: 9 | cs. CV Authors: Luozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang, Zhifei Zhang, He Zhang, Zhe Lin, Yingcong Chen Title: TransPixar: Advancing Text-to-Video Generation with Transparency Arxiv: http://arxiv.org/abs/2501.03006v1 Abstract: Text-to-video generative models have made significant strides, enabling diverse applications in entertainment, advertising, and education. Howeve...

AutoPresent: Designing Structured Visuals from Scratch 08.01.2025

🤗 Upvotes: 7 | cs. CV, cs. CL Authors: Jiaxin Ge, Zora Zhiruo Wang, Xuhui Zhou, Yi-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, Trevor Darrell Title: AutoPresent: Designing Structured Visuals from Scratch Arxiv: http://arxiv.org/abs/2501.00912v1 Abstract: Designing structured visuals such as presentation slides is essential for communicative need...

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation 07.01.2025

🤗 Upvotes: 41 | cs. RO, cs. CV, cs. LG Authors: Siyuan Huang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Zhengkai Jiang, Yue Hu, Peng Gao, Hongsheng Li, Maoqing Yao, Guanghui Ren Title: EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation Arxiv: http://arxiv.org/abs/2501.01895v1 Abstract: We introduce EnerVerse, a comprehensive framework for embodied future space generation spe...

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction 07.01.2025

🤗 Upvotes: 23 | cs. CV, cs. SD, eess. AS Authors: Chaoyou Fu, Haojia Lin, Xiong Wang, Yi-Fan Zhang, Yunhang Shen, Xiaoyu Liu, Yangze Li, Zuwei Long, Heting Gao, Ke Li, Xiawu Zheng, Rongrong Ji, Xing Sun, Caifeng Shan, Ran He Title: VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Arxiv: http://arxiv.org/abs/2501.01957v1 Abstract: Recent Multimodal Large Language Models (MLLM...

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation 07.01.2025

🤗 Upvotes: 12 | cs. CV Authors: Jiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang, Jiajun Xu, Yuan Wang, Wenbo Duan, Shen Yang, Qunlin Jin, Shurun Li, Jiayan Teng, Zhuoyi Yang, Wendi Zheng, Xiao Liu, Ming Ding, Xiaohan Zhang, Xiaotao Gu, Shiyu Huang, Minlie Huang, Jie Tang, Yuxiao Dong Title: VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation Arx...

Virgo: A Preliminary Exploration on Reproducing o1-like MLLM 07.01.2025

🤗 Upvotes: 12 | cs. CV, cs. AI Authors: Yifan Du, Zikang Liu, Yifan Li, Wayne Xin Zhao, Yuqi Huo, Bingning Wang, Weipeng Chen, Zheng Liu, Zhongyuan Wang, Ji-Rong Wen Title: Virgo: A Preliminary Exploration on Reproducing o1-like MLLM Arxiv: http://arxiv.org/abs/2501.01904v1 Abstract: Recently, slow-thinking reasoning systems, built upon large language models (LLMs), have garnered widespread atten...

SDPO: Segment-Level Direct Preference Optimization for Social Agents 07.01.2025

🤗 Upvotes: 10 | cs. AI, cs. CL Authors: Aobo Kong, Wentao Ma, Shiwan Zhao, Yongbin Li, Yuchuan Wu, Ke Wang, Xiaoqian Liu, Qicheng Li, Yong Qin, Fei Huang Title: SDPO: Segment-Level Direct Preference Optimization for Social Agents Arxiv: http://arxiv.org/abs/2501.01821v1 Abstract: Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling c...

Graph Generative Pre-trained Transformer 07.01.2025

🤗 Upvotes: 9 | cs. LG, cs. AI Authors: Xiaohui Chen, Yinkai Wang, Jiaxing He, Yuanqi Du, Soha Hassoun, Xiaolin Xu, Li-Ping Liu Title: Graph Generative Pre-trained Transformer Arxiv: http://arxiv.org/abs/2501.01073v1 Abstract: Graph generation is a critical task in numerous domains, including molecular design and social network analysis, due to its ability to model complex relationships and struct...

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models 07.01.2025

🤗 Upvotes: 7 | cs. CL, cs. IR Authors: Hieu Man, Nghia Trung Ngo, Viet Dac Lai, Ryan A. Rossi, Franck Dernoncourt, Thien Huu Nguyen Title: LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models Arxiv: http://arxiv.org/abs/2501.00874v1 Abstract: Recent advancements in large language models (LLMs) based embedding models have established new sta...

BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery 07.01.2025

🤗 Upvotes: 5 | cs. LG, cs. AI Authors: Kanishk Gandhi, Michael Y. Li, Lyle Goodyear, Louise Li, Aditi Bhaskar, Mohammed Zaman, Noah D. Goodman Title: BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery Arxiv: http://arxiv.org/abs/2501.01540v1 Abstract: Understanding the world and explaining it with scientific theories is a central aspiration of artificial intelli...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.