Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints 13.12.2024 21:09
🤗 Upvotes: 36 | cs. CV Authors: Jianhong Bai, Menghan Xia, Xintao Wang, Ziyang Yuan, Xiao Fu, Zuozhu Liu, Haoji Hu, Pengfei Wan, Di Zhang Title: SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints Arxiv: http://arxiv.org/abs/2412.07760v1 Abstract: Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and main...
LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations 13.12.2024 21:28
🤗 Upvotes: 28 | cs. CV Authors: Zejian Li, Chenye Meng, Yize Li, Ling Yang, Shengyuan Zhang, Jiarui Ma, Jiayi Li, Guang Yang, Changyuan Yang, Zhiyuan Yang, Jinxiong Chang, Lingyun Sun Title: LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations Arxiv: http://arxiv.org/abs/2412.08580v1 Abstract: Recent advances in text-to-image (T2I) generatio...
POINTS1.5: Building a Vision-Language Model towards Real World Applications 13.12.2024 24:23
🤗 Upvotes: 25 | cs. CV, cs. MM Authors: Yuan Liu, Le Tian, Xiao Zhou, Xinyu Gao, Kavio Yu, Yang Yu, Jie Zhou Title: POINTS1.5: Building a Vision-Language Model towards Real World Applications Arxiv: http://arxiv.org/abs/2412.08443v1 Abstract: Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character recognitio...
Learning Flow Fields in Attention for Controllable Person Image Generation 13.12.2024 21:04
🤗 Upvotes: 16 | cs. CV Authors: Zijian Zhou, Shikun Liu, Xiao Han, Haozhe Liu, Kam Woh Ng, Tian Xie, Yuren Cong, Hang Li, Mengmeng Xu, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang, Miaojing Shi, Sen He Title: Learning Flow Fields in Attention for Controllable Person Image Generation Arxiv: http://arxiv.org/abs/2412.08486v2 Abstract: Controllable person image generation aims to generate a person...
StyleMaster: Stylize Your Video with Artistic Generation and Translation 13.12.2024 23:20
🤗 Upvotes: 14 | cs. CV Authors: Zixuan Ye, Huijuan Huang, Xintao Wang, Pengfei Wan, Di Zhang, Wenhan Luo Title: StyleMaster: Stylize Your Video with Artistic Generation and Translation Arxiv: http://arxiv.org/abs/2412.07744v1 Abstract: Style control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to...
StreamChat: Chatting with Streaming Video 13.12.2024 19:44
🤗 Upvotes: 12 | cs. CV Authors: Jihao Liu, Zhiding Yu, Shiyi Lan, Shihao Wang, Rongyao Fang, Jan Kautz, Hongsheng Li, Jose M. Alvare Title: StreamChat: Chatting with Streaming Video Arxiv: http://arxiv.org/abs/2412.08646v1 Abstract: This paper presents StreamChat, a novel approach that enhances the interaction capabilities of Large Multimodal Models (LMMs) with streaming video content. In streami...
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark 13.12.2024 25:04
🤗 Upvotes: 11 | cs. CV Authors: Wufei Ma, Haoyu Chen, Guofeng Zhang, Celso M de Melo, Alan Yuille, Jieneng Chen Title: 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark Arxiv: http://arxiv.org/abs/2412.07825v1 Abstract: 3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to d...
Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction 13.12.2024 22:43
🤗 Upvotes: 11 | cs. CV, cs. GR Authors: Seungtae Nam, Xiangyu Sun, Gyeongjin Kang, Younggeun Lee, Seungjun Oh, Eunbyung Park Title: Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction Arxiv: http://arxiv.org/abs/2412.06234v2 Abstract: Generalized feed-forward Gaussian models have achieved significant progress in sparse-view 3D reconstruction b...
The BrowserGym Ecosystem for Web Agent Research 13.12.2024 25:16
🤗 Upvotes: 11 | cs. LG, cs. AI, cs. SE Authors: Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin, Massimo Caccia, Léo Boisvert, Megh Thakkar, Tom Marty, Rim Assouel, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, Dehan Kong, Frank F. Xu, Siva Reddy, Quentin Cappart, Graham Neubig, Ruslan Salakhutdinov, Nicolas Chapados, Alexandre Lacoste Title: The BrowserGym...
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation 12.12.2024 22:09
🤗 Upvotes: 31 | cs. CV Authors: Jianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng, Xiangtai Li, Yunhai Tong Title: DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Arxiv: http://arxiv.org/abs/2412.07589v1 Abstract: Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-to-image generation models....
Hidden in the Noise: Two-Stage Robust Watermarking for Images 12.12.2024 21:29
🤗 Upvotes: 20 | cs. CV, cs. AI, cs. LG Authors: Kasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde, Niv Cohen Title: Hidden in the Noise: Two-Stage Robust Watermarking for Images Arxiv: http://arxiv.org/abs/2412.04653v2 Abstract: As the quality of image generators continues to improve, deepfakes become a topic of considerable societal debate. Image watermarking allows responsible model ow...
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 12.12.2024 19:48
🤗 Upvotes: 19 | cs. CV Authors: Tong Wu, Yinghao Xu, Ryan Po, Mengchen Zhang, Guandao Yang, Jiaqi Wang, Ziwei Liu, Dahua Lin, Gordon Wetzstein Title: FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models Arxiv: http://arxiv.org/abs/2412.07674v1 Abstract: Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications....
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics 12.12.2024 23:56
🤗 Upvotes: 18 | cs. CV Authors: Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou, Soo Ye Kim, Qing Liu, Yijun Li, Jianming Zhang, Nanxuan Zhao, Yilin Wang, Hui Ding, Zhe Lin, Hengshuang Zhao Title: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics Arxiv: http://arxiv.org/abs/2412.07774v1 Abstract: We introduce UniReal, a unified framework designed to address various ima...
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation 12.12.2024 23:46
🤗 Upvotes: 17 | cs. CV Authors: Xiao Fu, Xian Liu, Xintao Wang, Sida Peng, Menghan Xia, Xiaoyu Shi, Ziyang Yuan, Pengfei Wan, Di Zhang, Dahua Lin Title: 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation Arxiv: http://arxiv.org/abs/2412.07759v1 Abstract: This paper aims to manipulate multi-entity 3D motions in video generation. Previous methods on controllable video...
Mobile Video Diffusion 12.12.2024 24:39
🤗 Upvotes: 16 | cs. CV, cs. AI Authors: Haitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas, Amir Ghodrati, Amirhossein Habibian Title: Mobile Video Diffusion Arxiv: http://arxiv.org/abs/2412.07583v1 Abstract: Video diffusion models have achieved impressive realism and controllability but are limited by high computational demands, restricting their use on mobile devices. This paper introduces the...
Granite Guardian 12.12.2024 21:00
🤗 Upvotes: 16 | cs. CL Authors: Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling, Martín Santillán Cooper, Kieran Fraser, Giulio Zizzo, Muhammad Zaid Hameed, Mark Purcell, Michael Desmond, Qian Pan, Inge Vejsbjerg, Elizabeth M. Daly, Michael Hind, Werner Geyer, Ambrish Rawat, Kush R. Varshney, Prasan...
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation 11.12.2024 18:54
🤗 Upvotes: 54 | cs. LG, cs. AI Authors: Egor Cherepanov, Nikita Kachaev, Artem Zholus, Alexey K. Kovalev, Aleksandr I. Panov Title: Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation Arxiv: http://arxiv.org/abs/2412.06531v1 Abstract: The incorporation of memory into agents is essential for numerous tasks within the domain of Reinforcement Learning (RL)...
ProcessBench: Identifying Process Errors in Mathematical Reasoning 11.12.2024 21:22
🤗 Upvotes: 38 | cs. AI, cs. CL, cs. LG Authors: Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin Title: ProcessBench: Identifying Process Errors in Mathematical Reasoning Arxiv: http://arxiv.org/abs/2412.06559v2 Abstract: As language models regularly make mistakes when solving math problems, automated identification of errors in th...
Training Large Language Models to Reason in a Continuous Latent Space 11.12.2024 22:02
🤗 Upvotes: 25 | cs. CL Authors: Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian Title: Training Large Language Models to Reason in a Continuous Latent Space Arxiv: http://arxiv.org/abs/2412.06769v1 Abstract: Large language models (LLMs) are restricted to reason in the "language space", where they typically express the reasoning process with a chain-of-t...
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation 11.12.2024 23:51
🤗 Upvotes: 10 | cs. CV Authors: Yuying Ge, Yizhuo Li, Yixiao Ge, Ying Shan Title: Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Arxiv: http://arxiv.org/abs/2412.04432v1 Abstract: In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to expl...
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation 11.12.2024 22:19
🤗 Upvotes: 9 | cs. CV, cs. LG Authors: Nicolas Dufour, David Picard, Vicky Kalogeiton, Loic Landrieu Title: Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation Arxiv: http://arxiv.org/abs/2412.06781v1 Abstract: Global visual geolocation predicts where an image was captured on Earth. Since images vary in how precisely they can be localized, this task inherently inv...
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models 11.12.2024 22:31
🤗 Upvotes: 8 | cs. CV, cs. CL, cs. LG Authors: Xiao Xu, Tianhao Niu, Yuxi Xie, Libo Qin, Wanxiang Che, Min-Yen Kan Title: Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models Arxiv: http://arxiv.org/abs/2412.05939v1 Abstract: Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., imag...
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale 11.12.2024 19:54
🤗 Upvotes: 7 | cs. CV Authors: Baorui Ma, Huachen Gao, Haoge Deng, Zhengxiong Luo, Tiejun Huang, Lulu Tang, Xinlong Wang Title: You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale Arxiv: http://arxiv.org/abs/2412.06699v1 Abstract: Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their perfor...
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations 11.12.2024 20:24
🤗 Upvotes: 7 | cs. CV, cs. AI, cs. IR Authors: Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, Conghui He Title: OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations Arxiv: http://arxiv.org/abs...
Robust Multi-bit Text Watermark with LLM-based Paraphrasers 11.12.2024 17:59
🤗 Upvotes: 5 | cs. AI Authors: Xiaojun Xu, Jinghan Jia, Yuanshun Yao, Yang Liu, Hang Li Title: Robust Multi-bit Text Watermark with LLM-based Paraphrasers Arxiv: http://arxiv.org/abs/2412.03123v1 Abstract: We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasi...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.