Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

SketchAgent: Language-Driven Sequential Sketch Generation 28.11.2024

🤗 Paper Upvotes: 13 | cs. CV Authors: Yael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao, Judith E Fan, Antonio Torralba Title: SketchAgent: Language-Driven Sequential Sketch Generation Arxiv: http://arxiv.org/abs/2411.17673v1 Abstract: Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While art...

TEXGen: a Generative Diffusion Model for Mesh Textures 28.11.2024

🤗 Paper Upvotes: 12 | cs. CV, cs. AI, cs. GR Authors: Xin Yu, Ze Yuan, Yuan-Chen Guo, Ying-Tian Liu, JianHui Liu, Yangguang Li, Yan-Pei Cao, Ding Liang, Xiaojuan Qi Title: TEXGen: a Generative Diffusion Model for Mesh Textures Arxiv: http://arxiv.org/abs/2411.14740v1 Abstract: While high-quality texture maps are essential for realistic 3D asset rendering, few studies have explored learning direct...

VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models 28.11.2024

🤗 Paper Upvotes: 8 | cs. CV, cs. CL Authors: Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, Lingpeng Kong, Qi Liu Title: VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models Arxiv: http://arxiv.org/abs/2411.17451v1 Abstract: Vision-language generative reward models (VL-GenRMs) play a crucia...

Learning 3D Representations from Procedural 3D Programs 28.11.2024

🤗 Paper Upvotes: 8 | cs. CV Authors: Xuweiyi Chen, Zezhou Cheng Title: Learning 3D Representations from Procedural 3D Programs Arxiv: http://arxiv.org/abs/2411.17467v1 Abstract: Self-supervised learning has emerged as a promising approach for acquiring transferable 3D representations from unlabeled 3D point clouds. Unlike 2D images, which are widely accessible, acquiring 3D assets requires specia...

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE 28.11.2024

🤗 Paper Upvotes: 7 | cs. CV Authors: Yongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang, XIngang Pan Title: SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE Arxiv: http://arxiv.org/abs/2411.16856v1 Abstract: Autoregressive models have demonstrated remarkable success across various fields, from large language models (LLMs) to large multimodal models (LMMs) a...

Material Anything: Generating Materials for Any 3D Object via Diffusion 27.11.2024

🤗 Paper Upvotes: 33 | cs. CV, cs. GR Authors: Xin Huang, Tengfei Wang, Ziwei Liu, Qing Wang Title: Material Anything: Generating Materials for Any 3D Object via Diffusion Arxiv: http://arxiv.org/abs/2411.15138v1 Abstract: We present Material Anything, a fully-automated, unified diffusion framework designed to generate physically-based materials for 3D objects. Unlike existing methods that rely on...

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator 27.11.2024

🤗 Paper Upvotes: 28 | cs. CV Authors: Chaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh Yoon Title: Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator Arxiv: http://arxiv.org/abs/2411.15466v1 Abstract: Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characte...

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge 27.11.2024

🤗 Paper Upvotes: 19 | cs. AI, cs. CL Authors: Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, Huan Liu Title: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge Arxiv: http://arxiv.org/abs/2411.16594v1 Abstract: Assessment and evaluation have long been criti...

O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? 27.11.2024

🤗 Paper Upvotes: 18 | cs. CL, cs. AI Authors: Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, Pengfei Liu Title: O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? Arxiv: http://arxiv.org/abs/2411.16489v1 Abstract: This paper presents a critical examination of current a...

MH-MoE: Multi-Head Mixture-of-Experts 27.11.2024

🤗 Paper Upvotes: 17 | cs. CL Authors: Shaohan Huang, Xun Wu, Shuming Ma, Furu Wei Title: MH-MoE: Multi-Head Mixture-of-Experts Arxiv: http://arxiv.org/abs/2411.16205v2 Abstract: Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper,...

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI 27.11.2024

🤗 Paper Upvotes: 15 | cs. CV Authors: Tianbin Li, Yanzhou Su, Wei Li, Bin Fu, Zhe Chen, Ziyan Huang, Guoan Wang, Chenglong Ma, Ying Chen, Ming Hu, Yanjun Li, Pengcheng Chen, Xiaowei Hu, Zhongying Deng, Yuanfeng Ji, Jin Ye, Yu Qiao, Junjun He Title: GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI Arxiv: http://arxiv.org/ab...

DreamRunner: Fine-Grained Storytelling Video Generation with Retrieval-Augmented Motion Adaptation 27.11.2024

🤗 Paper Upvotes: 13 | cs. CV, cs. AI, cs. CL Authors: Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal Title: DreamRunner: Fine-Grained Storytelling Video Generation with Retrieval-Augmented Motion Adaptation Arxiv: http://arxiv.org/abs/2411.16657v1 Abstract: Storytelling video generation (SVG) has recently emerged as a task to create long, multi-motion, multi-scene videos that consistentl...

Knowledge Transfer Across Modalities with Natural Language Supervision 27.11.2024

🤗 Paper Upvotes: 13 | cs. CV, 68T45 (Primary) 68T50 (Secondary), I.2.6 Authors: Carlo Alberto Barbano, Luca Molinaro, Emanuele Aiello, Marco Grangetto Title: Knowledge Transfer Across Modalities with Natural Language Supervision Arxiv: http://arxiv.org/abs/2411.15611v1 Abstract: We present a way to learn novel concepts by only using their textual description. We call this method Knowledge Transfe...

One Diffusion to Generate Them All 27.11.2024

🤗 Paper Upvotes: 13 | cs. CV, cs. AI Authors: Duong H. Le, Tuan Pham, Sangho Lee, Christopher Clark, Aniruddha Kembhavi, Stephan Mandt, Ranjay Krishna, Jiasen Lu Title: One Diffusion to Generate Them All Arxiv: http://arxiv.org/abs/2411.16318v1 Abstract: We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding acr...

VisualLens: Personalization through Visual History 27.11.2024

🤗 Paper Upvotes: 13 | cs. CV Authors: Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong Title: VisualLens: Personalization through Visual History Arxiv: http://arxiv.org/abs/2411.16034v1 Abstract: We hypothesize that a user's visual history with images reflecting their daily life, offers valuable insights into...

TÜLU 3: Pushing Frontiers in Open Language Model Post-Training 26.11.2024

🤗 Paper Upvotes: 38 | cs. CL Authors: Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, Hannaneh Hajishirzi Title:...

Style-Friendly SNR Sampler for Style-Driven Generation 26.11.2024

🤗 Paper Upvotes: 28 | cs. CV Authors: Jooyoung Choi, Chaehun Shin, Yeongtak Oh, Heeseung Kim, Sungroh Yoon Title: Style-Friendly SNR Sampler for Style-Driven Generation Arxiv: http://arxiv.org/abs/2411.14793v1 Abstract: Recent large-scale diffusion models generate high-quality images but struggle to learn new, personalized artistic styles, which limits the creation of unique style templates. Fine...

OminiControl: Minimal and Universal Control for Diffusion Transformer 26.11.2024

🤗 Paper Upvotes: 22 | cs. CV, cs. AI, cs. LG Authors: Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, Xinchao Wang Title: OminiControl: Minimal and Universal Control for Diffusion Transformer Arxiv: http://arxiv.org/abs/2411.15098v1 Abstract: In this paper, we introduce OminiControl, a highly versatile and parameter-efficient framework that integrates image conditions into pre-trained Diffu...

A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection 26.11.2024

🤗 Paper Upvotes: 15 | cs. CL, cs. LG, 68T50, I.2.7 Authors: Gabriel Chua, Shing Yee Chan, Shaun Khoo Title: A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection Arxiv: http://arxiv.org/abs/2411.12946v1 Abstract: Large Language Models are prone to off-topic misuse, where users may prompt these models to perform tasks beyond their intended scope....

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games 26.11.2024

🤗 Paper Upvotes: 14 | cs. AI Authors: Davide Paglieri, Bartłomiej Cupiał, Samuel Coward, Ulyana Piterbarg, Maciej Wolczyk, Akbir Khan, Eduardo Pignatelli, Łukasz Kuciński, Lerrel Pinto, Rob Fergus, Jakob Nicolaus Foerster, Jack Parker-Holder, Tim Rocktäschel Title: BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games Arxiv: http://arxiv.org/abs/2411.13543v1 Abstract: Large Language Models...

Large Multi-modal Models Can Interpret Features in Large Multi-modal Models 26.11.2024

🤗 Paper Upvotes: 12 | cs. CV, cs. CL Authors: Kaichen Zhang, Yifei Shen, Bo Li, Ziwei Liu Title: Large Multi-modal Models Can Interpret Features in Large Multi-modal Models Arxiv: http://arxiv.org/abs/2411.14982v1 Abstract: Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understan...

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection 26.11.2024

🤗 Paper Upvotes: 9 | cs. CV, cs. AI, cs. CL Authors: Songhao Han, Wei Huang, Hairong Shi, Le Zhuo, Xiu Su, Shifeng Zhang, Xu Zhou, Xiaojuan Qi, Yue Liao, Si Liu Title: VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Arxiv: http://arxiv.org/abs/2411.14794v1 Abstract: The advancement of Large Vision Language Models (LVLMs) has signific...

Efficient Long Video Tokenization via Coordinated-based Patch Reconstruction 26.11.2024

🤗 Paper Upvotes: 9 | cs. CV, cs. AI, cs. LG Authors: Huiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel, Younggyo Seo Title: Efficient Long Video Tokenization via Coordinated-based Patch Reconstruction Arxiv: http://arxiv.org/abs/2411.14762v1 Abstract: Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop...

MyTimeMachine: Personalized Facial Age Transformation 26.11.2024

🤗 Paper Upvotes: 8 | cs. CV Authors: Luchao Qi, Jiaye Wu, Bang Gong, Annie N. Wang, David W. Jacobs, Roni Sengupta Title: MyTimeMachine: Personalized Facial Age Transformation Arxiv: http://arxiv.org/abs/2411.14521v1 Abstract: Facial aging is a complex process, highly dependent on multiple factors like gender, ethnicity, lifestyle, etc., making it extremely challenging to learn a global aging pri...

Novel View Extrapolation with Video Diffusion Priors 26.11.2024

🤗 Paper Upvotes: 7 | cs. CV Authors: Kunhao Liu, Ling Shao, Shijian Lu Title: Novel View Extrapolation with Video Diffusion Priors Arxiv: http://arxiv.org/abs/2411.14208v1 Abstract: The field of novel view synthesis has made significant strides thanks to the development of radiance field methods. However, most radiance field techniques are far better at novel view interpolation than novel view ex...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.