Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
BrushEdit: All-In-One Image Inpainting and Editing 18.12.2024 27:48
🤗 Upvotes: 24 | cs. CV, cs. AI Authors: Yaowei Li, Yuxuan Bian, Xuan Ju, Zhaoyang Zhang, Ying Shan, Yuexian Zou, Qiang Xu Title: BrushEdit: All-In-One Image Inpainting and Editing Arxiv: http://arxiv.org/abs/2412.10316v2 Abstract: Image editing has advanced significantly with the development of diffusion models using both inversion-based and instruction-based methods. However, current inversion-b...
ColorFlow: Retrieval-Augmented Image Sequence Colorization 18.12.2024 22:32
🤗 Upvotes: 20 | cs. CV Authors: Junhao Zhuang, Xuan Ju, Zhaoyang Zhang, Yong Liu, Shiyi Zhang, Chun Yuan, Ying Shan Title: ColorFlow: Retrieval-Augmented Image Sequence Colorization Arxiv: http://arxiv.org/abs/2412.11815v1 Abstract: Automatic black-and-white image sequence colorization while preserving character and object identity (ID) is a complex task with significant market demand, such as in...
Smaller Language Models Are Better Instruction Evolvers 18.12.2024 23:17
🤗 Upvotes: 16 | cs. CL Authors: Tingfeng Hui, Lulu Zhao, Guanting Dong, Yaqi Zhang, Hua Zhou, Sen Su Title: Smaller Language Models Are Better Instruction Evolvers Arxiv: http://arxiv.org/abs/2412.11231v1 Abstract: Instruction tuning has been widely used to unleash the complete potential of large language models. Notably, complex and diverse instructions are of significant importance as they can...
Causal Diffusion Transformers for Generative Modeling 18.12.2024 23:47
🤗 Upvotes: 16 | cs. CV Authors: Chaorui Deng, Deyao Zhu, Kunchang Li, Shi Guang, Haoqi Fan Title: Causal Diffusion Transformers for Generative Modeling Arxiv: http://arxiv.org/abs/2412.12095v2 Abstract: We introduce Causal Diffusion as the autoregressive (AR) counterpart of Diffusion models. It is a next-token(s) forecasting framework that is friendly to both discrete and continuous modalities an...
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models 18.12.2024 23:05
🤗 Upvotes: 11 | cs. CL, cs. AI, cs. LG Authors: Jiale Cheng, Xiao Liu, Cunxiang Wang, Xiaotao Gu, Yida Lu, Dan Zhang, Yuxiao Dong, Jie Tang, Hongning Wang, Minlie Huang Title: SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Arxiv: http://arxiv.org/abs/2412.11605v1 Abstract: Instruction-following is a fundamental capability of language models,...
IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations 18.12.2024 20:29
🤗 Upvotes: 11 | cs. CV Authors: Zhibing Li, Tong Wu, Jing Tan, Mengchen Zhang, Jiaqi Wang, Dahua Lin Title: IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations Arxiv: http://arxiv.org/abs/2412.12083v1 Abstract: Capturing geometric and material information from images remains a fundamental challenge in computer vision and graphics. Traditional optimization-based me...
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs 18.12.2024 21:15
🤗 Upvotes: 10 | cs. RO, cs. AI, cs. CV Authors: Xinli Xu, Wenhang Ge, Dicong Qiu, ZhiFei Chen, Dongyu Yan, Zhuoyun Liu, Haoyu Zhao, Hanfeng Zhao, Shunsi Zhang, Junwei Liang, Ying-Cong Chen Title: GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs Arxiv: http://arxiv.org/abs/2412.11258v1 Abstract: Estimating physical properties for visual data is a crucial task in computer...
Apollo: An Exploration of Video Understanding in Large Multimodal Models 17.12.2024 24:59
🤗 Upvotes: 91 | cs. CV, cs. AI Authors: Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, Xide Xia Title: Apollo: An Exploration of Video Understanding in Large Multimodal Models Arxiv: http://arxiv.org/abs/2412.10360v1 Abstract: Despite the rapid integration of video perception capabi...
GenEx: Generating an Explorable World 17.12.2024 21:28
🤗 Upvotes: 65 | cs. CV, cs. RO Authors: Taiming Lu, Tianmin Shu, Junfei Xiao, Luoxin Ye, Jiahao Wang, Cheng Peng, Chen Wei, Daniel Khashabi, Rama Chellappa, Alan Yuille, Jieneng Chen Title: GenEx: Generating an Explorable World Arxiv: http://arxiv.org/abs/2412.09624v1 Abstract: Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development...
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 17.12.2024 25:15
🤗 Upvotes: 29 | cs. CV Authors: Hao Li, Changyao Tian, Jie Shao, Xizhou Zhu, Zhaokai Wang, Jinguo Zhu, Wenhan Dou, Xiaogang Wang, Hongsheng Li, Lewei Lu, Jifeng Dai Title: SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Arxiv: http://arxiv.org/abs/2412.09604v1 Abstract: The remarkable success of Large Language Models (LLMs) has extended to...
BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities 17.12.2024 17:46
🤗 Upvotes: 24 | cs. CV Authors: Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled Aldahmani, Fahad Khan, Rao Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal Title: BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities Arxiv: http://arxiv.org/abs/2412.07769v1 Abstract: This paper introduces BiMediX2, a bilingual (Arabic-En...
Large Action Models: From Inception to Implementation 17.12.2024 22:15
🤗 Upvotes: 23 | cs. AI Authors: Lu Wang, Fangkai Yang, Chaoyun Zhang, Junting Lu, Jiaxu Qian, Shilin He, Pu Zhao, Bo Qiao, Ray Huang, Si Qin, Qisheng Su, Jiayi Ye, Yudi Zhang, Jian-Guang Lou, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang Title: Large Action Models: From Inception to Implementation Arxiv: http://arxiv.org/abs/2412.10047v1 Abstract: As AI continues to advance, there is a g...
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption 17.12.2024 21:01
🤗 Upvotes: 17 | cs. CV, cs. AI Authors: Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou, Zhenheng Yang, Chaoyou Fu, Xiang Li, Jian Yang, Ying Tai Title: InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Arxiv: http://arxiv.org/abs/2412.09283v1 Abstract: Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically...
FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion 17.12.2024 21:32
🤗 Upvotes: 13 | cs. CV Authors: Haonan Qiu, Shiwei Zhang, Yujie Wei, Ruihang Chu, Hangjie Yuan, Xiang Wang, Yingya Zhang, Ziwei Liu Title: FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion Arxiv: http://arxiv.org/abs/2412.09626v1 Abstract: Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack...
ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven Generation 17.12.2024 21:41
🤗 Upvotes: 10 | cs. CV Authors: Daniel Winter, Asaf Shul, Matan Cohen, Dana Berman, Yael Pritch, Alex Rav-Acha, Yedid Hoshen Title: ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven Generation Arxiv: http://arxiv.org/abs/2412.08645v1 Abstract: This paper introduces a tuning-free method for both object insertion and subject-driven generation. The task involves composing an obj...
FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing 17.12.2024 21:47
🤗 Upvotes: 8 | cs. CV Authors: Yingying Deng, Xiangyu He, Changwang Mei, Peisong Wang, Fan Tang Title: FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing Arxiv: http://arxiv.org/abs/2412.07517v1 Abstract: Though Rectified Flows (ReFlows) with distillation offers a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and fol...
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers 17.12.2024 18:48
🤗 Upvotes: 7 | cs. CV Authors: Yusuf Dalva, Kavana Venkatesh, Pinar Yanardag Title: FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers Arxiv: http://arxiv.org/abs/2412.09611v1 Abstract: Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visua...
Phi-4 Technical Report 14.12.2024 22:12
🤗 Upvotes: 40 | cs. CL, cs. AI Authors: Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu, C...
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions 14.12.2024 24:28
🤗 Upvotes: 30 | cs. CV, cs. AI, cs. CL Authors: Jiarui Zhang, Ollie Liu, Tianyu Yu, Jinyi Hu, Willie Neiswanger Title: Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions Arxiv: http://arxiv.org/abs/2412.08737v1 Abstract: Multimodal large language models (MLLMs) have made rapid progress in recent years, yet continue to struggle with low-level visual perception (...
Multimodal Latent Language Modeling with Next-Token Diffusion 14.12.2024 22:35
🤗 Upvotes: 21 | cs. CL, cs. CV, cs. LG Authors: Yutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng, Li Dong, Shaohan Huang, Jianyong Wang, Furu Wei Title: Multimodal Latent Language Modeling with Next-Token Diffusion Arxiv: http://arxiv.org/abs/2412.08635v1 Abstract: Multimodal generative models require a unified approach to handle both discrete data (e.g., text and code) and continuous data (e.g....
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM 14.12.2024 21:54
🤗 Upvotes: 17 | cs. CV Authors: Zhuofan Zong, Dongzhi Jiang, Bingqi Ma, Guanglu Song, Hao Shao, Dazhong Shen, Yu Liu, Hongsheng Li Title: EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM Arxiv: http://arxiv.org/abs/2412.09618v1 Abstract: Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods most...
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials 14.12.2024 18:51
🤗 Upvotes: 16 | cs. CL Authors: Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, Tao Yu Title: AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials Arxiv: http://arxiv.org/abs/2412.09605v1 Abstract: Graphical User Interface (GUI) agents hold great potential for automating complex tasks across diverse digital environments, from web appli...
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training 14.12.2024 19:08
🤗 Upvotes: 14 | cs. CV Authors: Dongting Hu, Jierun Chen, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Rajesh Singh, Yerlan Idelbayev, Junli Cao, Yanyu Li, Kwang-Ting Cheng, S. -H. Gary Chan, Mingming Gong, Sergey Tulyakov, Anil Kag, Yanwu Xu, Jian Ren Title: SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architect...
Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion 14.12.2024 22:37
🤗 Upvotes: 13 | cs. CV Authors: Zexin He, Tengfei Wang, Xin Huang, Xingang Pan, Ziwei Liu Title: Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion Arxiv: http://arxiv.org/abs/2412.09593v1 Abstract: Recovering the geometry and materials of objects from a single image is challenging due to its under-constrained nature. In this paper, we present Neu...
JuStRank: Benchmarking LLM Judges for System Ranking 14.12.2024 21:10
🤗 Upvotes: 9 | cs. CL, cs. AI, cs. LG Authors: Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai Title: JuStRank: Benchmarking LLM Judges for System Ranking Arxiv: http://arxiv.org/abs/2412.09569v1 Abstract: Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations availabl...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.