Jingwen Liang, Gengyu Wang
Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art
Author
Jingwen Liang, Gengyu Wang
Category
Podcast website
Latest episode
Jul 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 04.01.2025 23:53
🤗 Upvotes: 45 | cs. CV, cs. CL, cs. LG Authors: Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun, Yongliang Shen, Weiming Lu, Deli Zhao, Yueting Zhuang, Lidong Bing Title: 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Arxiv: http://arxiv.org/abs/2501.00958v1 Abstract: Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand t...
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings 04.01.2025 23:32
🤗 Upvotes: 30 | cs. CL Authors: Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, Yibo Miao, Yunlong Feng, Zekun Wang, Jian Yang, Zeyu Cui, Yang Fan, Yichang Zhang, Binyuan Hui, Junyang Lin Title: CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Arxiv: http://arxiv.org/abs/2501.01257v1 Abstract: With...
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control 04.01.2025 19:15
🤗 Upvotes: 30 | cs. CV Authors: Yuanpeng Tu, Hao Luo, Xi Chen, Sihui Ji, Xiang Bai, Hengshuang Zhao Title: VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Arxiv: http://arxiv.org/abs/2501.01427v1 Abstract: Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appea...
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models 04.01.2025 24:49
🤗 Upvotes: 25 | cs. CV, cs. LG Authors: Jingfeng Yao, Xinggang Wang Title: Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Arxiv: http://arxiv.org/abs/2501.01423v1 Abstract: Latent diffusion models with Transformer architectures excel at generating high-fidelity images. However, recent studies reveal an optimization dilemma in this two-stage design: while inc...
ProgCo: Program Helps Self-Correction of Large Language Models 04.01.2025 20:19
🤗 Upvotes: 17 | cs. CL, cs. AI, cs. LG Authors: Xiaoshuai Song, Yanan Wu, Weixun Wang, Jiaheng Liu, Wenbo Su, Bo Zheng Title: ProgCo: Program Helps Self-Correction of Large Language Models Arxiv: http://arxiv.org/abs/2501.01264v1 Abstract: Self-Correction aims to enable large language models (LLMs) to self-verify and self-refine their initial responses without external feedback. However, LLMs oft...
MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models 04.01.2025 25:32
🤗 Upvotes: 16 | cs. CL Authors: Mahir Labib Dihan, Md Tanvir Hassan, Md Tanvir Parvez, Md Hasebul Hasan, Md Almash Alam, Muhammad Aamir Cheema, Mohammed Eunus Ali, Md Rizwan Parvez Title: MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models Arxiv: http://arxiv.org/abs/2501.00316v1 Abstract: Recent advancements in foundation models have enhanced AI systems' capabilities in...
A3: Android Agent Arena for Mobile GUI Agents 04.01.2025 23:35
🤗 Upvotes: 15 | cs. AI Authors: Yuxiang Chai, Hanhao Li, Jiayu Zhang, Liang Liu, Guozhi Wang, Shuai Ren, Siyuan Huang, Hongsheng Li Title: A3: Android Agent Arena for Mobile GUI Agents Arxiv: http://arxiv.org/abs/2501.01149v1 Abstract: AI agents have become increasingly prevalent in recent years, driven by significant advancements in the field of large language models (LLMs). Mobile GUI agents, a...
MLLM-as-a-Judge for Image Safety without Human Labeling 04.01.2025 22:20
🤗 Upvotes: 14 | cs. CV, cs. CL, cs. CY, cs. LG Authors: Zhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li, Ligong Han, Harihar Subramanyam, Li Chen, Jianfa Chen, Nan Jiang, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas, Ankit Jain Title: MLLM-as-a-Judge for Image Safety without Human Labeling Arxiv: http://arxiv.org/abs/2501.00192v1 Abstract: Image content safety has...
Dynamic Scaling of Unit Tests for Code Reward Modeling 04.01.2025 21:52
🤗 Upvotes: 13 | cs. CL, cs. SE Authors: Zeyao Ma, Xiaokang Zhang, Jing Zhang, Jifan Yu, Sijia Luo, Jie Tang Title: Dynamic Scaling of Unit Tests for Code Reward Modeling Arxiv: http://arxiv.org/abs/2501.01054v1 Abstract: Current large language models (LLMs) often struggle to produce accurate responses on the first attempt for complex reasoning tasks like code generation. Prior research tackles th...
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis 03.01.2025 22:38
🤗 Upvotes: 52 | cs. AI, cs. CL, cs. CV, cs. HC Authors: Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, Zhiyong Wu Title: OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis Arxiv: http://arxiv.org/abs/2412.19723v1 Abstract: Graphical User Int...
Xmodel-2 Technical Report 03.01.2025 17:16
🤗 Upvotes: 13 | cs. AI Authors: Wang Qun, Liu Yang, Lin Qingquan, Qu Zhijiu, Jiang Ling Title: Xmodel-2 Technical Report Arxiv: http://arxiv.org/abs/2412.19638v1 Abstract: Xmodel-2 is a 1.2-billion-parameter large language model designed specifically for reasoning tasks. Its architecture enables different model scales to share a unified set of hyperparameters, allowing for extensive experimentati...
Are Vision-Language Models Truly Understanding Multi-vision Sensor? 03.01.2025 24:50
🤗 Upvotes: 9 | cs. CV Authors: Sangyun Chung, Youngjoon Yu, Youngchae Chee, Se Yeon Kim, Byung-Kwan Lee, Yong Man Ro Title: Are Vision-Language Models Truly Understanding Multi-vision Sensor? Arxiv: http://arxiv.org/abs/2412.20750v1 Abstract: Large-scale Vision-Language Models (VLMs) have advanced by aligning vision inputs with text, significantly improving performance in computer vision tasks. M...
HUNYUANPROVER: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving 03.01.2025 20:48
🤗 Upvotes: 4 | cs. AI, cs. CL Authors: Yang Li, Dong Du, Linfeng Song, Chen Li, Weikang Wang, Tao Yang, Haitao Mi Title: HUNYUANPROVER: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving Arxiv: http://arxiv.org/abs/2412.20735v2 Abstract: We introduce HunyuanProver, an language model finetuned from the Hunyuan 7B for interactive automatic theorem proving with...
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control 03.01.2025 22:06
🤗 Upvotes: 2 | cs. CV Authors: Shaojin Wu, Fei Ding, Mengqi Huang, Wei Liu, Qian He Title: VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control Arxiv: http://arxiv.org/abs/2412.20800v1 Abstract: While diffusion models show extraordinary talents in text-to-image generation, they may still fail to generate highly aesthetic images. More specifically, there is still a gap...
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs 02.01.2025 20:07
🤗 Upvotes: 13 | cs. CL Authors: Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, Dong Yu Title: Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Arxiv: http://arxiv.org/abs/2412.21187v1 Abstract: The remarkable performance of models like the OpenAI o1 can be attribut...
OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System 02.01.2025 18:53
🤗 Upvotes: 11 | cs. CL, cs. AI, cs. DB, cs. IR, cs. LG Authors: Yujie Luo, Xiangyuan Ru, Kangwei Liu, Lin Yuan, Mengshu Sun, Ningyu Zhang, Lei Liang, Zhiqiang Zhang, Jun Zhou, Lanning Wei, Da Zheng, Haofen Wang, Huajun Chen Title: OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System Arxiv: http://arxiv.org/abs/2412.20005v1 Abstract: We introduce OneKE, a dockerized schema...
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization 01.01.2025 25:04
🤗 Upvotes: 39 | cs. CV Authors: Yang Shen, Xiu-Shen Wei, Yifan Sun, Yuxin Song, Tao Yuan, Jian Jin, Heyang Xu, Yazhou Yao, Errui Ding Title: Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization Arxiv: http://arxiv.org/abs/2412.18525v2 Abstract: Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language...
On the Compositional Generalization of Multimodal LLMs for Medical Imaging 01.01.2025 22:45
🤗 Upvotes: 29 | cs. CV, cs. AI, cs. CL, cs. LG Authors: Zhenyang Cai, Junying Chen, Rongsheng Wang, Weihong Wang, Yonglin Deng, Dingjie Song, Yize Chen, Zixu Zhang, Benyou Wang Title: On the Compositional Generalization of Multimodal LLMs for Medical Imaging Arxiv: http://arxiv.org/abs/2412.20070v1 Abstract: Multimodal large language models (MLLMs) hold significant potential in the medical field,...
Bringing Objects to Life: 4D generation from 3D objects 01.01.2025 21:48
🤗 Upvotes: 24 | cs. CV Authors: Ohad Rahamim, Ori Malca, Dvir Samuel, Gal Chechik Title: Bringing Objects to Life: 4D generation from 3D objects Arxiv: http://arxiv.org/abs/2412.20422v1 Abstract: Recent advancements in generative modeling now enable the creation of 4D content (moving 3D objects) controlled with text prompts. 4D generation has large potential in applications like virtual worlds, m...
Efficiently Serving LLM Reasoning Programs with Certaindex 01.01.2025 20:19
🤗 Upvotes: 20 | cs. LG, cs. CL Authors: Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Aurick Qiao, Hao Zhang Title: Efficiently Serving LLM Reasoning Programs with Certaindex Arxiv: http://arxiv.org/abs/2412.20993v1 Abstract: The rapid evolution of large language models (LLMs) has unlocked their capabilities in advanced reasoning tasks like mathematical problem-solving, code gener...
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization 01.01.2025 21:15
🤗 Upvotes: 14 | cs. SD, cs. AI, cs. CL, eess. AS Authors: Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Rafael Valle, Bryan Catanzaro, Soujanya Poria Title: TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Arxiv: http://arxiv.org/abs/2412.21037v1 Abstract: We introduce TangoFlux, an efficient Text-to-Audio (TTA)...
Edicho: Consistent Image Editing in the Wild 01.01.2025 22:47
🤗 Upvotes: 13 | cs. CV Authors: Qingyan Bai, Hao Ouyang, Yinghao Xu, Qiuyu Wang, Ceyuan Yang, Ka Leong Cheng, Yujun Shen, Qifeng Chen Title: Edicho: Consistent Image Editing in the Wild Arxiv: http://arxiv.org/abs/2412.21079v1 Abstract: As a verified need, consistent editing across in-the-wild images remains a technical challenge arising from various unmanageable factors, like object poses, light...
Facilitating large language model Russian adaptation with Learned Embedding Propagation 01.01.2025 22:12
🤗 Upvotes: 6 | cs. CL, cs. AI Authors: Mikhail Tikhomirov, Daniil Chernyshev Title: Facilitating large language model Russian adaptation with Learned Embedding Propagation Arxiv: http://arxiv.org/abs/2412.21140v1 Abstract: Rapid advancements of large language model (LLM) technologies led to the introduction of powerful open-source instruction-tuned LLMs that have the same text generation quality...
Training Software Engineering Agents and Verifiers with SWE-Gym 01.01.2025 26:54
🤗 Upvotes: 6 | cs. SE, cs. CL Authors: Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, Yizhe Zhang Title: Training Software Engineering Agents and Verifiers with SWE-Gym Arxiv: http://arxiv.org/abs/2412.21139v1 Abstract: We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task in...
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation 01.01.2025 20:54
🤗 Upvotes: 5 | cs. SE, cs. CL Authors: Zhaojian Yu, Yilun Zhao, Arman Cohan, Xiao-Ping Zhang Title: HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Arxiv: http://arxiv.org/abs/2412.21199v1 Abstract: We introduce self-invoking code generation, a new task designed to evaluate the progressive reasoning and problem-solving capabilities of LLMs. In this ta...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.