Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Latent Diffusion Model without Variational Autoencoder 21.10.2025

🤗 Upvotes: 30 | cs. CV, cs. AI Authors: Minglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan, Xiaoshi Wu, Xintao Wang, Pengfei Wan, Jie Zhou, Jiwen Lu Title: Latent Diffusion Model without Variational Autoencoder Arxiv: http://arxiv.org/abs/2510.15301v2 Abstract: Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs)...

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA 18.10.2025

🤗 Upvotes: 81 | cs. CL Authors: Elisei Rykov, Kseniia Petrushina, Maksim Savkin, Valerii Olisov, Artem Vazhentsev, Kseniia Titova, Alexander Panchenko, Vasily Konovalov, Julia Belikova Title: When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA Arxiv: http://arxiv.org/abs/2510.04849v1 Abstract: Hallucination detection remains a fundamental challenge for the safe...

Agentic Entropy-Balanced Policy Optimization 18.10.2025

🤗 Upvotes: 78 | cs. LG, cs. AI, cs. CL, cs. IR Authors: Guanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao, Xiaoxi Li, Jiajie Jin, Jinghan Yang, Hangyu Mao, Fuzheng Zhang, Kun Gai, Guorui Zhou, Yutao Zhu, Ji-Rong Wen, Zhicheng Dou Title: Agentic Entropy-Balanced Policy Optimization Arxiv: http://arxiv.org/abs/2510.14545v1 Abstract: Recently, Agentic Reinforcement Learning (Agentic RL) has m...

WithAnyone: Towards Controllable and ID Consistent Image Generation 18.10.2025

🤗 Upvotes: 65 | cs. CV, cs. AI Authors: Hengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang, Shuhan Wu, Rui Wang, Xianfang Zeng, Daxin Jiang, Gang Yu, Xingjun Ma, Yu-Gang Jiang Title: WithAnyone: Towards Controllable and ID Consistent Image Generation Arxiv: http://arxiv.org/abs/2510.14975v1 Abstract: Identity-consistent generation has become an important focus in text-to-image research, with recent m...

AI for Service: Proactive Assistance with AI Glasses 18.10.2025

🤗 Upvotes: 60 | cs. AI, cs. CL, cs. CV Authors: Zichen Wen, Yiyu Wang, Chenfei Liao, Boxue Yang, Junxian Li, Weifeng Liu, Haocong He, Bolong Feng, Xuyang Liu, Yuanhuiyi Lyu, Xu Zheng, Xuming Hu, Linfeng Zhang Title: AI for Service: Proactive Assistance with AI Glasses Arxiv: http://arxiv.org/abs/2510.14359v1 Abstract: In an era where AI is evolving from a passive tool into an active and adaptive...

From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 18.10.2025

🤗 Upvotes: 51 | cs. CV, cs. AI Authors: Haiwen Diao, Mingxuan Li, Silei Wu, Linjun Dai, Xiaohua Wang, Hanming Deng, Lewei Lu, Dahua Lin, Ziwei Liu Title: From Pixels to Words -- Towards Native Vision-Language Primitives at Scale Arxiv: http://arxiv.org/abs/2510.14979v1 Abstract: The edifice of native Vision-Language Models (VLMs) has emerged as a rising contender to typical modular VLMs, shaped b...

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints 18.10.2025

🤗 Upvotes: 46 | cs. CV Authors: Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, Chen Zhu, Bingze Song, Fangyuan Mao, Jiahong Wu, Xiangxiang Chu, Kaiqi Huang Title: ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints Arxiv: http://arxiv.org/abs/2510.14847v1 Abstract: Video generation models have achieved remarkable progress, particularly excelling...

Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents 18.10.2025

🤗 Upvotes: 30 | cs. CL, cs. AI, cs. LG Authors: Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying Title: Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents Arxiv: http://arxiv.org/abs/2510.14967v1 Abstract: Large language model (LLM)-based agents are increasingly trained with reinforcement learning (RL)...

LaSeR: Reinforcement Learning with Last-Token Self-Rewarding 18.10.2025

🤗 Upvotes: 30 | cs. CL, cs. AI, cs. LG Authors: Wenkai Yang, Weijie Liu, Ruobing Xie, Yiju Guo, Lulu Wu, Saiyong Yang, Yankai Lin Title: LaSeR: Reinforcement Learning with Last-Token Self-Rewarding Arxiv: http://arxiv.org/abs/2510.14943v1 Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a core paradigm for enhancing the reasoning capabilities of Large Langua...

TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar 18.10.2025

🤗 Upvotes: 27 | cs. CL, cs. AI, cs. LG, cs. PL, cs. SE Authors: Yinxi Li, Yuntian Deng, Pengyu Nie Title: TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar Arxiv: http://arxiv.org/abs/2510.14972v1 Abstract: Large language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but dr...

BitNet Distillation 18.10.2025

🤗 Upvotes: 26 | cs. LG, cs. CL Authors: Xun Wu, Shaohan Huang, Wenhui Wang, Ting Song, Li Dong, Yan Xia, Furu Wei Title: BitNet Distillation Arxiv: http://arxiv.org/abs/2510.13998v1 Abstract: In this paper, we present BitNet Distillation (BitDistill), a lightweight pipeline that fine-tunes off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (i.e., ternary weights {-1, 0, 1}) fo...

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model 16.10.2025

🤗 Upvotes: 133 | cs. RO Authors: Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, Haoang Li Title: Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model Arxiv: http://arxiv.org/abs/2510.12276v1 Abstract: Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instruc...

Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training 16.10.2025

🤗 Upvotes: 91 | cs. CV Authors: Jiachen Lei, Keli Liu, Julius Berner, Haiming Yu, Hongkai Zheng, Jiahong Wu, Xiangxiang Chu Title: Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training Arxiv: http://arxiv.org/abs/2510.12586v1 Abstract: Pixel-space generative models are often more difficult to train and generally underperform compared to their latent-space counterpa...

DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation 16.10.2025

🤗 Upvotes: 90 | cs. CL Authors: Enze Zhang, Jiaying Wang, Mengxi Xiao, Jifei Liu, Ziyan Kuang, Rui Dong, Eric Dong, Sophia Ananiadou, Min Peng, Qianqian Xie Title: DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation Arxiv: http://arxiv.org/abs/2510.09116v2 Abstract: Large language models (LLMs) have substantially advanced machine translation (MT), yet their effective...

Scaling Language-Centric Omnimodal Representation Learning 16.10.2025

🤗 Upvotes: 82 | cs. CL, cs. AI, cs. CV Authors: Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong Title: Scaling Language-Centric Omnimodal Representation Learning Arxiv: http://arxiv.org/abs/2510.11693v1 Abstract: Recent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising res...

Robot Learning: A Tutorial 16.10.2025

🤗 Upvotes: 44 | cs. RO, cs. LG Authors: Francesco Capuano, Caroline Pascal, Adil Zouitine, Thomas Wolf, Michel Aractingi Title: Robot Learning: A Tutorial Arxiv: http://arxiv.org/abs/2510.12403v1 Abstract: Robot learning is at an inflection point, driven by rapid advancements in machine learning and the growing availability of large-scale robotics data. This shift from classical, model-based meth...

Detect Anything via Next Point Prediction 16.10.2025

🤗 Upvotes: 34 | cs. CV Authors: Qing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong, Zhaoyang Zeng, Yihao Chen, Tianhe Ren, Junzhi Yu, Lei Zhang Title: Detect Anything via Next Point Prediction Arxiv: http://arxiv.org/abs/2510.12798v1 Abstract: Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts hav...

A Survey of Vibe Coding with Large Language Models 16.10.2025

🤗 Upvotes: 31 | cs. AI Authors: Yuyao Ge, Lingrui Mei, Zenghao Duan, Tianhao Li, Yujia Zheng, Yiwei Wang, Lexin Wang, Jiayu Yao, Tianyu Liu, Yujun Cai, Baolong Bi, Fangda Guo, Jiafeng Guo, Shenghua Liu, Xueqi Cheng Title: A Survey of Vibe Coding with Large Language Models Arxiv: http://arxiv.org/abs/2510.12399v1 Abstract: The advancement of large language models (LLMs) has catalyzed a paradigm sh...

FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution 16.10.2025

🤗 Upvotes: 30 | cs. CV Authors: Junhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li, Yihao Liu, Chun Yuan, Tianfan Xue Title: FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution Arxiv: http://arxiv.org/abs/2510.12747v1 Abstract: Diffusion models have recently advanced video restoration, but applying them to real-world video super-resolution (VSR) remains challenging due to high l...

Dr.LLM: Dynamic Layer Routing in LLMs 16.10.2025

🤗 Upvotes: 27 | cs. CL, cs. AI, cs. LG Authors: Ahmed Heakl, Martin Gubri, Salman Khan, Sangdoo Yun, Seong Joon Oh Title: Dr. LLM: Dynamic Layer Routing in LLMs Arxiv: http://arxiv.org/abs/2510.12773v1 Abstract: Large Language Models (LLMs) process every token through all layers of a transformer stack, causing wasted computation on simple queries and insufficient flexibility for harder ones that...

Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models 16.10.2025

🤗 Upvotes: 26 | cs. LG, cs. AI Authors: Youngrok Park, Hojung Jung, Sangmin Bae, Se-Young Yun Title: Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models Arxiv: http://arxiv.org/abs/2510.11057v1 Abstract: Diffusion models have achieved remarkable success as generative models. However, even a well-trained model can accumulate errors throughout the generation process. These errors...

QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs 15.10.2025

🤗 Upvotes: 106 | cs. LG, cs. CL, cs. CV Authors: Wei Huang, Yi Ge, Shuai Yang, Yicheng Xiao, Huizi Mao, Yujun Lin, Hanrong Ye, Sifei Liu, Ka Chun Cheung, Hongxu Yin, Yao Lu, Xiaojuan Qi, Song Han, Yukang Chen Title: QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs Arxiv: http://arxiv.org/abs/2510.11696v1 Abstract: We propose QeRL, a Quantization-enhanced Reinforcem...

Diffusion Transformers with Representation Autoencoders 15.10.2025

🤗 Upvotes: 93 | cs. CV, cs. LG Authors: Boyang Zheng, Nanye Ma, Shengbang Tong, Saining Xie Title: Diffusion Transformers with Representation Autoencoders Arxiv: http://arxiv.org/abs/2510.11690v1 Abstract: Latent generative modeling, where a pretrained autoencoder maps pixels into a latent space for the diffusion process, has become the standard strategy for Diffusion Transformers (DiT); however,...

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs 15.10.2025

🤗 Upvotes: 39 | cs. AI Authors: Caorui Li, Yu Chen, Yiyan Ji, Jin Xu, Zhenyu Cui, Shihao Li, Yuanxing Zhang, Jiafu Tang, Zhenghao Song, Dingling Zhang, Ying He, Haoxiang Liu, Yuxuan Wang, Qiufeng Wang, Zhenhe Wu, Jiehui Luo, Zhiyu Pan, Weihao Xie, Chenchen Zhang, Zhaohui Wang, Jiayi Tian, Yanghai Wang, Zhe Cao, Minxin Dai, Ke Wang, Runzhe Wen, Yinghao Ma, Yaning Pan, Sungkyun Chang, Termeh Taheri...

Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States 15.10.2025

🤗 Upvotes: 37 | cs. CL Authors: Qinglin Zhu, Yizhen Yao, Runcong Zhao, Yanzheng Xiang, Amrutha Saseendran, Chen Jin, Philip Alexander Teare, Bin Liang, Yulan He, Lin Gui Title: Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States Arxiv: http://arxiv.org/abs/2510.11052v1 Abstract: Autoregressive (AR) models remain the standard for natural language generat...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.