Jingwen Liang, Gengyu Wang

Daily Paper Cast

Science EN ↓ 2000 episodes

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.comCreator:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236Cover Image by Kawen Kuang https://kawen.art

Author

Jingwen Liang, Gengyu Wang

Category

Science

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Building Trust: Foundations of Security, Safety and Transparency in AI 21.11.2024

🤗 Paper Upvotes: 8 | cs. CY, cs. AI, cs. CL Authors: Huzaifa Sidhpurwala, Garth Mollett, Emily Fox, Mark Bestavros, Huamin Chen Title: Building Trust: Foundations of Security, Safety and Transparency in AI Arxiv: http://arxiv.org/abs/2411.12275v1 Abstract: This paper explores the rapidly evolving ecosystem of publicly available AI models, and their potential implications on the security and safet...

SEAGULL: No-reference Image Quality Assessment for Regions of Interest via Vision-Language Instruction Tuning 21.11.2024

🤗 Paper Upvotes: 5 | cs. CV Authors: Zewen Chen, Juan Wang, Wen Wang, Sunhan Xu, Hang Xiong, Yun Zeng, Jian Guo, Shuxun Wang, Chunfeng Yuan, Bing Li, Weiming Hu Title: SEAGULL: No-reference Image Quality Assessment for Regions of Interest via Vision-Language Instruction Tuning Arxiv: http://arxiv.org/abs/2411.10161v1 Abstract: Existing Image Quality Assessment (IQA) methods achieve remarkable suc...

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages 21.11.2024

🤗 Paper Upvotes: 3 | cs. CL, cs. AI Authors: S. Tamang, D. J. Bora Title: Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Arxiv: http://arxiv.org/abs/2411.12240v1 Abstract: Large Language Models (LLMs) based on transformer architectures have revolutionized a variety of domains, with tokenization playing a pivotal role in their pre-processing and fine-tun...

Generative World Explorer 20.11.2024

🤗 Paper Upvotes: 38 | cs. CV Authors: Taiming Lu, Tianmin Shu, Alan Yuille, Daniel Khashabi, Jieneng Chen Title: Generative World Explorer Arxiv: http://arxiv.org/abs/2411.11844v2 Abstract: Planning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their environment to update their bel...

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices 20.11.2024

🤗 Paper Upvotes: 31 | cs. CV, cs. CL Authors: Xudong Lu, Yinghao Chen, Cheng Chen, Hui Tan, Boheng Chen, Yina Xie, Rui Hu, Guanxin Tan, Renshou Wu, Yan Hu, Yi Zeng, Lei Wu, Liuyang Bian, Zhaoxiong Wang, Long Liu, Yanzhou Yang, Han Xiao, Aojun Zhou, Yafei Wen, Xiaoxin Chen, Shuai Ren, Hongsheng Li Title: BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Dev...

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering 20.11.2024

🤗 Paper Upvotes: 13 | cs. AI, cs. CL, stat. ML Authors: Xinyan Guan, Yanjiang Liu, Xinyu Lu, Boxi Cao, Ben He, Xianpei Han, Le Sun, Jie Lou, Bowen Yu, Yaojie Lu, Hongyu Lin Title: Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Arxiv: http://arxiv.org/abs/2411.11504v1 Abstract: The evolution of machine learning has increasi...

AnimateAnything: Consistent and Controllable Animation for Video Generation 20.11.2024

🤗 Paper Upvotes: 12 | cs. CV Authors: Guojun Lei, Chi Wang, Hong Li, Rong Zhang, Yikai Wang, Weiwei Xu Title: AnimateAnything: Consistent and Controllable Animation for Video Generation Arxiv: http://arxiv.org/abs/2411.10836v1 Abstract: We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions,...

Top-$nσ$: Not All Logits Are You Need 20.11.2024

🤗 Paper Upvotes: 12 | cs. LG Authors: Chenxia Tang, Jianchun Liu, Hongli Xu, Liusheng Huang Title: Top-$nσ$: Not All Logits Are You Need Arxiv: http://arxiv.org/abs/2411.07641v1 Abstract: Large language models (LLMs) typically employ greedy decoding or low-temperature sampling for reasoning tasks, reflecting a perceived trade-off between diversity and accuracy. We challenge this convention by int...

Drowning in Documents: Consequences of Scaling Reranker Inference 20.11.2024

🤗 Paper Upvotes: 10 | cs. IR, cs. CL, cs. LG Authors: Mathew Jacob, Erik Lindgren, Matei Zaharia, Michael Carbin, Omar Khattab, Andrew Drozdov Title: Drowning in Documents: Consequences of Scaling Reranker Inference Arxiv: http://arxiv.org/abs/2411.11767v1 Abstract: Rerankers, typically cross-encoders, are often used to re-score the documents retrieved by cheaper initial IR systems. This is becau...

SlimLM: An Efficient Small Language Model for On-Device Document Assistance 20.11.2024

🤗 Paper Upvotes: 10 | cs. CL Authors: Thang M. Pham, Phat T. Nguyen, Seunghyun Yoon, Viet Dac Lai, Franck Dernoncourt, Trung Bui Title: SlimLM: An Efficient Small Language Model for On-Device Document Assistance Arxiv: http://arxiv.org/abs/2411.09944v1 Abstract: While small language models (SLMs) show promises for mobile deployment, their real-world performance and applications on smartphones rem...

Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts 20.11.2024

🤗 Paper Upvotes: 8 | cs. CV Authors: Jinqiang Long, Yanqi Dai, Guoxing Yang, Hongpeng Lin, Nanyi Fei, Yizhao Gao, Zhiwu Lu Title: Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts Arxiv: http://arxiv.org/abs/2411.10669v1 Abstract: As the research of Multimodal Large Language Models (MLLMs) becomes popular, an advancing MLLM model is typically required to handle variou...

SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers 20.11.2024

🤗 Paper Upvotes: 8 | cs. LG Authors: Joseph Liu, Joshua Geddes, Ziyu Guo, Haomiao Jiang, Mahesh Kumar Nandwana Title: SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers Arxiv: http://arxiv.org/abs/2411.10510v1 Abstract: Diffusion Transformers (DiT) have emerged as powerful generative models for various tasks, including image, video, and speech synthesis. However,...

LLäMmlein: Compact and Competitive German-Only Language Models from Scratch 20.11.2024

🤗 Paper Upvotes: 7 | cs. CL, cs. AI, cs. LG Authors: Jan Pfister, Julia Wunderle, Andreas Hotho Title: LLäMmlein: Compact and Competitive German-Only Language Models from Scratch Arxiv: http://arxiv.org/abs/2411.11171v1 Abstract: We create two German-only decoder models, LL\"aMmlein 120M and 1B, transparently from scratch and publish them, along with the training data, for the German NLP research...

LLaVA-o1: Let Vision Language Models Reason Step-by-Step 19.11.2024

🤗 Paper Upvotes: 64 | cs. CV Authors: Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, Li Yuan Title: LLaVA-o1: Let Vision Language Models Reason Step-by-Step Arxiv: http://arxiv.org/abs/2411.10440v1 Abstract: Large language models have demonstrated substantial advancements in reasoning capabilities, particularly through inference-time scaling, as illustrated by models such as OpenAI's o1. H...

GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation 19.11.2024

🤗 Paper Upvotes: 19 | cs. CV, cs. AI, cs. GR Authors: Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong, Shuai Yang, Bo Dai, Xingang Pan, Chen Change Loy Title: GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation Arxiv: http://arxiv.org/abs/2411.08033v1 Abstract: While 3D content generation has advanced significantly, existing methods still face challenges with input...

Xmodel-1.5: An 1B-scale Multilingual LLM 19.11.2024

🤗 Paper Upvotes: 7 | cs. CL Authors: Wang Qun, Liu Yang, Lin Qingquan, Jiang Ling Title: Xmodel-1.5: An 1B-scale Multilingual LLM Arxiv: http://arxiv.org/abs/2411.10083v1 Abstract: We introduce Xmodel-1.5, a novel 1-billion-parameter multilingual large model pretrained on approximately 2 trillion tokens. The model demonstrates strong performance across several languages, with particularly notable...

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models 16.11.2024

🤗 Paper Upvotes: 32 | cs. LG, cs. AI, cs. CL, cs. CV, 68T05, I.3.5; I.2.10; I.2.6 Authors: Zhengyi Wang, Jonathan Lorraine, Yikai Wang, Hang Su, Jun Zhu, Sanja Fidler, Xiaohui Zeng Title: LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models Arxiv: http://arxiv.org/abs/2411.09595v1 Abstract: This work explores expanding the capabilities of large language models (LLMs) pretrained on text to...

MagicQuill: An Intelligent Interactive Image Editing System 16.11.2024

🤗 Paper Upvotes: 31 | cs. CV Authors: Zichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Wen Wang, Zhiheng Liu, Qifeng Chen, Yujun Shen Title: MagicQuill: An Intelligent Interactive Image Editing System Arxiv: http://arxiv.org/abs/2411.09703v1 Abstract: Image editing involves a variety of complex tasks and requires efficient and precise manipulation techniques. In this paper, we present...

Cut Your Losses in Large-Vocabulary Language Models 16.11.2024

🤗 Paper Upvotes: 15 | cs. LG, cs. CL Authors: Erik Wijmans, Brody Huval, Alexander Hertzberg, Vladlen Koltun, Philipp Krähenbühl Title: Cut Your Losses in Large-Vocabulary Language Models Arxiv: http://arxiv.org/abs/2411.09009v1 Abstract: As language models grow ever larger, so do their vocabularies. This has shifted the memory footprint of LLMs during training disproportionately to one single la...

ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction? 16.11.2024

🤗 Paper Upvotes: 9 | cs. CL Authors: Canyu Chen, Jian Yu, Shan Chen, Che Liu, Zhongwei Wan, Danielle Bitterman, Fei Wang, Kai Shu Title: ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction? Arxiv: http://arxiv.org/abs/2411.06469v1 Abstract: Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text...

Sharingan: Extract User Action Sequence from Desktop Recordings 16.11.2024

🤗 Paper Upvotes: 3 | cs. CV, cs. AI Authors: Yanting Chen, Yi Ren, Xiaoting Qin, Jue Zhang, Kehong Yuan, Lu Han, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, Qi Zhang Title: Sharingan: Extract User Action Sequence from Desktop Recordings Arxiv: http://arxiv.org/abs/2411.08768v1 Abstract: Video recordings of user activities, particularly desktop recordings, offer a rich source of data for underst...

Hermes: A Large Language Model Framework on the Journey to Autonomous Networks 16.11.2024

🤗 Paper Upvotes: 2 | cs. AI, cs. NI Authors: Fadhel Ayed, Ali Maatouk, Nicola Piovesan, Antonio De Domenico, Merouane Debbah, Zhi-Quan Luo Title: Hermes: A Large Language Model Framework on the Journey to Autonomous Networks Arxiv: http://arxiv.org/abs/2411.06490v1 Abstract: The drive toward automating cellular network operations has grown with the increasing complexity of these systems. Despite...

Inconsistencies In Consistency Models: Better ODE Solving Does Not Imply Better Samples 16.11.2024

🤗 Paper Upvotes: 2 | cs. LG, cs. AI Authors: Noël Vouitsis, Rasa Hosseinzadeh, Brendan Leigh Ross, Valentin Villecroze, Satya Krishna Gorti, Jesse C. Cresswell, Gabriel Loaiza-Ganem Title: Inconsistencies In Consistency Models: Better ODE Solving Does Not Imply Better Samples Arxiv: http://arxiv.org/abs/2411.08954v1 Abstract: Although diffusion models can generate remarkably high-quality samples,...

Direct Preference Optimization Using Sparse Feature-Level Constraints 15.11.2024

🤗 Paper Upvotes: 10 | cs. AI, cs. CL Authors: Qingyu Yin, Chak Tou Leong, Hongbo Zhang, Minjun Zhu, Hanqi Yan, Qiang Zhang, Yulan He, Wenjie Li, Jun Wang, Yue Zhang, Linyi Yang Title: Direct Preference Optimization Using Sparse Feature-Level Constraints Arxiv: http://arxiv.org/abs/2411.07618v1 Abstract: The alignment of large language models (LLMs) with human preferences remains a key challenge....

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection 15.11.2024

🤗 Paper Upvotes: 8 | cs. CL Authors: Wissam Antoun, Francis Kulumba, Rian Touchent, Éric de la Clergerie, Benoît Sagot, Djamé Seddah Title: CamemBERT 2.0: A Smarter French Language Model Aged to Perfection Arxiv: http://arxiv.org/abs/2411.08868v1 Abstract: French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models...

Listen to the Daily Paper Cast podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.