Pan Wu

Snacks Weekly on Data Science

This podcast is about making data science and machine learning knowledge accessible and less intimidating. Every week, I will handpick one selected industrial tech blog to break it down. We will discuss some key data science concepts and machine learning algorithms, and how they are applied in those real-world applications. Subscribe to the channel and enjoy Snacks Weekly on Data Science!

Author

Pan Wu

Category

Education

Podcast website

podcasters.spotify.com

Latest episode

Jul 6, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Enhance Customer Ticket Categorization with Generative AI [RazorPay] 27.01.2025

In this episode, we’ll explore RazorPay, its business, and the critical role customer service plays in its success. We'll dive into how RazorPay revolutionized customer ticket categorization using generative AI. By replacing customer-selected categories with an AI-driven system, they enabled automatic interpretation of ticket details. This approach incorporated pre-processing, prompt engineeri...

A Brief Overview of Causal Analysis [Microsoft] 20.01.2025

In this episode, we will explore the foundational concepts of causal analysis, focusing on its two main pillars: causal discovery and causal inference. We will discuss the types of questions these pillars aim to answer and provide illustrations of related methodologies to better clarify their concepts. For more details, you can refer to their published tech blog, linked here for your reference:  h...

Develop a chatbot using customized generative AI solution [Noom] 13.01.2025

In this episode, we will explore how Noom developed a customized AI assistant solution using generative AI models. By incorporating key components such as prompt engineering, dynamic personalization, and integration with their knowledge base, the team created a smarter, more reliable assistant to support their customers. This approach offers valuable insights for anyone looking to build tailored g...

Personal Data Classification [Airbnb] 06.01.2025

In this episode, we will explore Airbnb’s approach to personal data classification. We will begin by introducing the importance of protecting personal data. Next, we will examine the technical solution, where the three pillars—Catalog, Detection, and Reconciliation—form the backbone of their workflow. Finally, we will discuss how performance metrics are implemented to ensure the system remains rel...

Improving Search Engine Marketing Performance: My First Data Science Project 30.12.2024

In this New Year’s episode, I reflect on one of my earliest data science projects. My son will “interview” me about my experiences during that time. The key project discussed in this episode is also featured in a LinkedIn article I wrote a few years ago (⁠ https://www.linkedin.com/pulse/my-first-data-science-project-pan-wu/⁠⁠ ). I’d love for you to check it out, leave a 5-star review, and subscrib...

Reward engineering for better content recommendation [Netflix] 23.12.2024

In this episode, we will explore Netflix’s approach to content recommendation using contextual bandits and reward engineering. We will also discuss the important role of proxy reward functions and how Netflix leverages offline machine learning models to predict delayed customer feedback, enabling them to continuously improve their recommendation engine and deliver a more personalized viewing exper...

Adaptive Experimentation for Paid Marketing Optimization [Instacart] 16.12.2024

In this episode, we will introduce the concept of paid marketing. We’ll explore how Instacart’s team developed an adaptive experimentation framework that continuously balances exploration and exploitation, ultimately maximizing marketing efficiency. For more details, you can refer to their published tech blog, linked here for your reference:  https://tech.instacart.com/bandits-for-marketing-optimi...

Product bundle recommendation with Graph Learning and GPT [CVS Health] 09.12.2024

In this episode, we will introduce what CVS Health is and the importance of product recommendations for their business needs. We will delve into how their data science team leveraged advanced technologies, including Graph Neural Networks and generative AI models like GPT-4, to develop a prototype system to make product bundle recommendations. For more details, you can refer to their published tech...

Augmentation techniques for imbalanced text classification [Walmart] 02.12.2024

In this episode, we will introduce the issue of data imbalance and its impact on machine learning models, especially for text data. We will discuss a range of augmentation techniques and walked through how Walmart’s data science team built an automated augmentation module to apply these techniques consistently and effectively. For more details, you can refer to their published tech blog, linked he...

Optimize delivery picking process with mathematical modeling [Instamart] 25.11.2024

In this episode, we will discuss what is Instamart, its business model, and the need to optimize operational efficiency continuously. We will explore how the team tackled picker assignment issues, tested a multi-order batching solution, and applied modeling techniques to enhance Instamart’s speed and service quality. For more details, you can refer to their published tech blog, linked here for you...

Building Contextualised Moderation Classifier [GovTech Singapore] 18.11.2024

In this episode, we introduce GovTech Singapore and its reasons for tackling the content moderation problem. We discuss their innovative approach to building the moderation classifier, which involves using a consensus voting mechanism with existing commercial LLMs to improve labeling in the training dataset, providing a strong foundation for developing the machine learning model. For more details,...

Promotion aware demand forecasting for groceries [AFresh] 11.11.2024

In this episode, we will introduce Afresh, explore the concept of demand forecasting, and discuss the critical role of promotions in grocery forecasting. We will also examine Afresh’s technical solution, which combines deep learning with a carefully curated set of data inputs to create an accurate, promotion-aware demand forecasting system. For more details, you can refer to their published tech b...

Graph technology in fraud detection and prevention [Booking.com] 04.11.2024

In this episode, we will explore the business model of Booking.com and the unique challenges it faces in preventing fraud. We will discuss how graph technology can enhance fraud detection by representing data through relationships and how combining this with machine learning enables Booking.com to detect and stop fraudulent activity in real-time. For more details, you can refer to their published...

Marketing mix modeling in marketing Measurement [Qonto] 28.10.2024

In this episode, we will introduce a key marketing challenge—attribution—and explore how marketing mix modeling (MMM) can help solve it. We will also discuss how MMM works in practice and examined the trade-offs between using consultants for MMM development versus building an in-house solution. For more details, you can refer to their published tech blog, linked here for your reference:  https://m...

Causal machine learning to power data driven decisions [Urban Company] 21.10.2024

In this episode, we will explore Urban Company’s business needs and the role of causal machine learning in addressing them. We will delve into three key models—S-learner, T-learner, and X-learner—and used a simple example to illustrate how each one works. These models provide valuable insights by offering more accurate estimates of cause-and-effect relationships, helping companies make better data...

Advanced Product Categorization with Vision Language Models [Faire] 14.10.2024

In this episode, we will explore how Faire tackled the challenge of product categorization. They initially used the K-nearest neighbor algorithm with CLIP embeddings, which improved categorization but still required manual corrections. To further enhance accuracy, the team fine-tuned a vision-language model using their in-house dataset, increasing accuracy significantly. This solution showcases ho...

Leverage CUPED to reduce experimentation lifecycle [Walmart] 07.10.2024

In this episode, we will discuss why Walmart relies on online experimentation to drive data-driven decisions and how reducing variance is a crucial challenge to making these experiments more efficient. We will also introduce the CUPED methodology, explaining how Walmart leverages it to speed up its experimentation process, enabling faster, more reliable insights for continuous improvement. For mor...

Building video classifiers with vision language models and active learning [Netflix] 30.09.2024

In this episode, we will explore the challenge Netflix faces in building machine learning models for video understanding. We will examine Netflix’s solution—a self-service system with active learning that empowers video experts to participate in creating and refining machine learning classifiers through a streamlined, three-step process. For more details, you can refer to their published tech blog...

Measuring Marketing Incrementality with Geo Testing [Expedia] 23.09.2024

In this episode, we explore the concept of incrementality in marketing, and how the data science team at Expedia leverages geo-testing to successfully measure their marketing campaign’s  incrementality  For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/expedia-group-tech/measuring-marketing-success-the-power-of-incrementality-and-geo-t...

Personalized Out-of-App Marketing Strategy [Uber] 16.09.2024

In this episode, we explore the concept of out-of-app marketing and its importance for businesses like Uber. We'll discuss the challenges involved in personalizing out-of-app marketing messages, the essential components of their recommendation architecture, and how the team created customized solutions to improve personalization and relevance in their recommendations. For more details, you can...

Predicting Estimated Time of Arrival (ETA) Reliability [Lyft] 09.09.2024

In this episode, we will discuss the importance of ETA for ridesharing apps and the challenges of providing a reliable ETA to users upfront. We delved into the practices of the machine learning team at Lyft, examining how they developed a solution to address this unique challenge using a lightweight machine learning model. For more details, you can refer to their published tech blog, linked here f...

Moderating Inappropriate Video Content [Yelp] 02.09.2024

In this episode, we will explore how Yelp navigates the challenges of incorporating video reviews into its platform. We will discuss the use of machine learning to detect inappropriate content and the strategies to maintain the quality and integrity of its platform. For more details, you can refer to their published tech blog, linked here for your reference: https://engineeringblog.yelp.com/2024/0...

Measuring brand perception with social media data and deep learning [Airbnb] 26.08.2024

In this episode, we will introduce what a brand is and why measuring brand perception is crucial. We will discuss how social media data can be used to achieve this and look at how Airbnb leverages deep learning methods and word embedding technologies to quantify brand perception more effectively. For more details, you can refer to their published tech blog, linked here for your reference: https://...

Empower Decision Making with Regression Discontinuity Design [Instacart] 19.08.2024

In this episode, we will explore the concept of quasi-experimentation and its role in Instacart's decision-making process. We will take a closer look at one specific methodology, regression discontinuity design, explaining its key concepts and demonstrating how it works through an interesting example. For more details, you can refer to their published tech blog, linked here for your reference:...

Product Recommendation with Deep Learning and Reinforcement Learning [LinkedIn] 12.08.2024

In this episode, we will discuss the machine learning architecture built by LinkedIn for their premium product recommendation. We will explore the machine learning architecture, which includes a two-towered neural network and reinforcement learning as key components. For more details, you can refer to their published tech blog, linked here for your reference: https://www.linkedin.com/blog/engineer...

Listen to the Snacks Weekly on Data Science podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.