Pan Wu

Snacks Weekly on Data Science

This podcast is about making data science and machine learning knowledge accessible and less intimidating. Every week, I will handpick one selected industrial tech blog to break it down. We will discuss some key data science concepts and machine learning algorithms, and how they are applied in those real-world applications. Subscribe to the channel and enjoy Snacks Weekly on Data Science!

Author

Pan Wu

Category

Education

Podcast website

podcasters.spotify.com

Latest episode

Jul 6, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Causal Inference with Double Machine Learning [Microsoft] 21.07.2025

In this episode, we explore how causal inference helps companies like Microsoft answer high‑stakes product and business questions when A/B testing isn’t possible. We dive into Double Machine Learning—a technique that leverages ML models to control for confounding variables and isolate true causal effects. The result is a flexible, rigorous framework that every data scientist should have in their t...

Scalable and Blendable Feed Construction [Whatnot] 14.07.2025

In this episode, we explore how Whatnot tackled the challenge of scaling feed recommendation systems across a rapidly growing platform. We dive into WhataMix—a DAG-based framework that enables teams to build, test, and deploy feed logic using reusable, modular components. It’s a great example of how thoughtful system design can accelerate development while maintaining high standards in machine lea...

Using Generative and Traditional AI to Enhance Travel Experience [Expedia] 07.07.2025

In this episode, we explore how Expedia is integrating both generative and traditional AI to enhance the travel experience. The company’s approach leverages generative models for open-ended, natural language tasks, and relies on traditional models for structured, mission-critical problems. By playing to the strengths of each, Expedia is able to build smarter, more adaptable AI systems without over...

Ensuring Data Quality at Petabyte Scale [Glassdoor] 30.06.2025

In this episode, we dive into how Glassdoor addresses the challenge of maintaining data quality at a petabyte scale. By treating data as a product, the engineering team built a centralized, scalable platform that enables proactive validation, continuous monitoring, and cross-team collaboration. From data contracts and static code analysis to LLM-based logic checks and anomaly detection, we unpack...

Building a Travel Assistant with LLMs [Agoda] 23.06.2025

In this episode, we explore how Agoda used large language models (LLMs) to improve user experience through building a conversational AI product. By focusing on prompt engineering, grounding data, and smart evaluation, the team built a scalable assistant that adds real value to the user journey. For more details, you can refer to their published tech blog, linked here for your reference: https://me...

Setting Goals at Scale with the Goal Map [Meta] 16.06.2025

In this episode, we explore how Meta tackles the complex challenge of setting aligned, measurable, and high-impact goals across a vast organization. Whether you’re in data science, analytics, or product leadership, this episode offers practical insights into building a more effective goal-setting system. For more details, you can refer to their published tech blog, linked here for your reference:...

Predicting user actions with transformer-based models [Hike] 09.06.2025

In this episode, we will explore how Hike applied transformer-based models to predict user behavior in their Rush Gaming Universe. We will look at the business motivation and break down the technical solution, from input features to prediction and evaluation. This case is a good example of how modern deep learning techniques can drive real impact in improving user experience. For more details, you...

Quantization Techniques for Language Model [EsperantoTech] 02.06.2025

In this episode, we will explore quantization techniques for language models. We will look at the business motivation—making large language models more efficient—and unpack the technical solutions that make this possible.  For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/@EsperantoTech/quantization-and-mixed-mode-techniques-for-small-...

Key Ingredients for a Secure Agentic AI Future [Intuit] 26.05.2025

In this episode, we’ll explore the unique security challenges posed by agentic AI systems and why embedding trust and safety into these systems from the ground up is critical. We’ll review a few key ingredients for building a secure agentic AI future. For more details, you can refer to the blog, linked here for your reference: https://medium.com/intuit-engineering/owasp-dishes-out-key-ingredients-...

Lessons and Best Practices in Online Experimentation [Oda] 19.05.2025

In this episode, we explore how Oda scaled its A/B testing practices alongside its business growth, focusing not only on building a technical platform but also on creating a culture that supports high-quality, reliable experimentation. For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/oda-product-tech/odas-online-experimentation-journe...

Predict Booking Cancellations with Survival Modeling [Booking.com] 12.05.2025

In this episode, we explore how Booking.com tackled the challenge of predicting reservation cancellations in an ever-changing travel landscape. By shifting from a traditional classification model to a survival modeling approach, the team developed more time-sensitive and flexible predictions that better support their business needs and decision-making. For more details, you can refer to their publ...

Measure the unit cost of GenAI Features [Workday] 05.05.2025

In this episode, we will explore how Workday tackle the challenge of measuring the cost of GenAI features. We looked at why LLM-powered features require a new approach to cost tracking, and how the team engineered a telemetry-driven system to make those costs visible, actionable, and fair. For more details, you can refer to their published tech blog, linked here for your reference: https://medium....

Adopt recommendation for property search [Expedia] 28.04.2025

In this episode, we will discuss how Expedia’s recommendation system is designed to handle both standard destination searches and property-specific searches. While traditional ranking models optimize for broad search behavior, Expedia’s team refines their learning-to-rank approach by integrating property similarity, ensuring travelers get recommendations that align with their intent. For more deta...

Evaluate LLM-based chatbots performance [Microsoft] 21.04.2025

In this episode, we will explore why evaluating LLM-based chatbots is critical for businesses, the limitations of traditional evaluation methods, and what could be a good robust evaluation framework covering both search performance and LLM-specific metrics.  For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/data-science-at-microsoft/ev...

Algorithmic Content Recommendation with Editorial Judgment [NYTimes] 14.04.2025

In this episode, we will explore how The New York Times balances algorithmic recommendations with editorial judgment. We discuss their business challenge, examine their hybrid content recommendation system, and look at refinements designed to improve the reader experience. For more details, you can refer to their published tech blog, linked here for your reference: https://open.nytimes.com/how-the...

Enhancing Conversational AI with LLMs [Airbnb] 07.04.2025

In this episode, we will explore how Airbnb upgraded its conversational AI system, leveraging LLMs in a controlled and predictable way. We will first examine their business needs, highlighting why traditional chatbot-based workflows were no longer sufficient. Then, we will break down their technical solution, which combines structured workflows with AI-powered reasoning, context management, and a...

Emerging Economy of Large Language Models (LLMs) [Wix] 31.03.2025

In this episode, we will explore the importance of the Large Language Model (LLM) and the forces shaping the LLM economy: competition among AI giants, GPU scarcity, and tokens as the new currency. These dynamics drive innovation and challenge businesses to optimize resources and costs strategically. For more details, you can refer to their published tech blog, linked here for your reference: https...

Global Holdout Groups [Klaviyo] 24.03.2025

In this episode, we will explore why Klaviyo developed its global holdout group feature and how its engineering team overcame the technical challenges. This feature helps Klaviyo’s customers run fair and unbiased experiments across multiple marketing channels, ultimately enhancing the accuracy of their marketing performance insights. For more details, you can refer to their published tech blog, li...

Scaling Code Reviews with LLMs [Faire] 17.03.2025

In this episode, we will explore why code reviews are critical for a fast-growing marketplace like Faire and the challenges that come with scaling them manually as the engineering team expands. We’ll dive into how Large Language Models (LLMs) offer a game-changing solution—automating code reviews by providing instant, context-aware feedback, enforcing coding best practices, and integrating seamles...

Estimating Long-Run Treatment Effects Using Surrogate Indices [Instacart] 10.03.2025

In this episode, we will explore how Instacart uses data science to optimize its incentive promotions. We will discuss the business challenge, introduce the concept of surrogate indices, and walk through the step-by-step process of building and applying one. For more details, you can refer to their published tech blog, linked here for your reference: https://tech.instacart.com/instacarts-economics...

Enabling ML Productivity and Efficiency at Scale [Meta] 03.03.2025

In this episode, we will explore Meta’s AI scaling challenges and how the company leverages productivity and efficiency to optimize its computing power. We also discuss how analytics insights help identify active levers to improve AI development. For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/@AnalyticsAtMeta/innovation-demands-comp...

Elevating Product Machine Learning Models with LLMs [Coupang] 24.02.2025

In this episode, we will explore how Coupang integrates Large Language Models (LLMs) to enhance its machine learning ecosystem. We'll break down Coupang’s business model, key machine learning categories, and the role of Foundation Models in improving efficiency and accuracy. Additionally, we'll walk through the LLM development lifecycle, discuss critical infrastructure decisions, and exami...

Ranking Lodgings: Machine Learning Behind the Booking Experience [Expedia] 17.02.2025

In this episode, we will explore how Expedia ranks lodging options to optimize both customer experience and business objectives. We will discuss the business problem—how ranking impacts Expedia’s success and the challenges of hotel recommendations. Then, we will break down the data science solution, covering the objective functions, features/signals, machine learning architectures, and the evaluat...

Machine Learning Solution to Personalize Recommendations [Thumbtack] 10.02.2025

In this episode, we will explore Thumbtack’s business model and the importance of recommendations in helping professionals grow their businesses. We will share the team’s machine learning solution architecture, which involved building two sub-models and deploying them using offline inference to meet the business needs cost-efficiently. This quick and nimble approach to machine learning demonstrate...

Enhancing Machine Learning model quality with model excellence score framework [Uber] 03.02.2025

In this episode, we will explore Uber’s Model Excellence Scores (MES) framework, a robust system designed to maintain and enhance the quality of machine learning models at scale. We will unpack its core components—indicators, objectives, and agreements—and explain how they work together to ensure model reliability and performance. This framework enables Uber’s ML ecosystem to operate seamlessly an...

Listen to the Snacks Weekly on Data Science podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.