Pan Wu

Snacks Weekly on Data Science

This podcast is about making data science and machine learning knowledge accessible and less intimidating. Every week, I will handpick one selected industrial tech blog to break it down. We will discuss some key data science concepts and machine learning algorithms, and how they are applied in those real-world applications. Subscribe to the channel and enjoy Snacks Weekly on Data Science!

Author

Pan Wu

Category

Education

Podcast website

podcasters.spotify.com

Latest episode

Jul 6, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Measure Semantic Relevance in Search with Large Language Models (LLMs) [Faire] 05.08.2024

In this episode, we will discuss why search relevance is important for Faire, how their data team quantifies semantic relevance in search, and how they leverage large language models (LLMs) to measure it efficiently. For more details, you can refer to their published tech blog, linked here for your reference: https://craft.faire.com/fine-tuning-llama3-to-measure-semantic-relevance-in-search-86a7b1...

Making Informed Decisions in A/B Tests with Multiple Metrics [Spotify] 29.07.2024

In this episode, we will touch on the importance of A/B testing in the product decision-making process. We will share the four types of metrics and the decision-making framework used by the Data Science team at Spotify, as well as the necessary statistical adjustments that need to be incorporated into experimentation to ensure a solid statistical foundation. For more details, you can refer to thei...

Improving ETA Predictions with Advanced Deep Learning Architecture [DoorDash] 22.07.2024

In this episode, we will discuss the importance of Estimated Time of Arrival (ETA) for DoorDash and how the company enhanced its machine learning model through three key directions: upgrading from a tree-based model to a deep-learning architecture, adopting a multi-task modeling approach, and leveraging probabilistic models. For more details, you can refer to their published tech blog, linked here...

Forecasting with the balance of art and science [Meta] 15.07.2024

In this episode, we will discuss the intricacy of balancing both the art and science aspects in forecasting. We will explore this through two key aspects: validation of forecasting and integrating product impact into forecasting, where combining the art and science can be crucial to enhancing forecasting performance. For more details, you can refer to their published tech blog, linked here for you...

Optimize Feature Selection with Generic Algorithm [JustEatTakeaway.com/Grubhub] 08.07.2024

In this episode, we will discuss what is feature selection and how JustEatTakeaway.com leverages generic algorithms as one practice to to optimize their feature selection. For more details, you can refer to their published tech blog, linked here for your reference: https://medium.com/justeattakeaway-tech/optimising-feature-selection-with-genetic-algorithms-an-easy-to-use-python-script-dde44cc9c053...

Monitoring Mechanisms for Recommendation Systems [Tubi TV] 01.07.2024

In this episode, we will discuss how to build an effective monitoring mechanism for a recommendation system. We will cover the basic components of various recommendation processes and explore how the Tubi TV Engineering team developed their monitoring flow. For more details, you can refer to their published tech blog, linked here for your reference: https://code.tubitv.com/how-to-monitor-a-recomme...

Determine Causal Effects through Adoptor Analysis [Walmart] 24.06.2024

In this episode, we will introduce the concept of causal inference, and discuss how the data scientist team from Walmart determines Causal Effects when A/B Tests are Infeasible through Adopter Analysis. Based on their published tech blog, with the link provided here for your reference:  https://medium.com/walmartglobaltech/how-to-determine-causal-effects-when-a-b-tests-are-infeasible-through-adopt...

Measure Web Performance with Composite Metric [Indeed] 17.06.2024

In this episode, we will discuss the importance of metrics for the business. We will share the journey of how Indeed started with a single metric to measure client-side performance, and eventually converged into a composite metric serving as a comprehensive measure Based on their published tech blog, with the link provided here for your reference:  https://engineering.indeedblog.com/blog/2024/01/c...

Developing Text-to-SQL Feature with Large Language Models (LLMs) [Pinterest] 10.06.2024

In this episode, we will explore how Pinterest uses generative AI technology to develop a text-to-SQL feature for their data analytics team. We will examine the general architecture of their two iterations, regarding how the team has enhanced the AI product to support better a wider range of analytics use cases with improved productivity. Based on their published tech blog, with the link provided...

A/B Testing with Cluster Experimentation Under Strong Network Effects [Meta] 03.06.2024

In this episode, we'll discuss what network effects are, how they introduce challenges in the standard A/B testing framework, and how the cluster experimentation method can be leveraged to address these challenges. We will also delve into the technical details of how clusters can be generated, and evaluated, and the associated trade-offs that need to be considered. Based on their published tec...

Measuring Marketing Effectiveness with Geo-experimentation [Grammarly] 27.05.2024

In this episode, we'll explore how the data science team from Grammarly developed their geo-experimentation to measure marketing effectiveness. We will cover about three components in designing an A/B testing experiment, as well as considerations regarding the opportunistic costs of the experimentation. Based on their published tech blog, with the link provided here for your reference: https:/...

Monte Carlo Simulatoin for Sampled Success Metrics [Shopify] 20.05.2024

In this episode, we'll explore how the data science team from Shopify leverages Monte Carlo Simulation to develop their sampled success metrics. We'll discuss what is sampled success metrics, the associated trade-offs needed to build them, and how Monte Carlo simulation can be used to inform decisions. Based on their published tech blog, with the link provided here for your reference: http...

Building Generative AI Product for Customer Segmentation [Klaviyo] 13.05.2024

In this episode, we'll explore how the data science team from Klaviyo developed a Generative AI Product that enhances experiences and enables efficient customer segmentation. We'll also discuss two key concepts in Generative AI: prompt chaining and few-shot learning. Based on their published tech blog, with the link provided here for your reference: https://klaviyo.tech/building-segments-a...

Machine Learning Solution for Failed Job Auto Remediation [Netflix] 06.05.2024

Description: In this episode, we will talk about the importance of remediating failed workflow jobs to reduce business infrastructure costs. We delve into Netflix's approach, which involves enhancing their existing rule-based error classifier with advanced machine learning models. This allowed for auto-remediation, improving the handling of memory configuration and unclassified errors, ultimat...

Measure Technical Debt in Software Engineering [Booking.com] 29.04.2024

In this episode, we will talk about what is technical debt in software engineering and its associated risks. We will also share a set of metrics to measure the status of technical debt and ways to help companies quantify their progress toward better software engineering efforts. Based on their published tech blog, with the link provided here for your reference: https://medium.com/booking-com-devel...

Improving Price Experimentation at Amazon [Amazon] 22.04.2024

In this episode, we will discuss how the Science team at Amazon designs pricing experiments to improve experimentation power. We will cover concepts like the carryover effect and spillover effect, as well as the solutions the team developed to overcome those challenges. Based on their published tech blog, with the link provided here for your reference: https://www.amazon.science/blog/the-science-o...

Tackle Position Bias in Uber Eats Feed Recommendation [Uber] 15.04.2024

In this episode, we will talk about what is position bias in recommendation systems, and how the applied scientist team at Uber tackled this challenge. Based on their published tech blog, with the link provided here for your reference: https://www.uber.com/blog/improving-uber-eats-home-feed-recommendations

Decision Making with Analytical Hierarchy Processing [New York Times] 08.04.2024

In this episode, we will take a look at one analytical approach in decision-making called Analytical Hierarchy Processing (AHP). The New York Times team leverages this AHP process to help with making better collective decisions in their technology choices for privacy. Based on their published tech blog, with the link provided here for your reference: https://open.nytimes.com/collective-decision-ma...

Leveraging Generative AI to Boost Data Analyst Productivity [Intuit] 01.04.2024

In this episode, we delve into a study conducted by Intuit on measuring the productivity impact of the Gen AI tool on their data analyst team. It demonstrates exciting positive productivity gains, offering an exciting outlook for many in the industry. Based on their published tech blog, with the link provided here for your reference:  ​​ https://medium.com/intuit-engineering/how-intuit-data-analys...

Perturbation analysis of Large Language Models (LLM) [Microsoft] 25.03.2024

Description: In this episode, we discuss how the data science team at Microsoft designed perturbation analysis to understand the Large Language Model’s performance on commonly seen tasks. Based on their published tech blog, with the link provided here for your reference:  https://medium.com/data-science-at-microsoft/perturbation-analysis-and-llms-how-sensitive-are-llms-to-their-input-91a8407a971f

Monte Carlo Simulation to Predict Tennis Game Outcomes [DraftKings] 18.03.2024

In this episode, we discuss how DraftKings leverages Monte Carlo simulation to predict tennis game outcomes and generate probabilities to power its online gambling service. Based on their published tech blog, with the link provided here for your reference:  https://medium.com/draftkings-engineering/building-a-tennis-simulation-d6afdaa97d19  

Two-Tower Neural Network Architecture for Candidate Generation in Recommendation System [Expedia] 11.03.2024

In this episode, we discuss the insights shared by Expedia's Machine Learning Engineering team on how they leverage the two-tower neural network architecture in the candidate generation stage of their recommendation system. Based on their published tech blog, with the link provided here for your reference: https://medium.com/expedia-group-tech/candidate-generation-using-a-two-tower-approach-wi...

Large Language Model (LLM) with Retrieval Augmented Generation (RAG) Technology for Efficient Agile Planning [Walmart] 04.03.2024

In this episode, we discuss how the Engineering team at Walmart created a customized Large Language Model (LLM) agent with Retrieval Augmented Generation (RAG) Technology to improve their agile planning efficiency. Based on their published tech blog, with the link provided here for your reference: https://medium.com/walmartglobaltech/an-autonomous-agent-for-agile-planning-98303e194e08

Automated Sanity Checks to Streamline Machine Learning Deployment [Intuit] 26.02.2024

In this episode, we discuss why and how Intuit leverages automated sanity checks to streamline its machine learning development process. Based on their published tech blog, with the link provided here for your reference: https://medium.com/intuit-engineering/how-to-streamline-ml-model-deployment-automated-sanity-checks-64a23166fdc5

Demand Forecasting with Machine Learning Models [Picnic] 19.02.2024

In this episode, we take a look at the learnings from an online grocery tech company regarding their development of machine learning models for scalable demand forecasting. Based on the tech blog from Picnic International, the link is provided here for your reference: https://blog.picnic.nl/running-demand-forecasting-machine-learning-models-at-scale-bd058c9d4aa7

Listen to the Snacks Weekly on Data Science podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.