Kyle Polich

Data Skeptic

The Data Skeptic Podcast features interviews and discussion of topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of approaches.

Author

Kyle Polich

Category

Technology

Podcast website

dataskeptic.com

Latest episode

Jul 2, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

The Secret and the Global Consciousness Project with Alex Boklin 21.11.2014

I'm joined this week by Alex Boklin to explore the topic of magical thinking especially in the context of Rhonda Byrne's "The Secret", and the similarities it bears to The Global Consciousness Project (GCP). The GCP puts forward the hypothesis that random number generators elicit statistically significant changes as a result of major world events.

[MINI] Monkeys on Typewriters 14.11.2014

What is randomness? How can we determine if some results are randomly generated or not? Why are random numbers important to us in our everyday life? These topics and more are discussed in this mini-episode on random numbers. Many readers will be vaguely familar with the idea of "X number of monkeys banging on Y number of typewriters for Z number of years" - the idea being that such a setup would p...

Mining the Social Web with Matthew Russell 07.11.2014

This week's episode explores the possibilities of extracting novel insights from the many great social web APIs available. Matthew Russell's  Mining the Social Web  is a fantastic exploration of the tools and methods, and we explore a few related topics. One helpful feature of the book is it's use of a  Vagrant  virtual machine. Using it, readers can easily reproduce the examples from the book, an...

[MINI] Is the Internet Secure? 31.10.2014

This episode explores the basis of why we can trust encryption.  Suprisingly, a discussion of looking up a word in the dictionary (binary search) and efficiently going wine tasting (the travelling salesman problem) help introduce computational complexity as well as the P ?= NP question, which is paramount to the trustworthiness RSA encryption. With a high level foundation of computational theory,...

Practicing and Communicating Data Science with Jeff Stanton 24.10.2014

Jeff Stanton  joins me in this episode to discuss his book  An Introduction to Data Science , and some of the unique challenges and issues faced by someone doing applied data science. A challenge to any data scientist is making sure they have a good input data set and apply any necessary data munging steps before their analysis. We cover some good advise for how to approach such problems.

[MINI] The T-Test 17.10.2014

The t-test is this week's mini-episode topic. The t-test is a statistical testing procedure used to determine if the mean of two datasets differs by a statistically significant amount. We discuss how a wine manufacturer might apply a t-test to determine if the sweetness, acidity, or some other property of two separate grape vines might differ in a statistically meaningful way. Check out more detai...

Data Myths with Karl Mamer 10.10.2014

This week I'm joined by Karl Mamer to discuss the data behind three well known urban legends. Did a large blackout in New York and surrounding areas result in a baby boom nine months later? Do subliminal messages affect our behavior? Is placing beer alongside diapers a recipe for generating more revenue than these products in separate locations? Listen as Karl and I explore these claims.

Contest Announcement 08.10.2014

The Data Skeptic Podcast is launching a contest- not one of chance, but one of skill. Listeners are encouraged to put their data science skills to good use, or if all else fails, guess! The contest works as follows. Below is some data about the cumulative number of downloads the podcast has achieved on a few given dates. Your job is to predict the date and time at which the podcast will recieve do...

[MINI] Selection Bias 03.10.2014

A discussion about conducting US presidential election polls helps frame a converation about selection bias.

[MINI] Confidence Intervals 26.09.2014

Commute times and BBQ invites help frame a discussion about the statistical concept of confidence intervals.

[MINI] Value of Information 19.09.2014

A discussion about getting ready in the morning, negotiating a used car purchase, and selecting the best AirBnB place to stay at help frame a conversation about the decision theoretic principal known as the Value of Information equation.

Game Science Dice with Louis Zocchi 17.09.2014

In this bonus episode, guest Louis Zocchi discusses his background in the gaming industry, specifically, how he became a manufacturer of dice designed to produce statistically uniform outcomes.  During the show Louis mentioned a two part video listeners might enjoy:  part 1  and  part 2  can both be found on youtube.  Kyle mentioned a robot capable of unnoticably cheating at Rock Paper Scissors /...

Data Science at ZestFinance with Marick Sinay 12.09.2014

Marick Sinay from ZestFianance is our guest this weel.  This episode explores how data science techniques are applied in the financial world, specifically in assessing credit worthiness.  

[MINI] Decision Tree Learning 05.09.2014

Linhda and Kyle talk about Decision Tree Learning in this miniepisode.  Decision Tree Learning is the algorithmic process of trying to generate an optimal decision tree to properly classify or forecast some future unlabeled element based by following each step in the tree.

Jackson Pollock Authentication Analysis with Kate Jones-Smith 29.08.2014

Our guest this week is  Hamilton physics professor Kate Jones-Smith  who joins us to discuss the evidence for the claim that drip paintings of Jackson Pollock contain fractal patterns. This hypothesis originates in a paper by Taylor, Micolich, and Jonas titled  Fractal analysis of Pollock's drip paintings  which appeared in Nature.  Kate and co-author  Harsh Mathur  wrote a paper titled  Revisitin...

[MINI] Noise!! 22.08.2014

Our topic for this week is "noise" as in signal vs. noise.  This is not a signal processing discussions, but rather a brief introduction to how the work noise is used to describe how much information in a dataset is useless (as opposed to useful). Also, Kyle announces having recently had the pleasure of appearing as a guest on The Conspiracy Skeptic Podcast  to discussion The Bible Code.  Please c...

Guerilla Skepticism on Wikipedia with Susan Gerbic 15.08.2014

Our guest this week is Susan Gerbic. Susan is a skeptical activist involved in many activities, the one we focus on most in this episode is  Guerrilla Skepticism on Wikipedia , an organization working to improve the content and citations of Wikipedia.  During the episode, Kyle recommended Susan's talk a The Amazing Meeting 9 which can be found  here .  Some noteworthy topics mentioned during the p...

[MINI] Ant Colony Optimization 08.08.2014

In this week's mini episode, Linhda and Kyle discuss Ant Colony Optimization - a numerical / stochastic optimization technique which models its search after the process ants employ in using random walks to find a goal (food) and then leaving a pheremone trail in their walk back to the nest.  We even find some way of relating the city of San Francisco and running a restaurant into the discussion.

Data in Healthcare IT with Shahid Shah 01.08.2014

Our guest this week is Shahid Shah. Shahid is CEO at Netspective , and writes three blogs: Health Care Guy , Shahid Shah , and HitSphere - the Healthcare IT Supersite . During the program, Kyle recommended a talk from the 2014 MIT Sloan CIO Symposium entitled Transforming "Digital Silos" to "Digital Care Enterprise" which was hosted by our guest Shahid Shah . In addition to his work in Healthcare...

[MINI] Cross Validation 25.07.2014

This miniepisode discusses the technique called Cross Validation - a process by which one randomly divides up a dataset into numerous small partitions. Next, (typically) one is held out, and the rest are used to train some model. The hold out set can then be used to validate how good the model does at describing/predicting new data.

Streetlight Outage and Crime Rate Analysis with Zach Seeskin 18.07.2014

This episode features a discussion with statistics PhD student Zach Seeskin about a project he was involved in as part of the Eric and Wendy Schmidt Data Science for Social Good Summer Fellowship.  The project involved exploring the relationship (if any) between streetlight outages and crime in the City of Chicago.  We discuss how the data was accessed via the City of Chicago data portal, how the...

[MINI] Experimental Design 11.07.2014

This episode loosely explores the topic of Experimental Design including hypothesis testing, the importance of statistical tests, and an everyday and business example.

The Right (big data) Tool for the Job with Jay Shankar 07.07.2014

In this week's episode, we discuss applied solutions to big data problem with big data engineer Jay Shankar.  The episode explores approaches and design philosophy to solving real world big data business problems, and the exploration of the wide array of tools available.  

[MINI] Bayesian Updating 27.06.2014

In this minisode, we discuss Bayesian Updating - the process by which one can calculate the most likely hypothesis might be true given one's older / prior belief and all new evidence.

Personalized Medicine with Niki Athanasiadou 20.06.2014

In the second full length episode of the podcast, we discuss the current state of personalized medicine and the advancements in genetics that have made it possible.

Listen to the Data Skeptic podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.