Kyle Polich
Data Skeptic
The Data Skeptic Podcast features interviews and discussion of topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of approaches.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[MINI] Auto-correlative functions and correlograms 22.04.2016 14:58
When working with time series data, there are a number of important diagnostics one should consider to help understand more about the data. The auto-correlative function, plotted as a correlogram, helps explain how a given observations relates to recent preceding observations. A very random process (like lottery numbers) would show very low values, while temperature (our topic in this episode) doe...
Early Identification of Violent Criminal Gang Members 15.04.2016 27:05
This week I spoke with Elham Shaabani and Paulo Shakarian ( @PauloShakASU ) about their recent paper Early Identification of Violent Criminal Gang Members (also available on arXiv ). In this paper, they use social network analysis techniques and machine learning to provide early detection of known criminal offenders who are in a high risk group for committing violent crimes in the future. Their te...
[MINI] Fractional Factorial Design 08.04.2016 11:09
A dinner party at Data Skeptic HQ helps teach the uses of fractional factorial design for studying 2-way interactions.
Machine Learning Done Wrong 01.04.2016 25:21
Cheng-tao Chu ( @chengtao_chu ) joins us this week to discuss his perspective on common mistakes and pitfalls that are made when doing machine learning. This episode is filled with sage advice for beginners and intermediate users of machine learning, and possibly some good reminders for experts as well. Our discussion parallels his recent blog post Machine Learning Done Wrong . Cheng-tao Chu is an...
Potholes 25.03.2016 41:22
Co-host Linh Da was in a biking accident after hitting a pothole. She sustained an injury that required stitches. This is the story of our quest to file a 311 complaint and track it through the City of Los Angeles's open data portal. My guests this episode are Chelsea Ursaner (LA City Open Data Team), Ben Berkowitz (CEO and founder of SeeClickFix), and Russ Klettke (Editor of pothole.info)
[MINI] The Elbow Method 18.03.2016 15:14
Certain data mining algorithms (including k-means clustering and k-nearest neighbors) require a user defined parameter k. A user of these algorithms is required to select this value, which raises the questions: what is the "best" value of k that one should select to solve their problem? This mini-episode explores the appropriate value of k to use when trying to estimate the cost of a house in Los...
Too Good to be True 11.03.2016 35:11
Today on Data Skeptic, Lachlan Gunn joins us to discuss his recent paper Too Good to be True . This paper highlights a somewhat paradoxical / counterintuitive fact about how unanimity is unexpected in cases where perfect measurements cannot be taken. With large enough data, some amount of error is expected. The "Too Good to be True" paper highlights three interesting examples which we discuss in t...
[MINI] R-squared 04.03.2016 13:20
How well does your model explain your data? R-squared is a useful statistic for answering this question. In this episode we explore how it applies to the problem of valuing a house. Aspects like the number of bedrooms go a long way in explaining why different houses have different prices. There's some amount of variance that can be explained by a model, and some amount that cannot be directly meas...
Models of Mental Simulation 26.02.2016 39:44
Jessica Hamrick joins us this week to discuss her work studying mental simulation. Her research combines machine learning approaches iwth behavioral method from cognitive science to help explain how people reason and predict outcomes. Her recent paper Think again? The amount of mental simulation tracks uncertainty in the outcome is the focus of our conversation in this episode. Lastly, Kyle in...
[MINI] Multiple Regression 19.02.2016 18:29
This episode is a discussion of multiple regression: the use of observations that are a vector of values to predict a response variable. For this episode, we consider how features of a home such as the number of bedrooms, number of bathrooms, and square footage can predict the sale price. Unlike a typical episode of Data Skeptic, these show notes are not just supporting material, but are actually...
Scientific Studies of People's Relationship to Music 12.02.2016 42:14
Samuel Mehr joins us this week to share his perspective on why people are musical, where music comes from, and why it works the way it does. We discuss a number of empirical studies related to music and musical cognition, and dispense a few myths about music along the way. Some of Sam's work discussed in this episode include Music in the Home: New Evidence for an Intergenerational Link , Two rando...
[MINI] k-d trees 05.02.2016 14:11
This episode reviews the concept of k-d trees: an efficient data structure for holding multidimensional objects. Kyle gives Linhda a dictionary and asks her to look up words as a way of introducing the concept of binary search. We actually spend most of the episode talking about binary search before getting into k-d trees, but this is a necessary prerequisite.
Auditing Algorithms 29.01.2016 42:58
Algorithms are pervasive in our society and make thousands of automated decisions on our behalf every day. The possibility of digital discrimination is a very real threat, and it is very plausible for discrimination to occur accidentally (i.e. outside the intent of the system designers and programmers). Christian Sandvig joins us in this episode to talk about his work and the concept of auditing a...
[MINI] The Bonferroni Correction 22.01.2016 14:29
Today's episode begins by asking how many left handed employees we should expect to be at a company before anyone should claim left handedness discrimination. If not lefties, let's consider eye color, hair color, favorite ska band, most recent grocery store used, and any number of characteristics could be studied to look for deviations from the norm in a company. When multiple comparisons are to b...
Detecting Pseudo-profound BS 15.01.2016 37:37
A recent paper in the journal of Judgment and Decision Making titled On the reception and detection of pseudo-profound bullshit explores empirical questions around a reader's ability to detect statements which may sound profound but are actually a collection of buzzwords that fail to contain adequate meaning or truth. These statements are definitively different from lies and nonesense, as we discu...
[MINI] Gradient Descent 08.01.2016 14:51
Today's mini episode discusses the widely known optimization algorithm gradient descent in the context of hiking in a foggy hillside.
Let's Kill the Word Cloud 01.01.2016 15:03
This episode is a discussion of data visualization and a proposed New Year's resolution for Data Skeptic listeners. Let's kill the word cloud.
2015 Holiday Special 25.12.2015 14:22
Today's episode is a reading of Isaac Asimov's The Machine that Won the War . I can't think of a story that's more appropriate for Data Skeptic.
Wikipedia Revision Scoring as a Service 18.12.2015 42:56
In this interview with Aaron Halfaker of the Wikimedia Foundation, we discuss his research and career related to the study of Wikipedia. In his paper The Rise and Decline of an open Collaboration Community , he highlights a trend in the declining rate of active editors on Wikipedia which began in 2007. I asked Aaron about a variety of possible hypotheses for the phenomenon, in particular, how auto...
[MINI] Term Frequency - Inverse Document Frequency 11.12.2015 10:17
Today's topic is term frequency inverse document frequency, which is a statistic for estimating the importance of words and phrases in a set of documents.
The Hunt for Vulcan 04.12.2015 41:31
Early astronomers could see several of the planets with the naked eye. The invention of the telescope allowed for further understanding of our solar system. The work of Isaac Newton allowed later scientists to accurately predict Neptune, which was later observationally confirmed exactly where predicted. It seemed only natural that a similar unknown body might explain anomalies in the orbit of Merc...
[MINI] The Accuracy Paradox 27.11.2015 17:04
Today's episode discusses the accuracy paradox. There are cases when one might prefer a less accurate model because it yields more predictive power or better captures the underlying causal factors describing the outcome variable you are interested in. This is especially relevant in machine learning when trying to predict rare events. We discuss how the accuracy paradox might apply if you were tryi...
Neuroscience from a Data Scientist's Perspective 20.11.2015 40:18
... or should this have been called data science from a neuroscientist's perspective? Either way, I'm sure you'll enjoy this discussion with Laurie Skelly . Laurie earned a PhD in Integrative Neuroscience from the Department of Psychology at the University of Chicago. In her life as a social neuroscientist, using fMRI to study the neural processes behind empathy and psychopathy, she learned the ro...
[MINI] Bias Variance Tradeoff 13.11.2015 13:35
A discussion of the expected number of cars at a stoplight frames today's discussion of the bias variance tradeoff. The central ideal of this concept relates to model complexity. A very simple model will likely generalize well from training to testing data, but will have a very high variance since it's simplicity can prevent it from capturing the relationship between the covariates and the output....
Big Data Doesn't Exist 06.11.2015 32:28
The recent opinion piece Big Data Doesn't Exist on Tech Crunch by Slater Victoroff is an interesting discussion about the usefulness of data both big and small. Slater joins me this episode to discuss and expand on this discussion. Slater Victoroff is CEO of indico Data Solutions, a company whose services turn raw text and image data into human insight. He, and his co-founders, studied at Olin Col...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.