Jack Waudby
Disseminate: The Computer Science Research Podcast
This podcast features interviews with Computer Science researchers. Hosted by Dr. Jack Waudby researchers are interviewed, highlighting the problem(s) they tackled, solutions they developed, and how their findings can be applied in practice. This podcast is for industry practitioners, researchers, and students, aims to further narrow the gap between research and practice, and to generally make awesome Computer Science research more accessible. We have 2 types of episode: (i) Cutting Edge (red/blue logo) where we talk to researchers about their latest work, and (ii) High Impact (gold/silver log...
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Mateusz Gienieczko | AnyBlox: A Framework for Self-Decoding Datasets | #69 17.03.2026 1:02:28
In this episode of Disseminate: The Computer Science Research Podcast, host Dr. Jack Waudby is joined by Mateusz Gienieczko, PhD researcher at TU Munich and co-author of the VLDB Best Paper Award winning paper AnyBlox. They dive deep into a fundamental problem in modern data systems: why cutting-edge data encodings and file formats rarely make it from research into real-world systems — and how Any...
Xiangyao Yu | Disaggregation: A New Architecture for Cloud Databases | #68 27.11.2025 42:12
In this episode of Disseminate: The Computer Science Research Podcast , host Jack Waudby sits down with Xiangyao Yu (UW–Madison), one of the leading voices shaping the next generation of cloud-native databases. We dive deep into disaggregation — the architectural shift transforming how modern data systems are built. Xiangyao breaks down: Why traditional shared-nothing databases struggle in cloud e...
Navid Eslami | Diva: Dynamic Range Filter for Var-Length Keys and Queries | #67 13.11.2025 46:50
In this episode of Disseminate: The Computer Science Research Podcast , Jack sits down with Navid Eslami , PhD researcher at the University of Toronto , to discuss his award-winning paper “DIVA: Dynamic Range Filter for Variable Length Keys and Queries” , which earned Best Research Paper at VLDB . Navid breaks down how range filters extend the power of traditional filters for modern databases and...
Adaptive Factorization in DuckDB with Paul Groß 06.11.2025 51:15
In this episode of the DuckDB in Research series , host Jack Waudby sits down with Paul Groß, PhD student at CWI Amsterdam, to explore his work on adaptive factorization and worst-case optimal joins - techniques that push the boundaries of analytical query performance. Paul shares insights from his CIDR'25 paper “Adaptive Factorization Using Linear Chained Hash Tables” , revealing how decades of d...
Parachute: Rethinking Query Execution and Bidirectional Information Flow in DuckDB - with Mihail Stoian 30.10.2025 36:34
In this episode of the DuckDB in Research series, host Jack Waudby sits down with Mihail Stoian , PhD student at the Data Systems Lab, University of Technology Nuremberg , to unpack the cutting-edge ideas behind Parachute , a new approach to robust query processing and bidirectional information passing in modern analytical databases. We explore how Parachute bridges theory and practice , combining...
Anarchy in the Database: Abigale Kim on DuckDB and DBMS Extensibility 23.10.2025 46:24
In this episode of the DuckDB in Research series , host Jack Waudby talks with Abigale Kim , PhD student at the University of Wisconsin–Madison and author of VLDB 2025 paper : “Anarchy in the Database: A Survey and Evaluation of DBMS Extensibility”. They explore how database extensibility is reshaping modern data systems — and why DuckDB is emerging as the gold standard for safe, flexible, and hig...
Recursive CTEs, Trampolines, and Teaching Databases with DuckDB - with Prof. Torsten Grust 16.10.2025 51:05
In this episode of the DuckDB in Research series , host Dr Jack Waudby talks with Professor Torsten Grust from the University of Tübingen. Torsten is one of the pioneers behind DuckDB’s implementation of recursive CTEs. In the episode they unpack: The power of recursive CTEs and how they turn SQL into a full-fledged programming language. The story behind adding recursion to DuckDB , including the...
DuckDB in Research S2 Coming Soon! 16.10.2025 2:06
Hey folks! The DuckDB in Research series is back for S2! In this season we chat with: Torsten Grust: Recursive CTEs Abigale Kim: Anarchy in the Database Mihail Stoian: Parachute: Single-Pass Bi-Directional Information Passing Paul Gross: Adaptive Factorization Using Linear-Chained Hash Tables Whether you're a researcher, engineer, or just curious about the intersection of databases and innovation...
Rohan Padhye & Ao Li | Fray: An Efficient General-Purpose Concurrency JVM Testing Platform | #66 06.10.2025 58:45
In this episode of Disseminate: The Computer Science Research Podcast, guest host Bogdan Stoica sits down with Ao Li and Rohan Padhye (Carnegie Mellon University) to discuss their OOPSLA 2025 paper: "Fray: An Efficient General-Purpose Concurrency Testing Platform for the JVM". We dive into: Why concurrency bugs remain so hard to catch -- even in "well-tested" Java projects. The design of Fray, a n...
Shrey Tiwari | It's About Time: A Study of Date and Time Bugs in Python Software | #65 23.09.2025 1:05:29
In this episode, Bogdan Stoica, Postdoctoral Research Associate in the SysNet group at the University of Illinois Urbana-Champaign (UIUC) steps in to guest host. Bogdan sits down with Shrey Tiwari, a PhD student in the Software and Societal Systems Department at Carnegie Mellon University and member of the PASTA Lab, advised by Prof. Rohan Padhye. Together, they dive into Shrey’s award-winning res...
Lessons Learned from Five Years of Artifact Evaluations at EuroSys | #64 30.07.2025 43:48
In this episode we are joined by Thaleia Doudali, Miguel Matos, and Anjo Vahldiek-Oberwagner to delve into five years of experience managing artifact evaluation at the EuroSys conference. They explain the goals and mechanics of artifact evaluation, a voluntary process that encourages reproducibility and reusability in computer systems research by assessing the supporting code, data, and documentat...
Dominik Winterer | Validating SMT Solvers for Correctness and Performance via Grammar-based Enumeration | #63 25.07.2025 43:38
In this episode of the Disseminate podcast, Dominik Winterer discusses his research on SMT (Satisfiability Modulo Theories) solvers and his recent OOPSLA paper titled "Validating SMT Solvers for Correction and Performance via Grammar Based Enumeration" . Dominik shares his academic journey from the University of Freiburg to ETH Zurich, and now to a lectureship at the University of Manchester. He i...
Haralampos Gavriilidis | Fast and Scalable Data Transfer across Data Systems | #62 16.06.2025 56:46
In this episode of Disseminate , we welcome Harry Gavrilidis back to the podcast to explore his latest research on fast and scalable data transfer across systems, soon to be presented at SIGMOD 2025. Building on his work with XDB, Harry introduces XDBC , a novel data transfer framework designed to balance performance and generalizability. They dive into the challenges of moving data across heterog...
Haralampos Gavriilidis | SheetReader: Efficient spreadsheet parsing 17.04.2025 40:53
In this episode of the DuckDB in Research series, Harry Gavriilidis (PhD student at TU Berlin) joins us to discuss Sheet Reader — a high-performance spreadsheet parser that dramatically outpaces traditional tools in both speed and memory efficiency. By taking advantage of the standardized structure of spreadsheet files and bypassing generic XML parsers, Sheet Reader delivers fast and lightweight p...
Arjen P. de Vries | faiss: An extension for vector data & search 10.04.2025 46:14
In this episode of the DuckDB in Research series, we’re joined by Arjen de Vries, Professor of Data Science at Radboud University. Arjen dives into his team’s development of a DuckDB extension for FAISS, a library originally developed at Facebook for efficient similarity search and vector operations. We explore the growing importance of embeddings and dense retrieval in modern information retrieva...
David Justen | POLAR: Adaptive and non-invasive join order selection via plans of least resistance 03.04.2025 51:08
In this episode, we sit down with David Justen to discuss his work on POLAR: Adaptive and Non-invasive Join Order Selection via Plans of Least Resistance which was implemented in DuckDB. David shares his journey in the database space, insights into performance optimization, and the challenges of working with modern analytical workloads. We dive into the intricacies of query compilation, vectorized...
Daniël ten Wolde | DuckPGQ: A graph extension supporting SQL/PGQ 20.03.2025 48:38
In this episode, we sit down with Daniël ten Wolde, a PhD researcher at CWI’s Database Architectures Group, to explore DuckPGQ—an extension to DuckDB that brings powerful graph querying capabilities to relational databases. Daniel shares his journey into database research, the motivations behind DuckPGQ, and how it simplifies working with graph data. We also dive into the technical challenges of i...
Till Döhmen | DuckDQ: A Python library for data quality checks in ML pipelines 13.03.2025 58:12
In this episode we kick off our DuckDB in Research series with Till Döhmen, a software engineer at MotherDuck, where he leads AI efforts. Till shares insights into DuckDQ , a Python library designed for efficient data quality validation in machine learning pipelines, leveraging DuckDB’s high-performance querying capabilities. We discuss the challenges of ensuring data integrity in ML workflows, th...
Disseminate x DuckDB Coming Soon... 06.03.2025 2:40
Hey folks! We have been collaborating with everyone's favourite in-process SQL OLAP database management system DuckDB to bring you a new podcast series - the DuckDB in Research series! At Disseminate our mission is to bridge the gap between research and industry by exploring research that has a real-world impact. DuckDB embodies this synergy—decades of research underpin its design, and now it’s ma...
High Impact in Databases with... Anastasia Ailamaki 03.03.2025 46:17
In this High Impact in Databases episode we talk to Anastasia Ailamaki . Anastasia is a Professor of Computer and Communication Sciences at the École Polytechnique Fédérale de Lausanne (EPFL). Tune in to hear Anastasia's story! The podcast is proudly sponsored by Pometry the developers behind Raphtory , the open source temporal graph analytics engine for Python and Rust. You can find Anastasia on:...
Anastasiia Kozar | Fault Tolerance Placement in the Internet of Things | #61 16.12.2024 49:02
In this episode, we chat with Anastasiia Kozar about her research on fault tolerance in resource-constrained environments. As IoT applications leverage sensors, edge devices, and cloud infrastructure, ensuring system reliability at the edge poses unique challenges. Unlike the cloud, edge devices operate without persistent backups or high availability standards, leading to increased vulnerability t...
Liana Patel | ACORN: Performant and Predicate-Agnostic Hybrid Search | #60 11.11.2024 52:49
In this episode, we chat with with Liana Patel to discuss ACORN, a groundbreaking method for hybrid search in applications using mixed-modality data. As more systems require simultaneous access to embedded images, text, video, and structured data, traditional search methods struggle to maintain efficiency and flexibility. Liana explains how ACORN, leveraging Hierarchical Navigable Small Worlds (HN...
High Impact in Databases with... David Maier 04.11.2024 1:02:24
In this High Impact episode we talk to David Maier . David is the Maseeh Professor Emeritus of Emerging Technologies at Portland State University. Tune in to hear David's story and learn about some of his most impactful work. The podcast is proudly sponsored by Pometry the developers behind Raphtory , the open source temporal graph analytics engine for Python and Rust. You can find David on: Homep...
Raunak Shah | R2D2: Reducing Redundancy and Duplication in Data Lakes | #59 28.10.2024 31:09
In this episode, Raunak Shah joins us to discuss the critical issue of data redundancy in enterprise data lakes, which can lead to soaring storage and maintenance costs. Raunak highlights how large-scale data environments, ranging from terabytes to petabytes, often contain duplicate and redundant datasets that are difficult to manage. He introduces the concept of "dataset containment" and explains...
High Impact in Databases with... Aditya Parameswaran 21.10.2024 58:57
In this High Impact episode we talk to Aditya Parameswaran about his some of his most impactful work. Aditya is an Associate Professor at the University of California, Berkeley. Tune in to hear Aditya's story! The podcast is proudly sponsored by Pometry the developers behind Raphtory , the open source temporal graph analytics engine for Python and Rust. Links: EPIC Data Lab Answering Queries using...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.