Jacob Haimes

Into AI Safety

The Into AI Safety podcast aims to make it easier for everyone, regardless of background, to get meaningfully involved with the conversations surrounding the rules and regulations which should govern the research, development, deployment, and use of the technologies encompassed by the term "artificial intelligence" or "AI"For better formatted show notes, additional resources, and more, go to https://kairos.fm/intoaisafety/

Author

Jacob Haimes

Category

Technology

Podcast website

kairos.fm

Latest episode

Jul 9, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Pretraining Safety w/ Ethan Roland 09.07.2026

What if the safest AI models weren't built by adding guardrails after training, but by shaping what gets learned in the first place? Ethan Roland, senior alignment researcher at AE Studio and first author on an ICML 2026 spotlight paper, joins Jacob to talk about gradient routing, a technique that routes dangerous capabilities into isolated parts of a model's architecture where they can be locked...

Reclaiming UBI in the AI Age w/ Joe Williams 01.06.2026

Today's episode does double duty as an interview and an announcement. Joe Williams, host of the new Kairos.fm show "Beyond the Paycheck: Reclaiming the Case for UBI in the Age of AI," joins Jacob to talk about his background as a freelance translator and how AI quietly dismantled his livelihood in 2025. From there the conversation expands into whether this moment is really different from past wave...

Building Asymmetric Defense w/ Zainab Majid 12.05.2026

Zainab Majid, co-founder of Asymmetric Security, joins Jacob for a conversation on the intersection between AI Safety and cybersecurity, as well as the future of digital forensics. Drawing from years of incident response work, she explains how cyber attacks actually unfold, why AI is changing both offense and defense, and how her team is building AI-native tools to investigate breaches faster and...

Drawing Red Lines w/ Su Cizem 06.04.2026

Technology has been moving faster than policy for some time now, and the advent of AI isn't changing that, so what can we do to maintain safety despite uncertainty? Su Cizem has spent the last few years trying to answer that question. As an analyst at the Future Society, she works on global AI governance, specifically on building international consensus around AI red lines: the thresholds we colle...

Thinking Through "Digital Minds" w/ Jacy Reese-Anthis 10.03.2026

Jacy Reese-Anthis, founder of Sentience Institute and researcher at Stanford, began his journey working for animal welfare, but is now finishing up his PhD with research in many different AI subfields at the intersection of neuroscience, philosophy, social science, and machine learning. While this may seem like an odd jump at first, Jacy shares how his work has all been centered around the idea of...

Scaling AI Safety Through Mentorship w/ Dr. Ryan Kidd 02.02.2026

What does it actually take to build a successful AI safety organization? I'm joined by Dr. Ryan Kidd, who has co-led MATS from a small pilot program to one of the field's premier talent pipelines. In this episode, he reveals the low-hanging fruit in AI safety field-building that most people are missing: the amplifier archetype. I pushed Ryan on some hard questions, from balancing funder priorities...

Sobering Up on AI Progress w/ Dr. Sean McGregor 29.12.2025

Sean McGregor and I discuss about why evaluating AI systems has become so difficult; we cover everything from the breakdown of benchmarking, how incentives shape safety work, and what approaches like BenchRisk (his recent paper at NeurIPS) and AI auditing aim to fix as systems move into the real world. We also talk about his history and journey in AI safety, including his PhD on ML for public poli...

Against 'The Singularity' w/ Dr. David Thorstad 24.11.2025

Philosopher Dr. David Thorstad tears into one of AI safety's most influential arguments: the singularity hypothesis. We discuss why the idea of recursive self-improvement leading to superintelligence doesn't hold up under scrutiny, how these arguments have redirected hundreds of millions in funding away from proven interventions, and why people keep backpedaling to weaker versions when challenged....

Getting Agentic w/ Alistair Lowe-Norris 20.10.2025

Alistair Lowe-Norris, Chief Responsible AI Officer at Iridius and co-host of The Agentic Insider podcast, joins to discuss AI compliance standards, the importance of narrowly scoping systems, and how procurement requirements could encourage responsible AI adoption across industries. We explore the gap between the empty promises companies provide and actual safety practices, as well as the importan...

Growing BlueDot's Impact w/ Li-Lian Ang 15.09.2025

I'm joined by my good friend, Li-Lian Ang, first hire and product manager at BlueDot Impact. We discuss how BlueDot has evolved from their original course offerings to a new "defense-in-depth" approach, which focuses on three core threat models: reduced oversight in high risk scenarios (e.g. accelerated warfare), catastrophic terrorism (e.g. rogue actors with bioweapons), and the concentration of...

Layoffs to Leadership w/ Andres Sepulveda Morales 04.08.2025

Andres Sepulveda Morales joins me to discuss his journey from three tech layoffs to founding Red Mage Creative and leading the Fort Collins chapter of the Rocky Mountain AI Interest Group (RMAIIG). We explore the current tech job market, AI anxiety in nonprofits, dark patterns in AI systems, and building inclusive tech communities that welcome diverse perspectives. Reach out to Andres on his Linke...

Getting Into PauseAI w/ Will Petillo 23.06.2025

Will Petillo, onboarding team lead at PauseAI , joins me to discuss the grassroots movement advocating for a pause on frontier AI model development. We explore PauseAI's strategy, talk about common misconceptions Will hears, and dig into how diverse perspectives still converge on the need to slow down AI development. Will's Links Personal blog on AI His mindmap of the AI x-risk debate Game demos A...

Making Your Voice Heard w/ Tristan & Felix de Simone 19.05.2025

I am joined by Tristan Williams and Felix de Simone to discuss their work on the potential of constituent communication, specifically in the context of AI legislation. These two worked as part of an AI Safety Camp team to understand whether or not it would be useful for more people to be sharing their experiences, concerns, and opinions with their government representative (hint, it is). Check out...

INTERVIEW: Scaling Democracy w/ (Dr.) Igor Krawczuk 03.06.2024

The almost Dr. Igor Krawczuk joins me for what is the equivalent of 4 of my previous episodes. We get into all the classics: eugenics, capitalism, philosophical toads... Need I say more? If you're interested in connecting with Igor, head on over to his website , or check out placeholder for thesis (it isn't published yet). Because the full show notes have a whopping 115 additional links, I'll high...

INTERVIEW: StakeOut.AI w/ Dr. Peter Park (3) 25.03.2024

As always, the best things come in 3s: dimensions, musketeers, pyramids, and... 3 installments of my interview with Dr. Peter Park, an AI Existential Safety Post-doctoral Fellow working with Dr. Max Tegmark at MIT. As you may have ascertained from the previous two segments of the interview, Dr. Park cofounded StakeOut. AI along with Harry Luk and one other cofounder whose name has been removed due...

INTERVIEW: StakeOut.AI w/ Dr. Peter Park (2) 18.03.2024

Join me for round 2 with Dr. Peter Park, an AI Existential Safety Postdoctoral Fellow working with Dr. Max Tegmark at MIT. Dr. Park was a cofounder of StakeOut. AI , a non-profit focused on making AI go well for humans , along with Harry Luk and one other individual, whose name has been removed due to requirements of her current position. In addition to the normal links, I wanted to include the li...

MINISODE: Restructure Vol. 2 11.03.2024

UPDATE: Contrary to what I say in this episode, I won't be removing any episodes that are already published from the podcast RSS feed. After getting some advice and reflecting more on my own personal goals, I have decided to shift the direction of the podcast towards accessible content regarding "AI" instead of the show's original focus. I will still be releasing what I am calling research ride-al...

INTERVIEW: StakeOut.AI w/ Dr. Peter Park (1) 04.03.2024

Dr. Peter Park is an AI Existential Safety Postdoctoral Fellow working with Dr. Max Tegmark at MIT. In conjunction with Harry Luk and one other cofounder, he founded ⁠StakeOut. AI , a non-profit focused on making AI go well for humans . 00:54 - Intro 03:15 - Dr. Park, x-risk, and AGI 08:55 - StakeOut. AI 12:05 - Governance scorecard 19:34 - Hollywood webinar 22:02 - Regulations.gov comments 23:48...

MINISODE: "LLMs, a Survey" 26.02.2024

Take a trip with me through the paper Large Language Models, A Survey , published on February 9th of 2024. All figures and tables mentioned throughout the episode can be found on the Into AI Safety podcast website . 00:36 - Intro and authors 01:50 - My takes and paper structure 04:40 - Getting to LLMs 07:27 - Defining LLMs & emergence 12:12 - Overview of PLMs 15:00 - How LLMs are built 18:52 -...

FEEDBACK: Applying for Funding w/ Esben Kran 19.02.2024

Esben reviews an application that I would soon submit for Open Philanthropy's Career Transitition Funding opportunity. Although I didn't end up receiving the funding, I do think that this episode can be a valuable resource for both others and myself when applying for funding in the future. Head over to Apart Research's website to check out their work, or the Alignment Jam website for information o...

MINISODE: Reading a Research Paper 12.02.2024

Before I begin with the paper-distillation based minisodes, I figured we would go over best practices for reading research papers. I go through the anatomy of typical papers, and some generally applicable advice. 00:56 - Anatomy of a paper 02:38 - Most common advice 05:24 - Reading sparsity and path 07:30 - Notes and motivation Links to all articles/papers which are mentioned throughout the episod...

HACKATHON: Evals November 2023 (2) 05.02.2024

Join our hackathon group for the second episode in the Evals November 2023 Hackathon subseries. In this episode, we solidify our goals for the hackathon after some preliminary experimentation and ideation. Check out Stellaric's website , or follow them on Twitter . 01:53 - Meeting starts 05:05 - Pitch: extension of locked models 23:23 - Pitch: retroactive holdout datasets 34:04 - Preliminary resul...

MINISODE: Portfolios 29.01.2024

I provide my thoughts and recommendations regarding personal professional portfolios. 00:35 - Intro to portfolios 01:42 - Modern portfolios 02:27 - What to include 04:38 - Importance of visual 05:50 - The "About" page 06:25 - Tools 08:12 - Future of "Minisodes" Links to all articles/papers which are mentioned throughout the episode can be found below, in order of their appearance. From Portafoglio...

INTERVIEW: Polysemanticity w/ Dr. Darryl Wright 22.01.2024

Darryl and I discuss his background, how he became interested in machine learning, and a project we are currently working on investigating the penalization of polysemanticity during the training of neural networks. Check out a diagram of the decoder task used for our research! 01:46 - Interview begins 02:14 - Supernovae classification 08:58 - Penalizing polysemanticity 20:58 - Our "toy model" 30:0...

MINISODE: Starting a Podcast 15.01.2024

A summary and reflections on the path I have taken to get this podcast started, including some resources recommendations for others who want to do something similar. Links to all articles/papers which are mentioned throughout the episode can be found below, in order of their appearance. LessWrong Spotify for Podcasters Into AI Safety podcast website Effective Altruism Global Open Broadcaster Softw...

Listen to the Into AI Safety podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.