BlueDot Impact

BlueDot Narrated

Audio versions of the core readings, blog posts, and papers from BlueDot courses.

Author

BlueDot Impact

Category

Technology

Podcast website

bluedot.org

Latest episode

Jan 16, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Is Power-Seeking AI an Existential Risk? 04.01.2025

Audio versions of blogs and papers from BlueDot courses. This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playi...

Measuring Progress on Scalable Oversight for Large Language Models 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract:  Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our...

Supervising Strong Learners by Amplifying Weak Experts 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract:  Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by demonstrating or judging performance, but this approach fails if the task is too complicated for a human...

Summarizing Books With Human Feedback 04.01.2025

Audio versions of blogs and papers from BlueDot courses. To safely deploy powerful, general-purpose artificial intelligence in the future, we need to ensure that machine learning models act in accordance with human intentions. This challenge has become known as the alignment problem. A scalable solution to the alignment problem needs to work on tasks where model outputs are difficult or time-consu...

Least-To-Most Prompting Enables Complex Reasoning in Large Language Models 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Chain-of-thought prompting has demonstrated remarkable performance on various natural language reasoning tasks. However, it tends to perform poorly on tasks which requires solving problems harder than the exemplars shown in the prompts. To overcome this challenge of easy-to-hard generalization, we propose a novel prompting strategy, least-to...

AI Safety via Debate 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract: To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and useful, but this approach can fail if the task is too complicated for a human to directly jud...

AI Safety via Red Teaming Language Models With Language Models 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract:  Language Models (LMs) often cannot be deployed because of their potential to harm users in ways that are hard to predict in advance. Prior work identifies harmful behaviors before deployment by using human annotators to hand-write test cases. However, human annotation is expensive, limiting the number and diversity of test cases....

Robust Feature-Level Adversaries Are Interpretability Tools 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract:  The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates the latent representations of image generators to create "feature-level" adversarial perturbations gives us an opportunity to explore percepti...

Debate Update: Obfuscated Arguments Problem 04.01.2025

Audio versions of blogs and papers from BlueDot courses. This is an update on the work on AI Safety via Debate that we previously wrote about here .  What we did:  We tested the debate protocol introduced in AI Safety via Debate with human judges and debaters. We found various problems and improved the mechanism to fix these issues (details of these are in the appendix). However, we discovered tha...

Introduction to Logical Decision Theory for Computer Scientists 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Decision theories differ on exactly how to calculate the expectation--the probability of an outcome, conditional on an action. This foundational difference bubbles up to real-life questions about whether to vote in elections, or accept a lowball offer at the negotiating table. When you're thinking about what happens if you don't vo...

High-Stakes Alignment via Adversarial Training [Redwood Research Report] 04.01.2025

Audio versions of blogs and papers from BlueDot courses. (Update: We think the tone of this post was overly positive considering our somewhat weak results. You can read our latest post with more takeaways and followup results here.)  This post motivates and summarizes this paper from Redwood Research, which presents results from the project first introduced here. We used adversarial training to im...

Takeaways From Our Robust Injury Classifier Project [Redwood Research] 04.01.2025

Audio versions of blogs and papers from BlueDot courses. With the benefit of hindsight, we have a better sense of our takeaways from our first adversarial training project (paper). Our original aim was to use adversarial training to make a system that (as far as we could tell) never produced injurious completions. If we had accomplished that, we think it would have been the first demonstration of...

Acquisition of Chess Knowledge in Alphazero 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract: What is learned by sophisticated neural network agents such as AlphaZero? This question is of both scientific and practical interest. If the representations of strong neural networks bear no resemblance to human concepts, our ability to understand faithful explanations of their decisions will be restricted, ultimately limiting what...

Feature Visualization 04.01.2025

Audio versions of blogs and papers from BlueDot courses. There is a growing sense that neural networks need to be interpretable to humans. The field of neural network interpretability has formed in response to these concerns. As it matures, two major threads of research have begun to coalesce: feature visualization and attribution. This article focuses on feature visualization. While feature visua...

Understanding Intermediate Layers Using Linear Classifier Probes 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Abstract: Neural network models have a reputation for being black boxes. We propose to monitor the features at every layer of a model and measure how suitable they are for classification. We use linear classifiers, which we refer to as "probes", trained entirely independently of the model itself.  This helps us better understand th...

Embedded Agents 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Suppose you want to build a robot to achieve some real-world goal for you—a goal that requires the robot to learn for itself and figure out a lot of things that you don’t already know. There’s a complicated engineering problem here. But there’s also a problem of figuring out what it even means to build a learning agent like that. What is it...

Logical Induction (Blog Post) 04.01.2025

Audio versions of blogs and papers from BlueDot courses. MIRI is releasing a paper introducing a new model of deductively limited reasoning: “Logical induction,” authored by Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, myself, and Jessica Taylor. Readers may wish to start with the abridged version.  Consider a setting where a reasoner is observing a deductive process (such as a community o...

Cooperation, Conflict, and Transformative Artificial Intelligence: Sections 1 & 2 — Introduction, Strategy and Governance 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Transformative artificial intelligence (TAI) may be a key factor in the long-run trajectory of civilization. A growing interdisciplinary community has begun to study how the development of TAI can be made safe and beneficial to sentient life (Bostrom 2014; Russell et al., 2015; OpenAI, 2018; Ortega and Maini, 2018; Dafoe, 2018). We present a...

Careers in Alignment 04.01.2025

Audio versions of blogs and papers from BlueDot courses. Richard Ngo compiles a number of resources for thinking about careers in alignment research. Original text: https://docs.google.com/document/d/1iFszDulgpu1aZcq_aYFG7Nmcr5zgOhaeSwavOMk1akw/edit#heading=h.4whc9v22p7tb Narrated for AI Safety Fundamentals by Perrin Walker of TYPE III AUDIO . --- A podcast by BlueDot Impact .

Progress on Causal Influence Diagrams 04.01.2025

Audio versions of blogs and papers from BlueDot courses. By Tom Everitt, Ryan Carey, Lewis Hammond, James Fox, Eric Langlois, and Shane Legg About 2 years ago, we released the first few papers on understanding agent incentives using causal influence diagrams. This blog post will summarize progress made since then. What are causal influence diagrams? A key problem in AI alignment is understanding a...

We Need a Science of Evals 04.01.2025

Audio versions of blogs and papers from BlueDot courses.  This lays out a number of open questions, in what the author calls a 'Science of Evals'. Original text: https://www.apolloresearch.ai/blog/we-need-a-science-of-evals Author(s): Apollo Research blog A podcast by BlueDot Impact .

Introduction to Mechanistic Interpretability 04.01.2025

Audio versions of blogs and papers from BlueDot courses.  Our introduction introduces common mech interp concepts, to prepare you for the rest of this session's resources. Original text: https://aisafetyfundamentals.com/blog/introduction-to-mechanistic-interpretability/ Author(s): Sarah Hastings-Woodhouse A podcast by BlueDot Impact .

Constitutional AI Harmlessness from AI Feedback 04.01.2025

Audio versions of blogs and papers from BlueDot courses.  This paper explains Anthropic’s constitutional AI approach, which is largely an extension on RLHF but with AIs replacing human demonstrators and human evaluators. A podcast by BlueDot Impact .

Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback 04.01.2025

Audio versions of blogs and papers from BlueDot courses.  This paper explains Anthropic’s constitutional AI approach, which is largely an extension on RLHF but with AIs replacing human demonstrators and human evaluators. A podcast by BlueDot Impact .

Illustrating Reinforcement Learning from Human Feedback (RLHF) 04.01.2025

Audio versions of blogs and papers from BlueDot courses.  This more technical article explains the motivations for a system like RLHF, and adds additional concrete details as to how the RLHF approach is applied to neural networks. While reading, consider which parts of the technical implementation correspond to the 'values coach' and 'coherence coach' from the previous video. A...

Listen to the BlueDot Narrated podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.