Redwood Research
Redwood Research Blog
Narrations of Redwood Research blog posts. Redwood Research is a research nonprofit based in Berkeley. We investigate risks posed by the development of powerful artificial intelligence and techniques for mitigating those risks.
Author
Redwood Research
Category
Podcast website
Latest episode
Jul 2, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
“Comparing risk from internally-deployed AI to insider and outsider threats from humans” by Buck Shlegeris 23.06.2025 5:33
Subtitle: And why I think insider threat from AI combines the hard parts of both problems.. I’ve been thinking a lot recently about the relationship between AI control and traditional computer security. Here's one point that I think is important. My understanding is that there's a big qualitative distinction between two ends of a spectrum of security work that organizations do, that I’ll call “sec...
“Prefix cache untrusted monitors: a method to apply after you catch your AI” by Ryan Greenblatt 20.06.2025 14:04
Subtitle: Training the policy to not do egregious bad actions we detect has downsides and we might be able to do better. We often discuss what training you should do after catching your AI doing something egregiously bad. It seems relatively clear that it's good to train monitors on the (proliferated) bad action (assuming you can get acceptable hyperparameters and resolve other minor issues). But,...
“Making deals with early schemers” by Julian Stastny, Olli Järviniemi, Buck Shlegeris 20.06.2025 29:29
Subtitle: ...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.. Consider the following vignette: It is March 2028. With their new CoCo-Q neuralese reasoning model, a frontier AI lab has managed to fully automate the process of software engineering. In AI R&D, most human engineers have lost their old jobs, and only a small number of researchers now coo...
“AI safety techniques leveraging distillation” by Ryan Greenblatt 19.06.2025 22:08
Subtitle: Distillation is cheap; how can we use it to improve safety?. It's currently possible to (mostly or fully) cheaply reproduce the performance of a model by training another (initially weaker) model to imitate the stronger model's outputs.1 I'll refer to this as distillation. In the case of RL, distilling the learned capabilities is much, much cheaper than the RL itself (especially if you a...
“When does training a model change its goals?” by Vivek Hebbar, Ryan Greenblatt 12.06.2025 29:43
Subtitle: Can a scheming AI's goals really stay unchanged through training?. Here are two opposing pictures of how training interacts with deceptive alignment: “goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training object...
“The case for countermeasures to memetic spread of misaligned values” by Alex Mallen 28.05.2025 13:41
Subtitle: Defending against alignment problems that might come with long-term memory. As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coh...
“AIs at the current capability level may be important for future safety work” by Ryan Greenblatt 12.05.2025 7:01
Subtitle: Some reasons why relatively weak AIs might still be important when we have very powerful AIs. Sometimes people say that it's much less valuable to do AI safety research today than it will be in the future, because the current models are very different—in particular, much less capable—than the models that we’re actually worried about. I think that argument is mostly right, but it misses a...
“Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking” by Julian Stastny, Buck Shlegeris 08.05.2025 30:24
Subtitle: A new analysis of the risk of AIs intentionally performing poorly.. In the future, we will want to use powerful AIs on critical tasks such as doing AI safety and security R&D, dangerous capability evaluations, red-teaming safety protocols, or monitoring other powerful models. Since we care about models performing well on these tasks, we are worried about sandbagging: that if our mode...
“Training-time schemers vs behavioral schemers” by Alex Mallen 06.05.2025 13:23
Subtitle: Clarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.. People use the word “schemer” in two main ways: “Scheming” (or similar concepts: “deceptive alignment”, “alignment faking”) is often defined as a property of reasoning at training-time1. For example, Carlsmith defines a schemer a...
“What’s going on with AI progress and trends? (As of 5/2025)” by Ryan Greenblatt 03.05.2025 16:48
Subtitle: My views on what's driving AI progress and where it's headed.. AI progress is driven by improved algorithms and additional compute for training runs. Understanding what is going on with these trends and how they are currently driving progress is helpful for understanding the future of AI. In this post, I'll share a wide range of general takes on this topic as well as open questions. Be w...
“How can we solve diffuse threats like research sabotage with AI control?” by Vivek Hebbar 30.04.2025 15:45
Subtitle: Preventing research sabotage will require techniques very different from the original control paper.. Misaligned AIs might engage in research sabotage: making safety research go poorly by doing things like withholding their best ideas or putting subtle bugs in experiments. To mitigate this risk with AI control, we need very different techniques than those in the original control paper or...
“7+ tractable directions in AI control” by Ryan Greenblatt 29.04.2025 26:33
Subtitle: A list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute. In this post, we list 7 of our favorite easy-to-start directions in AI control. (Really, projects that are at least adjacent to AI control; We include directions which aren’t as centrally AI control and which also have other motivations.) This list is targeted at indep...
“Clarifying AI R&D threat models” by Josh Clymer 25.04.2025 9:54
Subtitle: (There are a few). A casual reader of one of the many AI company safety frameworks might be confused about why “AI R&D” is listed as a “threat model.” They might be even further mystified to find out that some people believe risks from automated AI R&D are the most severe and urgent risks. What's so concerning about AI writing code? A few months ago, I was frustrated by the lack...
“How training-gamers might function (and win)” by Vivek Hebbar 24.04.2025 32:48
Subtitle: A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.. In this post I present a model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents. This model suggests that given sufficiently rich training environments (and sufficient reasoning ability), models which terminally...
“Handling schemers if shutdown is not an option” by Buck Shlegeris 18.04.2025 24:43
Subtitle: What if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?. In most of our research and writing on AI control, we’ve emphasized the following situation: The AI developer is deploying a model that they think might be scheming, but they aren’t sure. The objective of the safety team is to ensure that if the model is scheming, it will be caugh...
“Ctrl-Z: Controlling AI Agents via Resampling” by Buck Shlegeris 16.04.2025 4:00
Subtitle: A new paper on AI control for agents.. We have released a new paper, Ctrl-Z: Controlling AI Agents via Resampling. This is the largest and most intricate study of control techniques to date: that is, techniques that aim to prevent catastrophic failures even if egregiously misaligned AIs attempt to subvert the techniques. We extend control protocols to a more realistic, multi-step setting...
“To be legible, evidence of misalignment probably has to be behavioral” by Ryan Greenblatt 15.04.2025 5:44
Subtitle: Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.. One key hope for mitigating risk from misalignment is inspecting the AI's behavior, noticing that it did something egregiously bad, converting this into legible evidence the AI is seriously misaligned, and then this triggering some strong and useful response (like spending relatively more re...
“Why do misalignment risks increase as AIs get more capable?” by Ryan Greenblatt 11.04.2025 7:21
Subtitle: A breakdown of how higher capabilities increase risk. It's generally agreed that as AIs get more capable, risks from misalignment increase. But there are a few different mechanisms by which more capable models are riskier, and distinguishing between those mechanisms is important when estimating the misalignment risk posed at a particular level of capabilities or by a particular model. Th...
“An overview of areas of control work” by Ryan Greenblatt 09.04.2025 54:35
Subtitle: What are all the research and implementation areas helpful for control?. In this post, I'll list all the areas of control research (and implementation) that seem promising to me. This references framings and abstractions discussed in Prioritizing threats for AI control. First, here is a division of different areas, though note that these areas have some overlaps: Developing and using set...
“An overview of control measures” by Ryan Greenblatt 06.04.2025 50:14
Subtitle: What methods can we use to ensure control?. We often talk about ensuring control, which in the context of this doc refers to preventing AIs from being able to cause existential problems, even if the AIs attempt to subvert our countermeasures. To better contextualize control, I think it's useful to discuss the main countermeasures (building on my prior discussion of the main threats). I'l...
“Buck on the 80,000 Hours podcast” by Buck Shlegeris 05.04.2025 1:19
Buck on the 80,000 Hours podcast My podcast with Rob Wiblin from 80,000 Hours just came out. I’m really happy with how it turned out. I talked about a bunch of stuff on the podcast that I don’t think we’ve written up before. Transcript + links + summary here. You can also get it as a podcast: Spotify: Apple: Subscribe to Redwood Research blog Launched a year ago We research catastrophic AI risks a...
“Notes on countermeasures for exploration hacking (aka sandbagging)” by Ryan Greenblatt 04.04.2025 15:26
Subtitle: How can we prevent AIs from intentionally underperforming on our metrics?. If we naively apply RL to a scheming AI, the AI may be able to systematically get low reward/performance while simultaneously not having this behavior trained out because it intentionally never explores into better behavior. As in, it intentionally puts very low probability on (some) actions which would perform ve...
“Notes on handling non-concentrated failures with AI control: high level methods and different regimes” by Ryan Greenblatt 03.04.2025 32:52
Subtitle: What are the methods and issues when failures occur diffusely over many actions?. In this post, I'll try to explain my current understanding of the high level methods for handling non-concentrated failures with control. I'll discuss the regimes produced by different methods and the failure modes of these different regimes. Non-concentrated failures are issues that arise from the AI doing...
“Prioritizing threats for AI control” by Ryan Greenblatt 19.03.2025 20:43
Subtitle: What are the main threats and how should we prioritize them?. We often talk about ensuring control, which in the context of this doc refers to preventing AIs from being able to cause existential problems, even if the AIs attempt to subvert our countermeasures. To better contextualize control, I think it's useful to discuss the main threats. I'll focus my discussion on threats induced by...
“How might we safely pass the buck to AI?” by Josh Clymer 19.02.2025 1:12:10
Subtitle: Developing AI employees that are safer than human ones. My goal as an AI safety researcher is to put myself out of a job. I don’t worry too much about how planet sized brains will shape galaxies in 100 years. That's something for AI systems to figure out. Instead, I worry about safely replacing human researchers with AI agents, at which point human researchers are “obsolete.” The situati...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.