Redwood Research

Redwood Research Blog

Narrations of Redwood Research blog posts. Redwood Research is a research nonprofit based in Berkeley. We investigate risks posed by the development of powerful artificial intelligence and techniques for mitigating those risks.

Author

Redwood Research

Category

Technology

Latest episode

Jul 2, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

“Takeaways from sketching a control safety case” by Josh Clymer 30.01.2025

Subtitle: Insights from a long technical paper compressed into a fun little commentary. Buck Shlegeris and I recently published a paper with UK AISI that sketches a safety case for “AI control” – measures that improve safety despite intentional subversion from AI systems. I would summarize this work as “turning crayon drawings of safety cases into blueprints.” It's a long and technical paper! So I...

“Planning for Extreme AI Risks” by Josh Clymer 29.01.2025

Subtitle: Are we ready for this?. This post should not be taken as a polished recommendation to AI companies and instead should be treated as an informal summary of a worldview. The content is inspired by conversations with a large number of people, so I cannot take credit for any of these ideas. For a summary of this post, see the thread on X. Many people write opinions about how to handle advanc...

“Ten people on the inside” by Buck Shlegeris 28.01.2025

Subtitle: A scary scenario that's worth planning for. (Many of these ideas developed in conversation with Ryan Greenblatt) In a shortform, I described some different levels of resources and buy-in for misalignment risk mitigations that might be present in AI labs: *The “safety case” regime.* Sometimes people talk about wanting to have approaches to safety such that if all AI developers followed th...

“When does capability elicitation bound risk?” by Josh Clymer 22.01.2025

For a summary of this post see the thread on X. The assumptions behind and limitations of capability elicitation have been discussed in multiple places (e.g. here, here, here, etc); however, none of these match my current picture. For example, Evan Hubinger's “When can we trust model evaluations?” cites “gradient hacking” (strategically interfering with the learning process) as the main reason sup...

“How will we update about scheming?” by Ryan Greenblatt 19.01.2025

Subtitle: A quantitative description of how I expect to change my mind.. [Cross-posted from LessWrong] I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released "Alignment Faking in Large Language Models", which provides empirical evidence for some components of the scheming thre...

“Thoughts on the conservative assumptions in AI control” by Buck Shlegeris 17.01.2025

Subtitle: Why are we so friendly to the red team?. Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper, Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats, How to prevent collusion when using untrusted models to monitor...

“Extending control evaluations to non-scheming threats” by Josh Clymer 13.01.2025

Buck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model psychologies are not well described by the standard notion of scheming. For example, models might display scheming-like behavior inconsistently or acquire harmful goals...

“Measuring whether AIs can statelessly strategize to subvert security measures” by Buck Shlegeris, Alex Mallen 20.12.2024

Subtitle: The complement to control evaluations. One way to show that risk from deploying an AI system is small is by showing that the model is not capable of subverting security measures enough to cause substantial harm. In AI control research so far, control evaluations have measured whether a red-team-created attack policy can defeat control measures, but they haven't measured how effectively m...

“Alignment Faking in Large Language Models” by Ryan Greenblatt, Buck Shlegeris 18.12.2024

Subtitle: In our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.. What happens when you tell an AI it is being trained to do something it doesn't want to do? We have a new paper (done in collaboration with Anthropic) demonstrating that, in our experiments, AIs will often strategically pretend to...

“Why imperfect adversarial robustness doesn’t doom AI control” by Buck Shlegeris 18.11.2024

Subtitle: There are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.. (thanks to Alex Mallen, Cody Rushing, Zach Stein-Perlman, Hoagy Cunningham, Vlad Mikulik, and Fabien Roger for comments) Sometimes I hear people argue against AI control as follows: if your control measures rely on getting good judgments from "trusted" AI models, you're doomed...

“Win/continue/lose scenarios and execute/replace/audit protocols” by Buck Shlegeris 15.11.2024

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective. In brief: Consider a deployment of an AI in a setting where it's going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we win, we lose, or the deployment c...

“Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren’t scheming” by Buck Shlegeris 10.10.2024

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest...

“A basic systems architecture for AI agents that do autonomous research” by Buck Shlegeris 26.09.2024

Subtitle: And diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.. A lot of threat models describing how AIs might escape our control (e.g. self-exfiltration, hacking the datacenter) start out with AIs that are acting as agents working autonomously on research tasks (especially AI R&D) in a datacenter controlled by the AI compan...

“How to prevent collusion when using untrusted models to monitor each other” by Buck Shlegeris 25.09.2024

Suppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions. You’re worried that this model is scheming, and you’re worried that it might only need to take a small number of actions to get to a dangerous and hard-to-reverse situation like exfiltrating its own weights. Problems like these that the AI can cau...

“Would catching your AIs trying to escape convince AI developers to slow down or undeploy?” by Buck Shlegeris 26.08.2024

Subtitle: I'm not so sure.. [Crossposted from LessWrong] I often talk to people who think that if frontier models were egregiously misaligned and powerful enough to pose an existential threat, you could get AI developers to slow down or undeploy models by producing evidence of their misalignment. I'm not so sure. As an extreme thought experiment, I’ll argue this could be hard even if you caught yo...

“Fields that I reference when thinking about AI takeover prevention” by Buck Shlegeris 13.08.2024

Subtitle: Is AI takeover like a nuclear meltdown? A coup? A plane crash?. My day job is thinking about safety measures that aim to reduce catastrophic risks from AI (especially risks from egregious misalignment). The two main themes of this work are the design of such measures (what's the space of techniques we might expect to be affordable and effective) and their evaluation (how do we decide whi...

“Getting 50% (SoTA) on ARC-AGI with GPT-4o” by Ryan Greenblatt 17.06.2024

Subtitle: You can just draw more samples. I recently got to 50%1 accuracy on the public test set for ARC-AGI by having GPT-4o generate a huge number of Python implementations of the transformation rule (around 8,000 per problem) and then selecting among these implementations based on correctness of the Python programs on the examples (if this is confusing, go to the next section)2. I use a variety...

Listen to the Redwood Research Blog podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.