Redwood Research
Redwood Research Blog
Narrations of Redwood Research blog posts. Redwood Research is a research nonprofit based in Berkeley. We investigate risks posed by the development of powerful artificial intelligence and techniques for mitigating those risks.
Author
Redwood Research
Category
Podcast website
Latest episode
Jul 2, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
“Notes on fatalities from AI takeover” by Ryan Greenblatt 23.09.2025 15:54
Subtitle: Large fractions of people will die, but literal human extinction seems unlikely. Suppose misaligned AIs take over. What fraction of people will die? I'll discuss my thoughts on this question and my basic framework for thinking about it. These are some pretty low-effort notes, the topic is very speculative, and I don't get into all the specifics, so be warned. I don't think moderate disag...
“Focus transparency on risk reports, not safety cases” by Ryan Greenblatt 22.09.2025 11:51
Subtitle: Transparency about just safety cases would have bad epistemic effects. There are many different things that AI companies could be transparent about. One relevant axis is transparency about the current understanding of risks and the current mitigations of these risks. I think transparency about this should take the form of a publicly disclosed risk report rather than the company making a...
“Prospects for studying actual schemers” by Ryan Greenblatt, Julian Stastny 19.09.2025 1:41:28
Subtitle: Studying actual schemers seems promising but tricky. One natural way to research scheming is to study AIs that are analogous to schemers. Research studying current AIs that are intended to be analogous to schemers won't yield compelling empirical evidence or allow us to productively study approaches for eliminating scheming, as these AIs aren't very analogous in practice. This difficulty...
“What training data should developers filter to reduce risk from misaligned AI?” by Alek Westover 17.09.2025 44:01
Subtitle: An initial narrow proposal. One potentially powerful way to change the properties of AI models is to change their training data. For example, Anthropic has explored filtering training data to mitigate bio misuse risk. What data, if any, should be filtered to reduce misalignment risk? In this post, I argue that the highest ROI data to filter is information about safety measures, and strat...
“AIs will greatly change engineering in AI companies well before AGI” by Ryan Greenblatt 09.09.2025 26:33
Subtitle: AIs that speed up engineering by 2x wouldn't accelerate AI progress that much. In response to my recent post arguing against above-trend progress from better RL environments, yet another argument for short(er) AGI timelines was raised to me: Sure, at the current rate of progress, it would take a while (e.g. 5 years) to reach AIs that can automate research engineering within AI companies...
“Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro” by Ryan Greenblatt 03.09.2025 14:05
Subtitle: Above trend progress due to a rapid increase in RL env quality is unlikely. I've recently written about how I've updated against seeing substantially faster than trend AI progress due to quickly massively scaling up RL on agentic software engineering. One response I've heard is something like: RL scale-ups so far have used very crappy environments due to difficulty quickly sourcing enoug...
“Attaching requirements to model releases has serious downsides (relative to a different deadline for these requirements)” by Ryan Greenblatt 27.08.2025 6:41
Subtitle: System cards are established but other approaches seem importantly better. Here's a relatively important question regarding transparency requirements for AI companies: At which points in time should AI companies be required to disclose information? (While I focus on transparency, this question is also applicable to other safety-relevant requirements, and is applicable to norms around vol...
“Notes on cooperating with unaligned AI” by Lukas Finnveden 24.08.2025 1:01:20
Subtitle: More thoughts on making deals with schemers. These are some research notes on whether we could reduce AI takeover risk by cooperating with unaligned AIs. I think the best and most readable public writing on this topic is “Making deals with early schemers”, so if you haven't read that post, I recommend starting there. These notes were drafted before that post existed, and the content is s...
“Being honest with AIs” by Lukas Finnveden 21.08.2025 36:20
Subtitle: When and why we should refrain from lying. In the future, we might accidentally create AIs with ambitious goals that are misaligned with ours. But just because we don’t have the same goals doesn’t mean we need to be in conflict. We could also cooperate with each other and pursue mutually beneficial deals. For previous discussion of this, you can read Making deals with early schemers. I r...
“My AGI timeline updates from GPT-5 (and 2025 so far)” by Ryan Greenblatt 20.08.2025 7:28
Subtitle: AGI before 2029 now seems substantially less likely. As I discussed in a prior post, I felt like there were some reasonably compelling arguments for expecting very fast AI progress in 2025 (especially on easily verified programming tasks). Concretely, this might have looked like reaching 8 hour 50% reliability horizon lengths on METR's task suite1 by now due to greatly scaling up RL and...
“Four places where you can put LLM monitoring” by Fabien Roger, Buck Shlegeris 09.08.2025 15:50
Subtitle: To wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems. To prevent potentially misaligned LLM agents from taking actions with catastrophic consequences, you can try to monitor LLM actions - that is, try to detect dangerous or malicious actions, and do something about it when you do (like blocking the action, starting an investigation, …).1 But where in your in...
“Four places where you can put LLM monitoring” by Fabien Roger, Buck Shlegeris 09.08.2025 15:10
Subtitle: To wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems. To prevent potentially misaligned LLM agents from taking actions with catastrophic consequences, you can try to monitor LLM actions - that is, try to detect dangerous or malicious actions, and do something about it when you do (like blocking the action, starting an investigation, …).1 But where in your in...
“Should we update against seeing relatively fast AI progress in 2025 and 2026?” by Ryan Greenblatt 28.07.2025 5:56
Subtitle: Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.. Around the early o3 announcement (and maybe somewhat before that?), I felt like there were some reasonably compelling arguments for putting a decent amount of weight on relatively fast AI progress in 2025 (and maybe in 2026): Maybe AI companies will be able to rapidly scale up RL further because RL...
“Should we update against seeing relatively fast AI progress in 2025 and 2026?” by Ryan Greenblatt 28.07.2025 5:47
Subtitle: Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.. Around the early o3 announcement (and maybe somewhat before that?), I felt like there were some reasonably compelling arguments for putting a decent amount of weight on relatively fast AI progress in 2025 (and maybe in 2026): Maybe AI companies will be able to rapidly scale up RL further because RL...
“Why it’s hard to make settings for high-stakes control research” by Buck Shlegeris 18.07.2025 7:20
Subtitle: It's like making challenging evals, but more constrained. One of our main activities at Redwood is writing follow-ups to previous papers on control like the original and Ctrl-Z, where we construct a setting with a bunch of tasks (e.g. APPS problems) and a notion of safety failure (e.g. backdoors according to our specific definition), then play the adversarial game where we develop protoc...
“Why it’s hard to make settings for high-stakes control research” by Buck Shlegeris 18.07.2025 7:03
Subtitle: It's like making challenging evals, but more constrained. One of our main activities at Redwood is writing follow-ups to previous papers on control like the original and Ctrl-Z, where we construct a setting with a bunch of tasks (e.g. APPS problems) and a notion of safety failure (e.g. backdoors according to our specific definition), then play the adversarial game where we develop protoc...
“Recent Redwood Research project proposals” by Ryan Greenblatt, Buck Shlegeris, Julian Stastny, Josh Clymer, Alex Mallen, Vivek Hebbar 14.07.2025 8:52
Subtitle: Empirical AI security/safety projects across a variety of areas. Previously, we've shared a few higher-effort project proposals relating to AI control in particular. In this post, we'll share a whole host of less polished project proposals. All of these projects excite at least one Redwood researcher, and high-quality research on any of these problems seems pretty valuable. They differ w...
“Recent Redwood Research project proposals” by Ryan Greenblatt, Buck Shlegeris, Julian Stastny, Josh Clymer, Alex Mallen, Vivek Hebbar 14.07.2025 8:30
Subtitle: Empirical AI security/safety projects across a variety of areas. Previously, we've shared a few higher-effort project proposals relating to AI control in particular. In this post, we'll share a whole host of less polished project proposals. All of these projects excite at least one Redwood researcher, and high-quality research on any of these problems seems pretty valuable. They differ w...
“What’s worse, spies or schemers?” by Buck Shlegeris, Julian Stastny 09.07.2025 10:33
Subtitle: And what if you have both at once?. Here are two potential problems you’ll face if you’re an AI lab deploying powerful AI: Spies: Some of your employees might be colluding to do something problematic with your AI, such as trying to steal its weights, use it for malicious intellectual labour (e.g. planning a coup or building a weapon), or install a secret loyalty. Schemers: Some of your A...
“Ryan on the 80,000 Hours podcast” by Buck Shlegeris 08.07.2025 1:42
Discover more from Redwood Research blog We research catastrophic AI risks and techniques that could be used to mitigate them. Over 2,000 subscribers Already have an account? Sign in Ryan on the 80,000 Hours podcast Ryan's podcast with Rob Wiblin has just come out! I think it turned out great. I particularly enjoyed Ryan's discussion of different pathways to AI takeover, which I don’t think has be...
“How much novel security-critical infrastructure do you need during the singularity?” by Buck Shlegeris 05.07.2025 10:08
Subtitle: And what does this mean for AI control?. I think a lot about the possibility of huge numbers of AI agents doing AI R&D inside an AI company (as depicted in AI 2027). I think particularly about what will happen if those AIs are scheming: coherently and carefully trying to grab power and take over the AI company, as a prelude to taking over the world. And even more particularly, I thin...
“Two proposed projects on abstract analogies for scheming” by Julian Stastny 04.07.2025 7:05
Subtitle: We should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.. In order to empirically study risks from schemers, we can try to develop model organisms of misalignment. Sleeper Agents and password-locked models, which train LLMs to behave in a benign or malign way depending on some feature of the input, are prominent examples of the m...
“There are two fundamentally different constraints on schemers” by Buck Shlegeris 02.07.2025 7:59
Subtitle: "They need to act aligned" often isn't precise enough. People (including me) often say that scheming models “have to act as if they were aligned”. This is an alright summary; it's accurate enough to use when talking to a lay audience. But if you want to reason precisely about threat models arising from schemers, or about countermeasures to scheming, I think it's important to make some fi...
“Jankily controlling superintelligence” by Ryan Greenblatt 27.06.2025 14:18
Subtitle: How much time can control buy us during the intelligence explosion?. When discussing AI control, we often talk about levels of AI capabilities where we think control can probably greatly lower risks and where we can probably estimate risks. However, I think it's plausible that an important application of control is modestly improving our odds of surviving significantly superhuman systems...
“What does 10x-ing effective compute get you?” by Ryan Greenblatt 24.06.2025 22:52
Subtitle: Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?. This is more speculative and confusing than my typical posts and I also think the content of this post could be substantially improved with more effort. But it's been sitting around in my drafts for a long time and I sometimes want to reference the arguments in it, so I thought I would go ahe...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.