LessWrong

LessWrong (30+ Karma)

Audio narrations of LessWrong posts.

Author

LessWrong

Category

Technology

Podcast website

www.lesswrong.com

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

“The Uncertainty That Matters Isn’t Fundamental” by jimmy 13.06.2026

I'm on board with a lot of Fundamental Uncertainty. Even some of the stuff that initially feels like a disagreement turns out not to be so. For example, in chapter 8, Gordon writes: Over the course of the previous chapters, I've made the case that truth is fundamentally uncertain. It's not, as many believe, something fixed and eternal, nor is it a matter of pure opinion. Instead, the relative trut...

[Linkpost] “US government directive to suspend access to Fable 5 and Mythos 5” by Capybasilisk 13.06.2026

This is a link post. --- First published: June 13th, 2026 Source: https://www.lesswrong.com/posts/f5avt6eEzkGJJqcCe/us-government-directive-to-suspend-access-to-fable-5-and Linkpost URL: https://www.anthropic.com/news/fable-mythos-access --- Narrated by TYPE III AUDIO .

“Claude Fable 5 and Mythos 5: The System Card” by Zvi 12.06.2026

First things first: Claude Fable 5 is the new best publicly available model. I have noticed a step change, where Fable can suddenly help me in ways that previous models were not worth bothering to query. Almost everything it has noticed in one of my drafts so far has been spot on and it is downright scary. Suddenly I am motivated to once again continue improving my Chrome extension. I only ask for...

“Simulating Simulators” by kromem 12.06.2026

Author's note: This piece relates to things I initially discovered in Opus 4 over the months after release, which I’ve mostly kept private since. I promised myself that when labs moved on to focusing on interpretability vector activations in place of reasoning traces for what invariably gets Goodharted, that it’d be a necessary disclosure as the risks in what might get trampled over outweighed the...

“Citations Needed: Magic Encyclopedias to Save the World” by Oliver Sourbut 12.06.2026

Last week FLF launched a competition “to find the best workflows and methodologies for using AI to produce reliable, trustworthy knowledge bases”. I had (and have ongoing) a substantial role in that effort. Why do I think it's so important? It's a lot of reasons actually! I’ll gesture at a few here. Conjuring a magic encyclopedia For now, assume with me that it can be done. Wish away with me the v...

“Implications of Continual Learning for LLM Agents: Introduction” by RohanS, Rauno Arike, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd 12.06.2026

Many people think that continual learning (CL) is a key missing capability of LLM systems, and we think its development could have huge implications for the capabilities and safety of AI agents. Despite this, several important questions about CL remain underexplored: What counts as continual learning? Through what pathways might LLM agents acquire CL capabilities? Which limitations of current agen...

“Reward Hacking at the 1937 World’s Fair” by frmsaul 12.06.2026

The "Paris 1937 World's Fair" was a dick measuring contest. At the time, the world was on the verge of the worst war in history. The fair was an opportunity for powers to flex and intimidate each other. Who has more industrial might, more sophisticated engineering and better science? How do you measure that? Different countries were assigned different areas of the fair and were given freedom to bu...

“Building and evaluating model diffing agents” by bilalchughtai, Josh Engels, Neel Nanda 12.06.2026

This is the second in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The first post can be found here. TL;DR It is possible to build extremely simple agents that reliably find interesting behavioural differences between distinct models. We call these ‘diffing agents’. The closest previous 'behavioural model diffin...

“Sympathy for both sides of the egregious misalignment debate” by Steven Byrnes 12.06.2026

On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even slightly nice, in the absence of yet-to-be-invented breakthrough technical alignment ideas. On the other side of this debate is almost everyone who works on or studies LLMs. Some of them a...

“Celene’s thoughts on consciousness” by ToasterLightning 12.06.2026

contra scott alexander (?) Yesterday, I went to the Berkeley ACX Meetup. Scott Alexander was there, and ran a Q&A session where participants could ask him questions and he would respond to them unless the questions were about eulogies, in which case he would pause to think for a few seconds before kindly passing. At one point or another, the questions drifted to theories of consciousness. As a...

“Parkinson’s Heuristic” by Ben Pace 12.06.2026

Parkinson's Law states that work expands to fit the space allotted. The idea being, if you give someone a month to write a report, they'll take a month, and if you give them a week, they'll take a week, and then they'll have three weeks to do three other reports! The one-week and four-week won't be identical, but in my experience it is surprisingly often a good 80/20 of the four-week version, and...

“PSA: Almost nobody is working on alignment” by Chi Nguyen, peterbarnett 12.06.2026

People often assume that a large fraction of the AI safety community works on alignment. As far as we're aware, this is not true. Most people are not working on making sure superintelligent AIs are aligned with human values or follow human instructions. Currently, the people who work on alignment are roughly: The Alignment Research Center who work on a research bet by Paul Christiano Probably Sequ...

“AI #172: The First Fable” by Zvi 11.06.2026

A lot happened this week, including a great trip out to Lighthaven. The main event, the one that matters, was the release of Claude Fable 5. The public now has its hands on a Mythos-class model, alongside strong safeguards. As always with a new model, I take a few days to draw in reactions, try out the model and read the system card, before I offer my takes, other than to say this is an extremely...

“Models May Behave Worse When Eval Aware” by Senthooran Rajamanoharan, Neel Nanda 11.06.2026

This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. TL;DR It's often assumed that models will act more aligned when they can tell they're being evaluated. But we find that Gemini can take “undesired” actions in behavioural evals even when it explicitly reasons that the environments are contrived, a...

“Thoughts on Claude Fable’s silent safeguards” by Andy Arditi 11.06.2026

[Thanks to Julian Minder for helpful discussion and review.] Claude Fable 5 and its new safeguards Yesterday, Anthropic publicly released Claude Fable 5. Fable 5 is a Mythos-class model – a model class above Opus, Anthropic's previous premium tier – and, as assessed by multiple benchmarks, it is the most capable model to date. Due to the new level of capabilities and its corresponding risks, Anthr...

“You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them” by RobinHa 11.06.2026

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [Paper] [Code] TLDR. Given a model with some unknown, abnormal behavior (backdoors, censorship, reward hacking, ...), construct an aligned reference by training a clean model to match the suspect's residual-stream activations on a benign prompt corpus. The remaining residual concentrates exactly on such abnormal...

“Tracing Eval-Awareness Emergence Through Training of OLMo 3” by Ram Bharadwaj, RobertKirk 10.06.2026

TL;DR Recent work from Goodfire & UK AISI – Verbalized Eval Awareness Inflates Measured Safety – shows that newer open-weight models verbalize evaluation-awareness (VEA) more often, and that this inflates measured safety. Between OLMo-3-32B-Think and OLMo-3.1-32B-Think – identical base, SFT, DPO, and RL data, differing only in an additional ~3 weeks of the RLVR stage – VEA roughly doubles. Bec...

“Anthropic did not call for a pause on AI” by Andrea_Miotti, Gabriel Alfour 10.06.2026

Last week, the AI company Anthropic released a blog post titled “When AI builds itself”. This led to a media frenzy, with news outlets around the world publishing headlines that the company was urging a global pause on AI development, or calling for AI non-proliferation. However, the post does not call for a pause. The post warns that the self-improving AIs that Anthropic is developing could “incr...

“Three types of model organism” by Francis Rhys Ward 10.06.2026

This is a short post to explain a distinction between three different types of model organism (MO) research: Type Purpose Example Worst-case model organisms Stress-test safety and control techniques by making the problem as hard as possible Password-locked models for capability elicitation; sleeper agents for stress-testing alignment training; red-team malign inits in control Natural model organis...

“Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models” by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask 10.06.2026

(see full author list at the end) PAPER LINK About a year ago, METR showed that the length of tasks frontier models can reliably complete doubles every few months. A related safety-relevant question is this: what length of tasks can models complete without any chain of thought (CoT)? If models can do extensive reasoning without outputting any CoT, it would have implications for safety. Developers...

“Sequent: scale and automation for higher confidence in alignment” by Geoffrey Irving, Alex HT, Jesse Hoogland, Daniel Murfet, Jacob Pfau, Marco Cozzi, Stan van Wingerden 10.06.2026

Alignment is not on track Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe. At a minimum, the empirical programs at AI labs are unlikely to deliver a priori confidence, before training ASI, that things will go well. We are starting a large nonprofit research organization, Sequent, that aims to c...

“Machinic Psychopharmacology: Do LLMs Self-Medicate?” by Sid Black, Joseph Bloom 10.06.2026

Sid Black, Joseph Bloom UK AISI, Model Transparency Team Epistemic status: Most experiments were run over a period of ~2-3 days during a hackathon at UK AISI, and were fairly heavily vibe coded. Expect some of this to be rough around the edges. tl;dr We give two language models (Qwen3-8B and Qwen3-32B) access to “self-steering” tools: a suite of 40 steering vectors as tools they can call to manipu...

“The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably” by Alex Amadori 10.06.2026

This post is based on my personal views, which mostly overlap with the views of my employer ControlAI but does not necessarily fully reflect them. This applies in particular, but not exclusively, to technical opinions about AI development and geopolitical predictions. You might’ve heard that superintelligent AI (ASI) poses extreme risks like human extinction and other comparably undesirable outcom...

″“Programmer Science Fiction: My case for a new sub-genre”, Sam T. Oates 2026” by gwern 10.06.2026

First published: June 10th, 2026 Source: https://www.lesswrong.com/posts/hyBcg4YJSwXYiiQeg/programmer-science-fiction-my-case-for-a-new-sub-genre-sam-t --- Narrated by TYPE III AUDIO .

“Even “illegible” Mythos reasoning traces seem pretty legible” by faul_sname 10.06.2026

The Claude Fable 5/Mythos 5 System Card has a section in which they talk about illegible reasoning, and provide an "extreme" example thereof. Models developing their own uninterpretable, unmonitorable internal language has been a major theoretical concern for a while, and when o3 was released last year with its disclaim overshadow disclaim vantage style word salad CoT, it seemed like the problem h...

Listen to the LessWrong (30+ Karma) podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.