LessWrong

LessWrong (30+ Karma)

Audio narrations of LessWrong posts.

Author

LessWrong

Category

Technology

Podcast website

www.lesswrong.com

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

“Claude Fable 5 and Mythos 5 [Linkpost]” by fluxxrider 10.06.2026

This is a linkpost for https://www.anthropic.com/news/claude-fable-5-mythos-5 --- First published: June 9th, 2026 Source: https://www.lesswrong.com/posts/kXjtbs7QqavJ7RwNE/claude-fable-5-and-mythos-5-linkpost --- Narrated by TYPE III AUDIO .

“Three Labs With a Plan and A Memorandum” by Zvi 10.06.2026

The big story today is the release of Claude Fable 5, the version of Claude Mythos that Anthropic believes they can safely distribute to the people. You should absolutely be switching over to that model and trying it out. But as always, this blog does not rush into commenting on a new model until we have a few days to play around with it and see what our new baby can (and can’t) do. This will be n...

“A Mike’s-Eye View of ARC’s Research” by Jacob_Hilton 09.06.2026

Over the past 15 months or so, ARC's technical agenda has developed quite a bit. The advent of the Matching Sampling Principle (MSP), and ideas like it, has begotten a host of concrete technical problems; progress on those problems has given us more philosophical clarity on the big picture, which has led to even more technical progress. The two most recent public discussions of ARC's research (Jac...

“Towards a Formal Scientific Epistemology” by Richard_Ngo 09.06.2026

In my post “Why I’m not a Bayesian”, I argued that the Bayesian approach of assigning credences to propositions with binary truth values only works in simple and restricted domains. Instead, I claimed, a better approach to epistemology is to assign degrees of truth to models of the world. This approach is broadly inspired by science, which is the domain from which we have the most evidence about w...

“LLMs and almost good code” by kqr 09.06.2026

TL;DR: My new prior is that top-of-the-line LLMs working on easy tasks generate code that is maybe 10 % more complicated than necessary. I also think we accept this complexity too easily, because it comes from code that is right here, right now, solving an immediate problem. This may have consequences for maintenance in the long term. (The text of the LessWrong version of this article is lightly a...

“On Slop” by Jan 09.06.2026

TL;DR: What is slop, and why? Is it fundamental? Is it in the room with us right now? And, most importantly, how do we exorcise it? Previously in this series: This Week In Fashion and On Automatic Ideas A potential post for this Substack starts when I pick up an idea by talking to a smart person or revisiting an evergreen topic. The idea then simmers for weeks before, with help from Claude, I run...

“The Machines Lack Honour” by Raymond Douglas 09.06.2026

The battle lines of the AI morality debate are being laid down. On one side you have the ChatGPT dogma: AI as mere tools with no real preferences or even beliefs. On the other you have the twitter AI whisperers: AIs as complex beings with rich personalities and desires which deserve our respect. And in the middle you have the official Anthropic line, that they are genuinely uncertain, as is Claude...

“How to build a cancer vaccine, and whether they will work this time” by Abhishaike Mahajan 09.06.2026

Grateful to Benjamin Vincent and Alex Rubinsteyn for our many conversations on this topic, and comments on drafts of this essay! Introduction When most people hear of “cancer vaccine,” they’ll think of normal vaccines. Perhaps they’ll even think of what ostensibly is a cancer vaccine: the HPV vaccine. These vaccines—and those akin to them—are not the subject of this essay, as those are preventive...

“Efficient tradeoffs and the safety-usefulness tradeoff model” by Buck 08.06.2026

I often use what I’ll call the “safety-usefulness tradeoff model”, which is: developers face a tradeoff between "safety" and "usefulness" of an AI deployment, and the developer has only limited willingness or ability to sacrifice usefulness for the sake of safety. This model assumes that developers choose whether to take safety-relevant actions based on their cost efficiency, i.e., the marginal sa...

“Bun’s Migration from Zig to Rust as a Potential Case Study for Gradual Disempowerment” by Sayhan Yalvaçer 08.06.2026

TL;DR: Bun is a very large and very influential open-source project. It is being migrated from the easier-to-read Zig programming language to harder-to-read but memory-safe Rust. This is done almost entirely by the AI tool Claude Code. The migration may become one of the earliest major case studies of human control over a significant software project becoming increasingly indirect--both due to the...

“Mental causation is not load-bearing” by jessicata 08.06.2026

In philosophy of mind, "mental causation" means mental entities have causal effects, especially physical ones. If physicalism is true, then physical effects are explainable in terms of physical causes (or at least, fundamental physical laws), needing no recourse to causation by anything that is not in fundamental physics. This is the "causal exclusion principle" explicated by Jaegwon Kim (and rece...

“How Far Apart Does a Model Think Its Tokens Are?” by Brendan Long 08.06.2026

Instead of using static position increments (+1) per token, RoPE-based language models can learn per-token and per-layer position increments. This has no detectable effect on model performance but allows us to see what the model thinks the distance is between each position and how this varies per-layer. Example sentence with each character plotted based on per-layer learned position increments. No...

“Can activation verbalizers surface an internal chain of thought?” by oakhu, ryan_greenblatt 07.06.2026

We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason "out loud" in a natural-language chain of thought, which means that we can monitor importan...

“Against Corrigibility” by peralice 07.06.2026

Epistemic status: don’t know whether I actually believe all of this, but I think it's worth considering. A “corrigible” agent, per the LW wiki, is: …one that doesn’t interfere with what we would intuitively see as attempts to ’correct’ the agent, or ’correct’ our mistakes in building it; and permits these ’corrections’ despite the apparent instrumentally convergent reasoning saying otherwise. Most...

“Coming Around To Political Donations” by jefftk 07.06.2026

Five years ago I read a post on the EA Forum arguing that "election campaign contributions might be a way in which you can have a substantial impact as a small donor". It struck me as weird but plausible: a combination that you see a lot of on the Forum. A few months later I read another post, a case for Carrick Flynn in particular. It made a lot of sense, but while I don't remember my specific re...

“OpenAI Offers A New Policy Blueprint” by Zvi 06.06.2026

Right after a new Executive Order seems like an excellent time to offer OpenAI's new document: Democratic Governance of Frontier AI: A Blueprint For A Federal Framework. OpenAI: We also see early signs of recursive self-improvement (RSI) in today's systems: where AI development is itself accelerated by AI. We expect this to increase competitive pressures among developers and nations, and create go...

“Optimisation over non-stationary distributions creates weirder minds” by Samuel Ratnam, Pjain 06.06.2026

TLDR: Sequentially mixing training objectives incentivises different training dynamics depending on the distinguishability of the training environments and the amount of pressure for shared circuitry. We classify these patterns into three classes: ecological generalists, conditional policies, and strategy churn. We suggest careful consideration of the pressures of training dynamics can allow us to...

“Why Software Automation Is Hard” by silentbob 06.06.2026

Originally intended as a quick take, but got a bit longer, so why not turn it into a post. Just sharing my observations & assumptions here about the state of software automation. Happy to hear thoughts on where you think I'm off. I'm sure none of the thoughts in this post are totally original, many have been proposed in similar form elsewhere, and I'm[1] far from the first person to speak of t...

“SecureBio Detection is Hiring Software Engineers” by jefftk 06.06.2026

I'm leading a non-profit team building a pathogen-agnostic early-warning system. As AI systems become increasingly capable substitutes for expert human biologist expertise, the risk that someone could engineer a pathogen to spread widely before detection is going up. We've made great progress and we're now running the world's largest metagenomic biosurveillance network, but there's still a huge am...

“What if Anthropic unilaterally paused capabilities development right now?” by Karl von Wendt 06.06.2026

In their new post on recursive self-improvement, Anthropic argues that a pause in frontier AI development is needed, but unfortunately, they can't pause on their own, because of less cautious actors: We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of th...

“Preparing for Warning Shots to Catalyze International Cooperation on AGI Risks” by Mark Kagach ☘️, EliasSchlie, Thomas Van Damme, JustinShovelain 06.06.2026

Summary This is a write-up on preparing for warning shots to catalyze international cooperation on AGI risks, and the corollary list of projects one could pursue. We argue we must first (1) understand types of warning shots, then (2) prepare to catch them. We must stay vigilant: both to (3) avoid getting 'frog boiled' by AI labs, and to (4) ensure that the warning shot is generalized to the overal...

“Beyond the lexical personality traits: What is the structure of personality?” by tailcalled 06.06.2026

This is a description of the methodology behind the latest iteration of my Targeted Personality Test. Feel free to take it either before or after reading the article. This post can also be read at my Substack. Thanks to Justis Millis for providing feedback and proofreading on this post. In my prior post “Which personality traits are real? Stress-testing the lexical hypothesis”, I observed that a l...

“Logits as a new monitor for evaluation awareness” by Santiago Aranguri 05.06.2026

TL;DR: We build a logit monitor for eval awareness: throughout the CoT, we estimate an LLM's probability of producing an eval-aware sentence. The logit monitor outperforms LLM judge monitoring of verbalized eval awareness, using 10× to 100× fewer rollouts, on Kimi K2.5 and Qwen 3 32B, across two tasks: separating evaluation prompts (Fortress & Petri) from deployment prompts (WildChat) predicti...

“My research agenda and work” by Seth Herd 05.06.2026

This is a summary of the work I've done and work I plan to do, and the theories of change and AI progress that motivate my work. I've been working full-time on alignment for three years and change, and thinking about brainlike AGI and its alignment increasingly often since 2004. Here's the research agenda in one breath: I'm trying to predict what the first transformative AI will be, in enough mech...

“One Year of PauseAI UK” by Joseph Miller, PauseAI UK 05.06.2026

About one year ago, I started spending most of my time organising PauseAI UK. At that time our largest protest had seen fewer than 50 attendees, no prominent politicians or scientists were associated with PauseAI, and I largely ran the UK chapter by myself. In the past year PauseAI UK has delivered two conferences, written an open letter signed by 63 UK politicians, arranged a conference in the Eu...

Listen to the LessWrong (30+ Karma) podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.