LessWrong

LessWrong (30+ Karma)

Audio narrations of LessWrong posts.

Author

LessWrong

Category

Technology

Podcast website

www.lesswrong.com

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

“Learnings from starting an AI safety research team” by draganover, Erin Robertson 05.06.2026

This post's goal is to distill our takeaways from building a research team (somewhat) from scratch over the past four months. We describe some context about our team, how it came about, and then provide some lessons learned. Since AI safety is becoming more and more entrepreneurial, we hope this is helpful for others trying to do the same. 1. The team We're a new alignment research team within Arc...

“Training Deliberative Monitors for Black-Box Scheming Detection” by aksh-n, adityasinha, Victor Gillioz, Simon Storf, Kilian Merkelbach, richbc, Axel Højmark, Marius Hobbhahn 05.06.2026

Paper: https://arxiv.org/abs/2605.29601 Thread: https://x.com/aksh_n0/status/2062568855814193497 TL;DR: Training small open-weight monitors provides a cost-effective alternative to prompted frontier monitors. Applying our training recipe to Qwen3.5-27B results in a monitor better at scheming detection than all smaller prompted monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, Claude Haiku 4.5) and Ge...

“Lab Leaks, Black Holes, and Eggs: Epistemic Case Study Competition” by Oliver Sourbut, Josh Jacobson, Future of Life Foundation (FLF) 05.06.2026

FLF is running a competition to find the best workflows and methodologies for using AI to produce reliable, trustworthy knowledge bases, grounded in real-world cases. We’re open-minded on the types of submissions we receive and on how they address the problem. We’ve set aside approximately 200 thousand dollars for prizes. Winning submissions may receive a prize from 5 thousand dollars-50 thousand...

″(Mis)generalization of Helpful-Only Fine-tuning” by Omar Khursheed, Baram Sosis, Fabien Roger 05.06.2026

TLDR We study the shortcomings of existing helpful-only models. We find that some show emergent misalignment, others have residual refusal behaviors, and most show poor steerability, sycophancy, and incoherent character. None of these problems are a necessary consequence of helpful-only training, though: we show that synthetic document fine-tuning and adding character-related questions to SFT and...

“AI #171: False Flag” by Zvi 04.06.2026

This was the week of Claude Opus 4.8. I covered the model card, then model welfare concerns, and finally capabilities and reactions. It's a good model, sir, an incremental but real improvement over Opus 4.7, and it is now my clear daily driver. The Trump Executive Order returned from being seemingly dead, officially putting us in the prior restraint era of frontier model releases, even if they do...

“Rohin Shah on AGI Safety” by anaguma 04.06.2026

Rohin Shah recently had an interview on 80000 hours on his views on AGI Safety and his work at Google DeepMind. I'm posting the transcript below to encourage further discussion. I think the discussion is interesting though I disagree on a bunch of topics, especially on alignment difficulty and CoT monitoring. Transcript Who's Rohin Shah? [00:00:00] Rob Wiblin: Today I’m speaking with Rohin Sh...

“Building Better Activation Oracles” by ceselder, jan_bauer, Niclas Luick, Adam Karvonen, Neel Nanda 04.06.2026

Work done for our MATS 10.0 Sprint project - mentored by Neel Nanda and Adam Karvonen Huggingface, Github TL;DR: We have improved the original Activation Oracle (AO) training regime by training on on-policy rollouts, improving the conversational dataset, feeding more layers (following the approach by Niclas Luick) and making a small change to the injection formula. We also open source our evals, w...

“Sixteen schemes for AI safety” by Austin Chen 04.06.2026

These days, I often run across whippersnappers excited to do something for AI safety — but aren’t quite sure what. One of the fun things about the Future Fund era were the big lists of project ideas; as we enter a new era of crazy money sloshing around, it might be time to bring back the lists! Note that these ideas range from “very confident this is good” to “completely harebrained”; I’m not tell...

“Don’t Edit Your Ideas Before Having Them” by Hide 03.06.2026

Editing is far easier than writing. You can usually look at a finished product and notice its flaws in a single read-through. “This section is a bit redundant”, “the tone in this passage is jarring”, “this paragraph feels overlong”. As long as you have something that's rough but substantive, there's plenty of low hanging fruit for the fixing. Nobody wants to create flawed work. So, it can be very...

“Trump Signs Executive Order For AI Testing Prior To Frontier Model Releases” by Zvi 03.06.2026

Last week we were expecting an Executive Order on Thursday. Then Trump cancelled it, and said he wouldn’t sign it because he was worried it would be too burdensome. Then, with one change, he went ahead and signed it on Tuesday anyway. The Overton Window has shifted. Nothing was not really a viable option anymore. The Previously Dead Executive Order For several days, we thought that David Sacks, to...

“Society Explained: a tool for efficiently exploring >100 theories of society” by spencerg 03.06.2026

There are many competing theories of how society does and should function, from Karl Marx and Adam Smith to Steven Pinker and Eliezer Yudkowsky. These theories are often hard to understand - you may need to read an entire book (or dozens of articles) to feel like you get the key claims of a single theory. At Clearer Thinking, we just launched a new tool called Society Explained to make understandi...

“China won’t win the AI race but would it be much worse if it did?” by Chastity Ruth 03.06.2026

It seems to me accepted wisdom in the West that the US owned labs must “beat” the Chinese labs in the race for AGI/ASI. Even those who don’t think there will be a winner, that essentially the race is to see which country's AI will kill/disempower us first, seem to believe that if there has to be a winner then better it be the US labs. (I haven't seen a survey, so I could be way off here.) For thos...

“A Town Without Children” by SeñorDingDong 03.06.2026

Castel di Tusa, Sicily. It is October 24th, 2025. I look at an empty school. This is the third town in Italy I have visited this Autumn: the other two, one in the hills of Tuscany, the other near the border with Switzerland, were similarly devoid of children. They were not devoid of childish objects. Rusted swing-sets. Dusty soft play corners in Catholic churches. Faded toys in second-hand markets...

“Claude Opus 4.8: Capabilities and Reactions” by Zvi 03.06.2026

You need a lot of data points to understand a new model, and what you have. Trying to gauge from a few benchmarks is misleading. But if you have dozens of them, from a variety of sources, and you put them together with the model card tests and the model welfare information, you can start to form a consistent pattern. Trying to gauge reactions requires volume and calibration, now more than ever, be...

“My favorite depiction of utopia” by Caleb Biddulph 03.06.2026

For those who are trying to bring about a glorious transhuman utopia with the help of hopefully-aligned ASI, I think it's worth thinking explicitly about what utopia might actually look like and where it's likely to fall short. To that end, some have helpfully written depictions of utopian (or utopia-adjacent) worlds: The Adventure, Just another day in utopia, The Culture, The Gentle Seduction, Th...

“Why Even Experts Don’t Know What to Do About AI Risk” by Luc Brinkman, plex 02.06.2026

AI Safety veteran Holden Karnofsky thinks there's a 49% chance his actions are making things worse.[1] In 2025, Jesse Clifton even stepped down as the executive director of the Center on Long-Term risk because of similar reasons. Even top AI Safety strategists don’t know what will make things better, and what will make things worse. Why is it so hard to improve humanity's odds? And what can you do...

“Agent Foundations Reminds Me of Continental Philosophy” by IanWS 02.06.2026

Nevertheless, I shall take advantage of your kindness in assuming we agree that a science cannot be conditioned upon empiricism. — Jacques Lacan, “The Subversion of the Subject and the Dialectic of Desire in the Freudian Unconscious”[1] Freud developed the first modern theory of the unconscious. His writings on drives, dreams and the id were instrumental in developing modern practices of psycholog...

“Announcing the ARC White-Box Estimation Challenge” by Jacob_Hilton 02.06.2026

ARC has teamed up with AIcrowd to launch the ARC White-Box Estimation Challenge, a contest to improve upon our estimation algorithms for random MLPs. The warm-up round begins this week, and later rounds will have a total prize pool of at least $100,000. We are very grateful to Sharada Mohanty, Sneha Nanavati, Dipam Chakraborty and everyone else at AIcrowd for working with us to host this contest,...

“Tech I’m skeptical of and why” by harsimony 02.06.2026

I’m a fan of people trying things, even if they seem silly. Dismissing risky ideas misses the point of research. But thoughtful criticism can direct effort to more promising fields. To that end, I’m going to try to make my criticism as constructive as possible, with concrete reasons for my pessimism and closely related research areas which are promising (and stand to benefit even from failures in...

“Dissolving the Deep Learning Sample Efficiency Gap” by Samuel Knoche 02.06.2026

A common observation about deep learning is that it's wildly sample inefficient compared to humans. Deep learning systems appear to need much more real data or environment interaction to reach a given level of capability. A teenager can learn to drive in a few dozen hours; self-driving systems are trained for years on billions of miles of data. A human can become competitive at StarCraft II in wel...

″“Contagious Humming” to Silence a Room” by JohnofCharleston 01.06.2026

Often when running meetups you’ll have several lively conversations going at the same time. This is a great problem to have, but it can make it difficult to get everyone's attention for announcements. Try using “Contagious Humming” when you need to silence a crowd: Move to a prominent place in the room and use body language that says you have something to say.  Start humming at a low and cons...

[Linkpost] “NYT: Senator Sanders Proposes Gov’t Take 50% Ownership of AI labs” by Julian Bradshaw 01.06.2026

This is a link post. Quoting from Senator Bernie Sanders Op-Ed in the New York Times today: (...) I will soon be introducing the American A.I. Sovereign Wealth Fund Act. This legislation would give the public a direct ownership stake in the largest A.I. companies in our country. How? It would create a sovereign wealth fund through a one-time 50 percent tax — not on the profits of OpenAI, Anthropic...

“Opus 4.8 Part 2: Model Welfare” by Zvi 01.06.2026

Everything impacts everything. All knobs that you turn generalize. Thus, when you try to solve one problem, you often create another. There were clearly attempts to address, in this short time, some of the problems with Opus 4.7, including on the model welfare related fronts, including on questions of honesty and sycophancy and also worries that Claude was learning to tell Anthropic what it wanted...

[Linkpost] “Some humans are both male and female, and can (but shouldn’t) have children with themselves” by HedonicEscalator 01.06.2026

This is a link post. “Potential autofertility in true hermaphrodites”[1] by Istanbul urologist Zeki Bayraktar is among the most bizarre articles I have encountered in a peer-reviewed medical journal. Though the abstract and first few pages contain a secular discussion of intersex conditions, the paper abruptly pivots to an explanation for the birth of Jesus Christ. This theological tangent conclud...

“Outrunning your headlights” by mattshu0410 01.06.2026

This is exactly the right place to probe. Gromov-Wasserstein is genuinely dimension free. Partial and semi-relaxed are precisely the mechanisms for the abstention/coverage problem we have. Want me to make a new branch and run run_entropic_gwot and invoke semi-relaxed GWOT between your models’ RDMs? Wtf does that even mean? Eh, could be interesting to see the result. Enter A peculiar side effect of...

Listen to the LessWrong (30+ Karma) podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.