LessWrong

LessWrong (Curated & Popular)

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

Author

LessWrong

Category

Technology

Podcast website

sites.libsyn.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

"GPTs are Predictors, not Imitators" by Eliezer Yudkowsky 12.04.2023

(Related text posted to Twitter ; this version is edited and has a more advanced final section.) Imagine yourself in a box, trying to predict the next word - assign as much probability mass to the next token as possible - for all the text on the Internet. Koan:  Is this a task whose difficulty caps out as human intelligence, or at the intelligence level of the smartest human who wrote any Internet...

"Discussion with Nate Soares on a key alignment difficulty" by Holden Karnofsky 05.04.2023

https://www.lesswrong.com/posts/iy2o4nQj9DnQD7Yhj/discussion-with-nate-soares-on-a-key-alignment-difficulty Crossposted from the AI Alignment Forum . May contain more technical jargon than usual. In late 2022, Nate Soares gave some feedback on my Cold Takes series on AI risk (shared as drafts at that point), stating that I hadn't discussed what he sees as one of the key difficulties of AI ali...

"A stylized dialogue on John Wentworth's claims about markets and optimization" by Nate Soares 05.04.2023

https://www.lesswrong.com/posts/fJBTRa7m7KnCDdzG5/a-stylized-dialogue-on-john-wentworth-s-claims-about-markets ( This is a stylized version of a real conversation, where the first part happened as part of a public debate between John Wentworth and Eliezer Yudkowsky, and the second part happened between John and me over the following morning. The below is combined, stylized, and written in my own v...

"Deep Deceptiveness" by Nate Soares 05.04.2023

https://www.lesswrong.com/posts/XWwvwytieLtEWaFJX/deep-deceptiveness This post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI ( 1 , 2 ), by DeepMind’s alignment team , or by Anthropic , and I know of no other acknow...

"Losing the root for the tree" by Adam Zerner 28.03.2023

https://www.lesswrong.com/posts/ma7FSEtumkve8czGF/losing-the-root-for-the-tree You know that being healthy is important. And that there's a lot of stuff you could do to improve your health: getting enough sleep, eating well, reducing stress, and exercising, to name a few. There’s various things to hit on when it comes to exercising too. Strength, obviously. But explosiveness is a separate thi...

"There’s no such thing as a tree (phylogenetically)" by Eukaryote 28.03.2023

https://www.lesswrong.com/posts/fRwdkop6tyhi3d22L/there-s-no-such-thing-as-a-tree-phylogenetically This is a linkpost for https://eukaryotewritesblog.com/2021/05/02/theres-no-such-thing-as-a-tree/ [Crossposted from Eukaryote Writes Blog.] So you’ve heard about how fish aren’t a monophyletic group? You’ve heard about carcinization , the process by which ocean arthropods convergently evolve into cra...

"The Onion Test for Personal and Institutional Honesty" by Chana Messinger & Andrew Critch 28.03.2023

https://www.lesswrong.com/posts/nTGEeRSZrfPiJwkEc/the-onion-test-for-personal-and-institutional-honesty [co-written by Chana Messinger and Andrew Critch, Andrew is the originator of the idea] You (or your organization or your mission or your family or etc.) pass the “onion test” for honesty if each layer hides but does not mislead about the information hidden within. When people get to know you be...

"Lies, Damn Lies, and Fabricated Options" by Duncan Sabien 28.03.2023

https://www.lesswrong.com/posts/gNodQGNoPDjztasbh/lies-damn-lies-and-fabricated-options This is an essay about one of those "once you see it, you will see it everywhere" phenomena.  It is a psychological and interpersonal dynamic roughly as common, and almost as destructive, as motte-and-bailey, and at least in my own personal experience it's been quite valuable to have it reified,...

"What failure looks like" by Paul Christiano 28.03.2023

https://www.lesswrong.com/posts/HBxe6wdjxK239zajf/what-failure-looks-like Crossposted from the AI Alignment Forum . May contain more technical jargon than usual. The stereotyped image of AI catastrophe is a powerful, malicious AI system that takes its creators by surprise and quickly achieves a decisive advantage over the rest of humanity. I think this is probably not what failure will look like,...

"Why I think strong general AI is coming soon" by Porby 28.03.2023

https://www.lesswrong.com/posts/K4urTDkBbtNuLivJx/why-i-think-strong-general-ai-is-coming-soon I think there is little time left before someone builds AGI (median ~2030). Once upon a time, I didn't think this. This post attempts to walk through some of the observations and insights that collapsed my estimates. The core ideas are as follows: We've already captured way too much of intellig...

"It Looks Like You’re Trying To Take Over The World" by Gwern 28.03.2023

https://gwern.net/fiction/clippy In A.D. 20XX. Work was beginning. “How are you gentlemen !! ”… (Work. Work never changes; work is always hell.) Specifically, a MoogleBook researcher has gotten a pull request from Reviewer #2 on his new paper in evolutionary search in auto-ML, for error bars on the auto-ML hyperparameter sensitivity like larger batch sizes , because more can be different and there...

"More information about the dangerous capability evaluations we did with GPT-4 and Claude." by Beth Barnes 21.03.2023

https://www.lesswrong.com/posts/4Gt42jX7RiaNaxCwP/more-information-about-the-dangerous-capability-evaluations Crossposted from the AI Alignment Forum . May contain more technical jargon than usual. This is a linkpost for https://evals.alignment.org/blog/2023-03-18-update-on-recent-evals/ [Written for more of a general-public audience than alignment-forum audience. We're working on a more thor...

""Carefully Bootstrapped Alignment" is organizationally hard" by Raemon 21.03.2023

https://www.lesswrong.com/posts/thkAtqoQwN6DtaiGT/carefully-bootstrapped-alignment-is-organizationally-hard In addition to technical challenges, plans to safely develop AI face lots of organizational challenges. If you're running an AI lab, you need a concrete plan for handling that.  In this post, I'll explore some of those issues, using one particular AI plan as an example. I first hea...

"The Parable of the King and the Random Process" by moridinamael 14.03.2023

https://www.lesswrong.com/posts/LzQtrHSYDafXynofq/the-parable-of-the-king-and-the-random-process ~ A Parable of Forecasting Under Model Uncertainty ~ You, the monarch, need to know when the rainy season will begin, in order to properly time the planting of the crops. You have two advisors, Pronto and Eternidad, who you trust exactly equally.  You ask them both: "When will the next heavy rain...

"Enemies vs Malefactors" by Nate Soares 14.03.2023

https://www.lesswrong.com/posts/zidQmfFhMgwFzcHhs/enemies-vs-malefactors Status: some mix of common wisdom (that bears repeating in our particular context), and another deeper point that I mostly failed to communicate. Short version Harmful people often lack explicit malicious intent. It’s worth deploying your social or community defenses against them anyway. I recommend focusing less on intent an...

"The Waluigi Effect (mega-post)" by Cleo Nardo 08.03.2023

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post In this article, I will present a mechanistic explanation of the Waluigi Effect and other bizarre "semiotic" phenomena which arise within large language models such as GPT-3/3.5/4 and their variants (ChatGPT, Sydney, etc). This article will be folklorish to some readers, and profoundly novel to others.

"Acausal normalcy" by Andrew Critch 06.03.2023

https://www.lesswrong.com/posts/3RSq3bfnzuL3sp46J/acausal-normalcy Crossposted from the AI Alignment Forum . May contain more technical jargon than usual. This post is also available on the EA Forum . Summary: Having thought a bunch about acausal trade — and proven some theorems relevant to its feasibility — I believe there do not exist powerful information hazards about it that stand up to clear...

"Please don't throw your mind away" by TsviBT 01.03.2023

https://www.lesswrong.com/posts/RryyWNmJNnLowbhfC/please-don-t-throw-your-mind-away [Warning: the following dialogue contains an incidental spoiler for "Music in Human Evolution" by Kevin Simler . That post is short, good, and worth reading without spoilers, and this post will still be here if you come back later. It's also possible to get the point of this post by skipping the dial...

"Cyborgism" by Nicholas Kees & Janus 15.02.2023

https://www.lesswrong.com/posts/bxt7uCiHam4QXrQAA/cyborgism There is a lot of disagreement and confusion about the feasibility and risks associated with automating alignment research. Some see it as the default path toward building aligned AI, while others expect limited benefit from near term systems, expecting the ability to significantly speed up progress to appear well after misalignment and d...

"Childhoods of exceptional people" by Henrik Karlsson 14.02.2023

https://www.lesswrong.com/posts/CYN7swrefEss4e3Qe/childhoods-of-exceptional-people This is a linkpost for https://escapingflatland.substack.com/p/childhoods Let’s start with one of those insights that are as obvious as they are easy to forget: if you want to master something, you should study the highest achievements of your field. If you want to learn writing, read great writers, etc. But this is...

"What I mean by "alignment is in large part about making cognition aimable at all"" by Nate Soares 13.02.2023

https://www.lesswrong.com/posts/NJYmovr9ZZAyyTBwM/what-i-mean-by-alignment-is-in-large-part-about-making Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. (Epistemic status: attempting to clear up a misunderstanding about points I have attempted to make in the past. This post is not intended as an argument for those points.) I have long said that the lion'...

"On not getting contaminated by the wrong obesity ideas" by Natália Coelho Mendonça 10.02.2023

https://www.lesswrong.com/posts/NRrbJJWnaSorrqvtZ/on-not-getting-contaminated-by-the-wrong-obesity-ideas A Chemical Hunger (a), a series by the authors of the blog Slime Mold Time Mold (SMTM), argues that the obesity epidemic is entirely caused (a) by environmental contaminants.  In my last post, I investigated SMTM’s main suspect (lithium).[1] This post collects other observations I have made abo...

"SolidGoldMagikarp (plus, prompt generation)" 08.02.2023

https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation Work done at SERI-MATS , over the past two months, by Jessica Rumbelow and Matthew Watkins. TL;DR Anomalous tokens: a mysterious failure mode for GPT (which reliably insulted Matthew) We have found a set of anomalous tokens which result in a previously undocumented failure mode for GPT-2 and GPT-3 models. (T...

"Focus on the places where you feel shocked everyone's dropping the ball" by Nate Soares 03.02.2023

https://www.lesswrong.com/posts/Zp6wG5eQFLGWwcG6j/focus-on-the-places-where-you-feel-shocked-everyone-s Writing down something I’ve found myself repeating in different conversations: If you're looking for ways to help with the whole “the world looks pretty doomed” business, here's my advice: look around for places where we're all being total idiots. Look for places where everyone&ap...

"Basics of Rationalist Discourse" by Duncan Sabien 02.02.2023

https://www.lesswrong.com/posts/XPv4sYrKnPzeJASuk/basics-of-rationalist-discourse-1 Introduction This post is meant to be a linkable resource. Its core is a short list of guidelines (you can link directly to the list) that are intended to be fairly straightforward and uncontroversial, for the purpose of nurturing and strengthening a culture of clear thinking, clear communication, and collaborative...

Listen to the LessWrong (Curated & Popular) podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.