LessWrong
LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
"How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions" by Jan Brauner et al. 03.10.2023 7:03
Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when instructed to output misinformation. Here, we develop a simple lie detector that requires neither access to the LLM's activations (black-box) nor ground-truth knowledge of the fact in quest...
"The Lighthaven Campus is open for bookings" by Habryka 03.10.2023 5:30
Lightcone Infrastructure (the organization that grew from and houses the LessWrong team) has just finished renovating a 7-building physical campus that we hope to use to make the future of humanity go better than it would otherwise. We're hereby announcing that it is generally available for bookings. We offer preferential pricing for projects we think are good for the world, but to cover oper...
"'Diamondoid bacteria' nanobots: deadly threat or dead-end? A nanotech investigation" by titotal 03.10.2023 37:13
A lot of people are highly concerned that a malevolent AI or insane human will, in the near future, set out to destroy humanity. If such an entity wanted to be absolutely sure they would succeed, what method would they use? Nuclear war? Pandemics? According to some in the x-risk community, the answer is this: The AI will invent molecular nanotechnology, and then kill us all with diamondoid bacteri...
"The King and the Golem" by Richard Ngo 29.09.2023 8:27
This is a linkpost for https://narrativeark.substack.com/p/the-king-and-the-golem Long ago there was a mighty king who had everything in the world that he wanted, except trust. Who could he trust, when anyone around him might scheme for his throne? So he resolved to study the nature of trust, that he might figure out how to gain it. He asked his subjects to bring him the most trustworthy thing in...
"Sparse Autoencoders Find Highly Interpretable Directions in Language Models" by Logan Riggs et al 27.09.2023 10:12
This is a linkpost for Sparse Autoencoders Find Highly Interpretable Directions in Language Models We use a scalable and unsupervised method called Sparse Autoencoders to find interpretable , monosemantic features in real LLMs (Pythia-70M/410M) for both residual stream and MLPs. We showcase monosemantic features, feature replacement for Indirect Object Identification (IOI), and use OpenAI's a...
"Inside Views, Impostor Syndrome, and the Great LARP" by John Wentworth 26.09.2023 8:46
Epistemic status: model which I find sometimes useful, and which emphasizes some true things about many parts of the world which common alternative models overlook. Probably not correct in full generality. Consider Yoshua Bengio, one of the people who won a Turing Award for deep learning research. Looking at his work, he clearly “knows what he’s doing”. He doesn’t know what the answers will be in...
"There should be more AI safety orgs" by Marius Hobbhahn 25.09.2023 29:35
I’m writing this in my own capacity. The views expressed are my own, and should not be taken to represent the views of Apollo Research or any other program I’m involved with. TL;DR: I argue why I think there should be more AI safety orgs. I’ll also provide some suggestions on how that could be achieved. The core argument is that there is a lot of unused talent and I don’t think existing orgs scal...
"The Talk: a brief explanation of sexual dimorphism" by Malmesbury 22.09.2023 30:06
Cross-posted from substack . "Everything in the world is about sex, except sex. Sex is about clonal interference." – Oscar Wilde (kind of) As we all know, sexual reproduction is not about reproduction. Reproduction is easy. If your goal is to fill the world with copies of your genes, all you need is a good DNA-polymerase to duplicate your genome, and then to divide into two copies of yo...
"A Golden Age of Building? Excerpts and lessons from Empire State, Pentagon, Skunk Works and SpaceX" by jacobjacob 20.09.2023 45:43
Patrick Collison has a fantastic list of examples of people quickly accomplishing ambitious things together since the 19th Century. It does make you yearn for a time that feels... different, when the lethargic behemoths of government departments could move at the speed of a racing startup: [...] last century, [the Department of Defense] innovated at a speed that puts modern Silicon Valley startu...
"AI presidents discuss AI alignment agendas" by TurnTrout & Garrett Baker 19.09.2023 23:39
This is a linkpost for https://www.youtube.com/watch?v=02kbWY5mahQ None of the presidents fully represent my (TurnTrout's) views. TurnTrout wrote the script. Garrett Baker helped produce the video after the audio was complete. Thanks to David Udell, Ulisse Mini, Noemi Chulo, and especially Rio Popper for feedback and assistance in writing the script. Source: https://www.lesswrong.com/posts/7M...
"UDT shows that decision theory is more puzzling than ever" by Wei Dai 18.09.2023 2:44
I feel like MIRI perhaps mispositioned FDT (their variant of UDT) as a clear advancement in decision theory, whereas maybe they could have attracted more attention/interest from academic philosophy if the framing was instead that the UDT line of thinking shows that decision theory is just more deeply puzzling than anyone had previously realized. Instead of one major open problem (Newcomb's, o...
"Sum-threshold attacks" by TsviBT 11.09.2023 19:14
How do you affect something far away, a lot, without anyone noticing? (Note: you can safely skip sections. It is also safe to skip the essay entirely, or to read the whole thing backwards if you like.) Source: https://www.lesswrong.com/posts/R3eDrDoX8LisKgGZe/sum-threshold-attacks Narrated for LessWrong by TYPE III AUDIO . Share feedback on this narration. [125+ Karma Post] ✓
"A list of core AI safety problems and how I hope to solve them" by Davidad 09.09.2023 12:06
Context: I sometimes find myself referring back to this tweet and wanted to give it a more permanent home. While I'm at it, I thought I would try to give a concise summary of how each distinct problem would be solved by an Open Agency Architecture (OAA) , if OAA turns out to be feasible. Source: https://www.lesswrong.com/posts/D97xnoRr6BHzo5HvQ/one-minute-every-moment Narrated for LessWrong b...
"Report on Frontier Model Training" by Yafah Edelman 09.09.2023 35:50
This is a linkpost for https://docs.google.com/document/d/1TsYkDYtV6BKiCN9PAOirRAy3TrNDu2XncUZ5UZfaAKA/edit?usp=sharing Understanding what drives the rising capabilities of AI is important for those who work to forecast, regulate, or ensure the safety of AI. Regulations on the export of powerful GPUs need to be informed by understanding of how these GPUs are used, forecasts need to be informed by...
"Defunding My Mistake" by ymeskhout 08.09.2023 11:06
Until about five years ago, I unironically parroted the slogan All Cops Are Bastards (ACAB) and earnestly advocated to abolish the police and prison system. I had faint inklings I might be wrong about this a long time ago, but it took a while to come to terms with its disavowal. What follows is intended to be not just a detailed account of what I used to believe but most pertinently, why . Despite...
"Sharing Information About Nonlinear" by Ben Pace 08.09.2023 56:26
Added (11th Sept): Nonlinear have commented that they intend to write a response , have written a short follow-up , and claim that they dispute 85 claims in this post. I'll link here to that if-and-when it's published. Added (11th Sept): One of the former employees, Chloe, has written a lengthy comment personally detailing some of her experiences working at Nonlinear and the aftermath. A...
"One Minute Every Moment" by abramdemski 08.09.2023 6:12
About how much information are we keeping in working memory at a given moment? "Miller's Law" dictates that the number of things humans can hold in working memory is " the magical number 7 ± 2 ". This idea is derived from Miller's experiments, which tested both random-access memory (where participants must remember call-response pairs, and give the correct response wh...
"What I would do if I wasn’t at ARC Evals" by LawrenceC 08.09.2023 25:08
In which: I list 9 projects that I would work on if I wasn’t busy working on safety standards at ARC Evals, and explain why they might be good to work on. Epistemic status: I’m prioritizing getting this out fast as opposed to writing it carefully. I’ve thought for at least a few hours and talked to a few people I trust about each of the following projects, but I haven’t done that much digging int...
"The U.S. is becoming less stable" by lc 04.09.2023 3:39
We focus so much on arguing over who is at fault in this country that I think sometimes we fail to alert on what's actually happening. I would just like to point out, without attempting to assign blame, that American political institutions appear to be losing common knowledge of their legitimacy, and abandoning certain important traditions of cooperative governance. It would be slightly hyper...
"Meta Questions about Metaphilosophy" by Wei Dai 04.09.2023 5:24
To quickly recap my main intellectual journey so far (omitting a lengthy side trip into cryptography and Cypherpunk land), with the approximate age that I became interested in each topic in parentheses: Source: https://www.lesswrong.com/posts/fJqP9WcnHXBRBeiBg/meta-questions-about-metaphilosophy Narrated for LessWrong by TYPE III AUDIO . Share feedback on this narration. [125+ Karma Post] ✓
"OpenAI API base models are not sycophantic, at any size" by Nostalgebraist 04.09.2023 4:43
In Discovering Language Model Behaviors with Model-Written Evaluations" (Perez et al 2022) , the authors studied language model "sycophancy" - the tendency to agree with a user's stated view when asked a question. The paper contained the striking plot reproduced below, which shows sycophancy increasing dramatically with model size while being largely independent of RLHF steps a...
"Dear Self; we need to talk about ambition" by Elizabeth 30.08.2023 13:02
I keep seeing advice on ambition, aimed at people in college or early in their career, that would have been really bad for me at similar ages. Rather than contribute ( more ) to the list of people giving poorly universalized advice on ambition, I have written a letter to the one person I know my advice is right for: myself in the past. Source: https://www.lesswrong.com/posts/uGDtroD26aLvHSoK2/dear...
"Book Launch: "The Carving of Reality," Best of LessWrong vol. III" by Raemon 28.08.2023 5:35
The Carving of Reality , third volume of the Best of LessWrong books is now available on Amazon (US) . The Carving of Reality includes 43 essays from 29 authors. We've collected the essays into four books, each exploring two related topics. The "two intertwining themes" concept was first inspired when as I looked over the cluster of "coordination" themed posts, and noting...
"Assume Bad Faith" by Zack_M_Davis 28.08.2023 12:20
I've been trying to avoid the terms "good faith" and "bad faith". I'm suspicious that most people who have picked up the phrase "bad faith" from hearing it used, don't actually know what it means—and maybe, that the thing it does mean doesn't carve reality at the joints . People get very touchy about bad faith accusations: they think that you shoul...
"Large Language Models will be Great for Censorship" by Ethan Edwards 23.08.2023 15:20
LLMs can do many incredible things. They can generate unique creative content, carry on long conversations in any number of subjects, complete complex cognitive tasks, and write nearly any argument. More mundanely, they are now the state of the art for boring classification tasks and therefore have the capability to radically upgrade the censorship capacities of authoritarian regimes throughout th...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.