LessWrong

LessWrong (Curated & Popular)

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

Author

LessWrong

Category

Technology

Podcast website

sites.libsyn.com

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

“When is a mind me?” by Rob Bensinger 08.07.2024

xlr8harder writes: In general I don’t think an uploaded mind is you, but rather a copy. But one thought experiment makes me question this. A Ship of Theseus concept where individual neurons are replaced one at a time with a nanotechnological functional equivalent. Are you still you? Presumably the question xlr8harder cares about here isn't semantic question of how linguistic communities use t...

“80,000 hours should remove OpenAI from the Job Board (and similar orgs should do similarly)” by Raemon 04.07.2024

I haven't shared this post with other relevant parties – my experience has been that private discussion of this sort of thing is more paralyzing than helpful. I might change my mind in the resulting discussion, but, I prefer that discussion to be public. I think 80,000 hours should remove OpenAI from its job board, and similar EA job placement services should do the same. (I personally believ...

[Linkpost] “introduction to cancer vaccines” by bhauth 02.07.2024

This is a linkpost for https://www.bhauth.com/blog/biology/cancer%20vaccines.html cancer neoantigens For cells to become cancerous, they must have mutations that cause uncontrolled replication and mutations that prevent that uncontrolled replication from causing apoptosis. Because cancer requires several mutations, it often begins with damage to mutation-preventing mechanisms. As such, cancers oft...

“Priors and Prejudice” by MathiasKB 02.07.2024

I Imagine an alternate version of the Effective Altruism movement, whose early influences came from socialist intellectual communities such as the Fabian Society, as opposed to the rationalist diaspora. Let's name this hypothetical movement the Effective Samaritans. Like the EA movement of today, they believe in doing as much good as possible, whatever this means. They began by evaluating exi...

“My experience using financial commitments to overcome akrasia” by William Howard 02.07.2024

About a year ago I decided to try using one of those apps where you tie your goals to some kind of financial penalty. The specific one I tried is Forfeit, which I liked the look of because it's relatively simple, you set single tasks which you have to verify you have completed with a photo. I’m generally pretty sceptical of productivity systems, tools for thought, mindset shifts, life hacks a...

“The Incredible Fentanyl-Detecting Machine” by sarahconstantin 01.07.2024

An NII machine in Nogales, AZ. (Image source)There's bound to be a lot of discussion of the Biden-Trump presidential debates last night, but I want to skip all the political prognostication and talk about the real issue: fentanyl-detecting machines. Joe Biden says: And I wanted to make sure we use the machinery that can detect fentanyl, these big machines that roll over everything that comes...

“AI catastrophes and rogue deployments” by Buck 01.07.2024

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.[Thanks to Aryan Bhatt, Ansh Radhakrishnan, Adam Kaufman, Vivek Hebbar, Hanna Gabor, Justis Mills, Aaron Scher, Max Nadeau, Ryan Greenblatt, Peter Barnett, Fabien Roger, and various people at a presentation of these arguments for comments. These ideas aren’t very original to me; many of the examples of threat mod...

“Loving a world you don’t trust” by Joe Carlsmith 01.07.2024

(Cross-posted from my website. Audio version here, or search for "Joe Carlsmith Audio" on your podcast app.) This is the final essay in a series that I'm calling "Otherness andcontrol in the age of AGI." I'm hoping that the individual essays can beread fairly well on their own, butsee here fora brief summary of the series as a whole. There's also a PDF of the who...

“Formal verification, heuristic explanations and surprise accounting” by paulfchristiano 27.06.2024

ARC's current research focus can be thought of as trying to combine mechanistic interpretability and formal verification. If we had a deep understanding of what was going on inside a neural network, we would hope to be able to use that understanding to verify that the network was not going to behave dangerously in unforeseen situations. ARC is attempting to perform this kind of verification,...

“LLM Generality is a Timeline Crux” by eggsyntax 25.06.2024

Summary Summary . LLMs may be fundamentally incapable of fully general reasoning, and if so, short timelines are less plausible. Longer summary There is ML research suggesting that LLMs fail badly on attempts at general reasoning, such as planning problems, scheduling, and attempts to solve novel visual puzzles. This post provides a brief introduction to that research, and asks: Whether this limit...

“SAE feature geometry is outside the superposition hypothesis” by jake_mendel 25.06.2024

Summary: Superposition-based interpretations of neural network activation spaces are incomplete. The specific locations of feature vectors contain crucial structural information beyond superposition, as seen in circular arrangements of day-of-the-week features and in the rich structures. We don’t currently have good concepts for talking about this structure in feature geometry, but it is likely ve...

“Connecting the Dots: LLMs can Infer & Verbalize Latent Structure from Training Data” by Johannes Treutlein, Owain_Evans 23.06.2024

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. This is a link post. TL;DR: We published a new paper on out-of-context reasoning in LLMs. We show that LLMs can infer latent information from training data and use this information for downstream tasks, without any in-context learning or CoT. For instance, we finetune GPT-3.5 on pairs (x,f(x)) for some unknown f...

“Boycott OpenAI” by PeterMcCluskey 21.06.2024

This is a link post. I have canceled my OpenAI subscription in protest over OpenAI's lack ofethics. In particular, I object to: threats to confiscate departing employees' equity unless thoseemployees signed a life-long non-disparagement contract Sam Altman's pattern of lying about important topics I'm trying to hold AI companies to higher standards than I use fortypical compani...

“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison 20.06.2024

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. This is a link post. New Anthropic model organisms research paper led by Carson Denison from the Alignment Stress-Testing Team demonstrating that large language models can generalize zero-shot from simple reward-hacks (sycophancy) to more complex reward tampering (subterfuge). Our results suggest that accidental...

“I would have shit in that alley, too” by Declan Molony 18.06.2024

After living in a suburb for most of my life, when I moved to a major U.S. city the first thing I noticed was the feces. At first I assumed it was dog poop, but my naivety didn’t last long. One day I saw a homeless man waddling towards me at a fast speed while holding his ass cheeks. He turned into an alley and took a shit. As I passed him, there was a moment where our eyes met. He sheepishly aver...

“Getting 50% (SoTA) on ARC-AGI with GPT-4o” by ryan_greenblatt 18.06.2024

ARC-AGI post Getting 50% (SoTA) on ARC-AGI with GPT-4o I recently got to 50%[1] accuracy on the public test set for ARC-AGI by having GPT-4o generate a huge number of Python implementations of the transformation rule (around 8,000 per problem) and then selecting among these implementations based on correctness of the Python programs on the examples (if this is confusing, go here)[2]. I use a varie...

“Why I don’t believe in the placebo effect” by transhumanist_atom_understander 15.06.2024

Have you heard this before? In clinical trials, medicines have to be compared to a placebo to separate the effect of the medicine from the psychological effect of taking the drug. The patient's belief in the power of the medicine has a strong effect on its own. In fact, for some drugs such as antidepressants, the psychological effect of taking a pill is larger than the effect of the drug. It...

“Safety isn’t safety without a social model (or: dispelling the myth of per se technical safety)” by Andrew_Critch 14.06.2024

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. As an AI researcher who wants to do technical work that helps humanity, there is a strong drive to find a research area that is definitely helpful somehow, so that you don’t have to worry about how your work will be applied, and thus you don’t have to worry about things like corporate ethics or geopolitics to ma...

“My AI Model Delta Compared To Christiano” by johnswentworth 13.06.2024

Preamble: Delta vs Crux This section is redundant if you already read My AI Model Delta Compared To Yudkowsky. I don’t natively think in terms of cruxes. But there's a similar concept which is more natural for me, which I’ll call a delta. Imagine that you and I each model the world (or some part of it) as implementing some program. Very oversimplified example: if I learn that e.g. it's c...

“My AI Model Delta Compared To Yudkowsky” by johnswentworth 10.06.2024

Preamble: Delta vs Crux I don’t natively think in terms of cruxes. But there's a similar concept which is more natural for me, which I’ll call a delta. Imagine that you and I each model the world (or some part of it) as implementing some program. Very oversimplified example: if I learn that e.g. it's cloudy today, that means the “weather” variable in my program at a particular time[1] ta...

“Response to Aschenbrenner’s ‘Situational Awareness’” by Rob Bensinger 07.06.2024

(Cross-posted from Twitter.) My take on Leopold Aschenbrenner's new report: I think Leopold gets it right on a bunch of important counts. Three that I especially care about: Full AGI and ASI soon. (I think his arguments for this have a lot of holes, but he gets the basic point that superintelligence looks 5 or 15 years off rather than 50+.) This technology is an overwhelmingly huge deal, and...

“Humming is not a free $100 bill” by Elizabeth 07.06.2024

Last month I posted about humming as a cheap and convenient way to flood your nose with nitric oxide (NO), a known antiviral. Alas, the economists were right, and the benefits were much smaller than I estimated. The post contained one obvious error and one complication. Both were caught by Thomas Kwa, for which he has my gratitude. When he initially pointed out the error I awarded him a $50 bounty...

“Announcing ILIAD — Theoretical AI Alignment Conference ” by Nora_Ammann, Alexander Gietelink Oldenziel 06.06.2024

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. We are pleased to announce ILIAD — a 5-day conference bringing together 100+ researchers to build strong scientific foundations for AI alignment. ***Apply to attend by June 30!*** When: Aug 28 - Sep 3, 2024 Where: @Lighthaven (Berkeley, US) What: A mix of topic-specific tracks, and unconference style programming...

“Non-Disparagement Canaries for OpenAI” by aysja, Adam Scholl 31.05.2024

Since at least 2017, OpenAI has asked departing employees to sign offboarding agreements which legally bind them to permanently—that is, for the rest of their lives—refrain from criticizing OpenAI, or from otherwise taking any actions which might damage its finances or reputation.[1] If they refused to sign, OpenAI threatened to take back (or make unsellable) all of their already-vested equity—a h...

“MIRI 2024 Communications Strategy” by Gretta Duleba 30.05.2024

As we explained in our MIRI 2024 Mission and Strategy update, MIRI has pivoted to prioritize policy, communications, and technical governance research over technical alignment research. This follow-up post goes into detail about our communications strategy. The Objective: Shut it Down[1] Our objective is to convince major powers to shut down the development of frontier AI systems worldwide before...

Listen to the LessWrong (Curated & Popular) podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.