LessWrong
LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
"Thoughts on the AI Safety Summit company policy requests and responses" by So8res 03.11.2023 21:27
Over the next two days, the UK government is hosting an AI Safety Summit focused on “the safe and responsible development of frontier AI”. They requested that seven companies (Amazon, Anthropic, DeepMind, Inflection, Meta, Microsoft, and OpenAI) “outline their AI Safety Policies across nine areas of AI Safety”. Below, I’ll give my thoughts on the nine areas the UK government described; I’ll note k...
[Human Voice] "Book Review: Going Infinite" by Zvi 31.10.2023 2:39:54
Support ongoing human narrations of curated posts: www.patreon.com/LWCurated Previously: Sadly, FTX I doubted whether it would be a good use of time to read Michael Lewis’s new book Going Infinite about Sam Bankman-Fried (hereafter SBF or Sam). What would I learn that I did not already know? Was Michael Lewis so far in the tank of SBF that the book was filled with nonsense and not to be trusted? I...
"Announcing Timaeus" by Jesse Hoogland et al. 30.10.2023 11:15
Timaeus is a new AI safety research organization dedicated to making fundamental breakthroughs in technical AI alignment using deep ideas from mathematics and the sciences. Currently, we are working on singular learning theory and developmental interpretability . Over time we expect to work on a broader research agenda, and to create understanding-based evals informed by our research. Source: http...
"Thoughts on responsible scaling policies and regulation" by Paul Christiano 30.10.2023 11:28
I am excited about AI developers implementing responsible scaling policies ; I’ve recently been spending time refining this idea and advocating for it. Most people I talk to are excited about RSPs, but there is also some uncertainty and pushback about how they relate to regulation. In this post I’ll explain my views on that: I think that sufficiently good responsible scaling policies could dramati...
"AI as a science, and three obstacles to alignment strategies" by Nate Soares 30.10.2023 18:23
AI used to be a science. In the old days (back when AI didn't work very well), people were attempting to develop a working theory of cognition. Those scientists didn’t succeed, and those days are behind us. For most people working in AI today and dividing up their work hours between tasks, gone is the ambition to understand minds. People working on mechanistic interpretability (and others att...
"Architects of Our Own Demise: We Should Stop Developing AI" by Roko 30.10.2023 6:05
Some brief thoughts at a difficult time in the AI risk debate. Imagine you go back in time to the year 1999 and tell people that in 24 years time, humans will be on the verge of building weakly superhuman AI systems. I remember watching the anime short series The Animatrix at roughly this time, in particular a story called The Second Renaissance I part 2 II part 1 II part 2 . For those who haven&a...
"At 87, Pearl is still able to change his mind" by rotatingpaguro 30.10.2023 9:36
Judea Pearl is a famous researcher, known for Bayesian networks (the standard way of representing Bayesian models), and his statistical formalization of causality. Although he has always been recommended reading here , he's less of a staple compared to, say, Jaynes. So the need to re-introduce him. My purpose here is to highlight a soothing, unexpected show of rationality on his part. One yea...
"We're Not Ready: thoughts on "pausing" and responsible scaling policies" by Holden Karnofsky 30.10.2023 11:48
Views are my own, not Open Philanthropy’s. I am married to the President of Anthropic and have a financial interest in both Anthropic and OpenAI via my spouse. Over the last few months, I’ve spent a lot of my time trying to help out with efforts to get responsible scaling policies adopted. In that context, a number of people have said it would be helpful for me to be publicly explicit about whethe...
[HUMAN VOICE] "Alignment Implications of LLM Successes: a Debate in One Act" by Zack M Davis 23.10.2023 26:18
Support ongoing human narrations of curated posts: www.patreon.com/LWCurated Doomimir : Humanity has made no progress on the alignment problem. Not only do we have no clue how to align a powerful optimizer to our "true" values, we don't even know how to make AI "corrigible"—willing to let us correct it. Meanwhile, capabilities continue to advance by leaps and bounds. All i...
"LoRA Fine-tuning Efficiently Undoes Safety Training from Llama 2-Chat 70B" by Simon Lermen & Jeffrey Ladish. 23.10.2023 33:12
Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Jeffrey Ladish. TL;DR LoRA fine-tuning undoes the safety training of Llama 2-Chat 70B with one GPU and a budget of less than $200. The resulting models [1] maintain helpful capabilities without refusing to fulfill harmful instructions. We show that, if model weights are released, safety...
"Holly Elmore and Rob Miles dialogue on AI Safety Advocacy" by jacobjacob, Robert Miles & Holly_Elmore 23.10.2023 50:12
Holly is an independent AI Pause organizer, which includes organizing protests (like this upcoming one ). Rob is an AI Safety YouTuber . I (jacobjacob) brought them together for this dialogue, because I've been trying to figure out what I should think of AI safety protests, which seems like a possibly quite important intervention; and Rob and Holly seemed like they'd have thoughtful and...
"Labs should be explicit about why they are building AGI" by Peter Barnett 19.10.2023 2:32
Three of the big AI labs say that they care about alignment and that they think misaligned AI poses a potentially existential threat to humanity. These labs continue to try to build AGI. I think this is a very bad idea. The leaders of the big labs are clear that they do not know how to build safe, aligned AGI. The current best plan is to punt the problem to a (different) AI, and hope that can solv...
[HUMAN VOICE] "Sum-threshold attacks" by TsviBT 18.10.2023 21:14
Support ongoing human narrations of curated posts: www.patreon.com/LWCurated How do you affect something far away, a lot, without anyone noticing? (Note: you can safely skip sections. It is also safe to skip the essay entirely, or to read the whole thing backwards if you like.) Source: https://www.lesswrong.com/posts/R3eDrDoX8LisKgGZe/sum-threshold-attacks Narrated for LessWrong by Perrin Walker ....
"Will no one rid me of this turbulent pest?" by Metacelsus 18.10.2023 17:24
Last year, I wrote about the promise of gene drives to wipe out mosquito species and end malaria. In the time since my previous writing, gene drives have still not been used in the wild, and over 600,000 people have died of malaria. Although there are promising new developments such as malaria vaccines , there have also been some pretty bad setbacks (such as mosquitoes and parasites developing res...
"RSPs are pauses done right" by evhub 15.10.2023 12:21
COI: I am a research scientist at Anthropic, where I work on model organisms of misalignment ; I was also involved in the drafting process for Anthropic’s RSP . Prior to joining Anthropic, I was a Research Fellow at MIRI for three years. Thanks to Kate Woolverton, Carson Denison, and Nicholas Schiefer for useful feedback on this post. Recently, there’s been a lot of discussion and advocacy around...
[HUMAN VOICE] "Inside Views, Impostor Syndrome, and the Great LARP" by John Wentworth 15.10.2023 9:45
Patreon to support human narration. (Narrations will remain freely available on this feed, but you can optionally support them if you'd like me to keep making them.) *** Epistemic status: model which I find sometimes useful, and which emphasizes some true things about many parts of the world which common alternative models overlook. Probably not correct in full generality. Consider Yoshua Ben...
"Cohabitive Games so Far" by mako yass 15.10.2023 32:10
A cohabitive game [1] is a partially cooperative, partially competitive multiplayer game that provides an anarchic dojo for development in applied cooperative bargaining, or negotiation. Applied cooperative bargaining isn't currently taught, despite being an infrastructural literacy for peace, trade, democracy or any other form of pluralism. We suffer for that. There are many good board games...
"Announcing MIRI’s new CEO and leadership team" by Gretta Duleba 15.10.2023 6:44
In 2023, MIRI has shifted focus in the direction of broad public communication—see, for example, our recent TED talk , our piece in TIME magazine “ Pausing AI Developments Isn’t Enough. We Need to Shut it All Down ”, and our appearances on various podcasts. While we’re continuing to support various technical research programs at MIRI, this is no longer our top priority, at least for the foreseeabl...
"Comparing Anthropic's Dictionary Learning to Ours" by Robert_AIZI 15.10.2023 8:54
Readers may have noticed many similarities between Anthropic's recent publication Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ( LW post ) and my team's recent publication Sparse Autoencoders Find Highly Interpretable Directions in Language Models ( LW post ). Here I want to compare our techniques and highlight what we did similarly or differently. My hop...
"Towards Monosemanticity: Decomposing Language Models With Dictionary Learning" by Zac Hatfield-Dodds 09.10.2023 4:55
Neural networks are trained on data, not programmed to follow rules. We understand the math of the trained network exactly – each neuron in a neural network performs simple arithmetic – but we don't understand why those mathematical operations result in the behaviors we see. This makes it hard to diagnose failure modes, hard to know how to fix them, and hard to certify that a model is truly s...
"Evaluating the historical value misspecification argument" by Matthew Barnett 09.10.2023 11:22
ETA: I'm not saying that MIRI thought AIs wouldn't understand human values. If there's only one thing you take away from this post, please don't take away that. Recently, many people have talked about whether some of the main MIRI people (Eliezer Yudkowsky, Nate Soares, and Rob Bensinger [1] ) should update on whether value alignment is easier than they thought given that GPT-4...
"Response to Quintin Pope’s Evolution Provides No Evidence For the Sharp Left Turn" by Zvi 09.10.2023 16:44
Response to: Evolution Provides No Evidence For the Sharp Left Turn , due to it winning first prize in The Open Philanthropy Worldviews contest . Quintin’s post is an argument about a key historical reference class and what it tells us about AI. Instead of arguing that the reference makes his point, he is instead arguing that it doesn’t make anyone’s point - that we understand the reasons for hum...
"Announcing Dialogues" by Ben Pace 09.10.2023 7:11
As of today, everyone is able to create a new type of content on LessWrong: Dialogues . In contrast with posts, which are for monologues, and comment sections, which are spaces for everyone to talk to everyone, a dialogue is a space for a few invited people to speak with each other . I'm personally very excited about this as a way for people to produce lots of in-depth explanations of their...
"Thomas Kwa's MIRI research experience" by Thomas Kwa and others 06.10.2023 52:05
Moderator note: the following is a dialogue using LessWrong’s new dialogue feature. The exchange is not completed: new replies might be added continuously, the way a comment thread might work. If you’d also be excited about finding an interlocutor to debate, dialogue, or getting interviewed by: fill in this dialogue matchmaking form . Hi Thomas, I'm quite curious to hear about your research...
"EA Vegan Advocacy is not truthseeking, and it’s everyone’s problem" by Elizabeth 03.10.2023 41:30
Effective altruism prides itself on truthseeking. That pride is justified in the sense that EA is better at truthseeking than most members of its reference category, and unjustified in that it is far from meeting its own standards. We’ve already seen dire consequences of the inability to detect bad actors who deflect investigation into potential problems, but by its nature you can never be sure yo...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.