zvi

LessWrong posts by zvi

Audio narrations of LessWrong posts by zvi

Koniecznie odwiedź stronę podcastu i wesprzyj twórcę: lesswrong.com

Autor

zvi

Kategoria

Technology

Strona podcastu

lesswrong.com

Ostatni odcinek

10 lip 2026

Gdzie słuchać?

Podcasty w aplikacji Replaio Radio Już wkrótce

Podcasty trafią do aplikacji już wkrótce. Zainstaluj teraz i jako pierwszy zobacz nowe podejście do podcastów

Pobierz z Google Play Zainstaluj za darmo Android 5 mln+ pobrań · ocena 4,8 iOS niedługo

Odcinki

“Selling H200s to China Is Unwise and Unpopular” by Zvi 09.12.2025

AI is the most important thing about the future. It is vital to national security. It will be central to economic, military and strategic supremacy. This is true regardless of what other dangers and opportunities AI might present. The good news is that America has many key advantages in AI. America's greatest advantage in AI is our vastly superior access to compute. We are in danger of selling a l...

“Little Echo” by Zvi 08.12.2025

I believe that we will win. An echo of an old ad for the 2014 US men's World Cup team. It did not win. I was in Berkeley for the 2025 Secular Solstice. We gather to sing and to reflect. The night's theme was the opposite: ‘I don’t think we’re going to make it.’ As in: Sufficiently advanced AI is coming. We don’t know exactly when, or what form it will take, but it is probably coming. When it does,...

“DeepSeek v3.2 Is Okay And Cheap But Slow” by Zvi 05.12.2025

DeepSeek v3.2 is DeepSeek's latest open model release with strong bencharks. Its paper contains some technical innovations that drive down cost. It's a good model by the standards of open models, and very good if you care a lot about price and openness, and if you care less about speed or whether the model is Chinese. It is strongest in mathematics. What it does not appear to be is frontier. It is...

“AI #145: You’ve Got Soul” by Zvi 04.12.2025

The cycle of language model releases is, one at least hopes, now complete. OpenAI gave us GPT-5.1 and GPT-5.1-Codex-Max. xAI gave us Grok 4.1. Google DeepMind gave us Gemini 3 Pro and Nana Banana Pro. Anthropic gave us Claude Opus 4.5. It is the best model, sir. Use it whenever you can. One way Opus 4.5 is unique is that it as what it refers to as a ‘soul document.’ Where OpenAI tries to get GPT-5...

“On Dwarkesh Patel’s Second Interview With Ilya Sutskever” by Zvi 03.12.2025

Some podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level. This was very clearly one of those. So here we go. Double click to interact with video As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. If I am quoting directly I use quote marks, otherwise assume paraphrases. What are the...

“Reward Mismatches in RL Cause Emergent Misalignment” by Zvi 02.12.2025

Learning to do misaligned-coded things anywhere teaches an AI (or a human) to do misaligned-coded things everywhere. So be sure you never, ever teach any mind to do what it sees, in context, as misaligned-coded things. If the optimal solution (as in, the one you most reinforce) to an RL training problem is one that the model perceives as something you wouldn’t want it to do, it will generally lear...

“Claude Opus 4.5 Is The Best Model Available” by Zvi 01.12.2025

Claude Opus 4.5 is the best model currently available. No model since GPT-4 has come close to the level of universal praise that I have seen for Claude Opus 4.5. It is the most intelligent and capable, most aligned and thoughtful model. It is a joy. There are some auxiliary deficits, and areas where other models have specialized, and even with the price cut Opus remains expensive, so it should not...

Claude Opus 4.5: Model Card, Alignment and Safety 28.11.2025

They saved the best for last. The contrast in model cards is stark. Google provided a brief overview of its tests for Gemini 3 Pro, with a lot of ‘we did this test, and we learned a lot from it, and we are not going to tell you the results.’ Anthropic gives us a 150 page book, including their capability assessments. This makes sense. Capability is directly relevant to safety, and also frontier cap...

AI #144: Thanks For the Models 27.11.2025

Thanks for everything. And I do mean everything. Everyone gave us a new model in the last few weeks. OpenAI gave us GPT-5.1 and GPT-5.1-Codex-Max. These are overall improvements, although there are worries around glazing and reintroducing parts of the 4o spirit. xAI gave us Grok 4.1, although few seem to have noticed and I haven’t tried it. Google gave us both by far the best image model in Nana B...

The Big Nonprofits Post 2025 27.11.2025

There remain lots of great charitable giving opportunities out there. I have now had three opportunities to be a recommender for the Survival and Flourishing Fund (SFF). I wrote in detail about my first experience back in 2021, where I struggled to find worthy applications. The second time around in 2024, there was an abundance of worthy causes. In 2025 there were even more high quality applicatio...

The Big Nonprofits Post 2025 26.11.2025

There remain lots of great charitable giving opportunities out there. I have now had three opportunities to be a recommender for the Survival and Flourishing Fund (SFF). I wrote in detail about my first experience back in 2021, where I struggled to find worthy applications. The second time around in 2024, there was an abundance of worthy causes. In 2025 there were even more high quality applicatio...

ChatGPT 5.1 Codex Max 25.11.2025

OpenAI has given us GPT-5.1-Codex-Max, their best coding model for OpenAI Codex. They claim it is faster, more capable and token-efficient and has better persistence on long tasks. It scores 77.9% on SWE-bench-verified, 79.9% on SWE-Lancer-IC SWE and 58.1% on Terminal-Bench 2.0, all substantial gains over GPT-5.1-Codex. It's triggering OpenAI to prepare for being high level in cybersecurity threat...

Gemini 3 Pro Is a Vast Intelligence With No Spine 24.11.2025

It's A Great Model, Sir One might even say the best model. It is for now my default weapon of choice. Google's official announcement of Gemini 3 Pro is full of big talk. Google tells us: Welcome to a new era of intelligence. Learn anything. Build anything. Plan anything. An agent-first development experience in Google Antigravity. Gemini Agent for your browser. It's terrific at everything. They ev...

Gemini 3: Model Card and Safety Framework Report 21.11.2025

Gemini 3 Pro is an excellent model, sir. This is a frontier model release, so we start by analyzing the model card and safety framework report. Then later I’ll look at capabilities. I found the safety framework highly frustrating to read, as it repeatedly ‘hides the football’ and withholds or makes it difficult to understand key information. I do not believe there is a frontier safety problem with...

“AI #143: Everything, Everywhere, All At Once” by Zvi 20.11.2025

Last week had the release of GPT-5.1, which I covered on Tuesday. This week included Gemini 3, Nana Banana Pro, Grok 4.1, GPT 5.1 Pro, GPT 5.1-Codex-Max, Anthropic making a deal with Microsoft and Nvidia, Anthropic disrupting a sophisticated cyberattack operation and what looks like an all-out attack by the White House to force through a full moratorium on and preemption of any state AI laws witho...

“Monthly Roundup #36: November 2025” by Zvi 19.11.2025

Happy Gemini Week to those who celebrate. Coverage of the new release will begin on Friday. Meanwhile, here's this month's things that don’t go anywhere else. Good News, Everyone Google has partnered with Polymarket to include Polymarket odds into Google Search and Google Finance. This is fantastic and suggests we should expand the number of related markets on Polymarket. In many ways Polymarket p...

“On Writing #2” by Zvi 18.11.2025

In honor of my dropping by Inkhaven at Lighthaven in Berkeley this week, I figured it was time for another writing roundup. You can find #1 here, from March 2025. I’ll be there from the 17th (the day I am publishing this) until the morning of Saturday the 22nd. I am happy to meet people, including for things not directly about writing. Table of Contents Table of Contents. How I Use AI For Writing...

“GPT 5.1 Follows Custom Instructions and Glazes” by Zvi 18.11.2025

There are other model releases to get to, but while we gather data on those, first things first. OpenAI has given us GPT-5.1: Same price including in the API, Same intelligence, better mundane utility? Their Announcement Sam Altman (CEO OpenAI): GPT-5.1 is out! It's a nice upgrade. I particularly like the improvements in instruction following, and the adaptive thinking. The intelligence and style...

“AI Craziness: Additional Suicide Lawsuits and The Fate of GPT-4o” by Zvi 14.11.2025

GPT-4o has been a unique problem for a while, and has been at the center of the bulk of mental health incidents involving LLMs that didn’t involve character chatbots. I’ve previously covered related issues in AI Craziness Mitigation Efforts, AI Craziness Notes, GPT-4o Responds to Negative Feedback, GPT-4o Sycophancy Post Mortem and GPT-4o Is An Absurd Sycophant. Discussions of suicides linked to A...

“AI #142: Common Ground” by Zvi 13.11.2025

The Pope offered us wisdom, calling upon us to exercise moral discernment when building AI systems. Some rejected his teachings. We mark this for future reference. The long anticipated Kimi K2 Thinking was finally released. It looks pretty good, but it's too soon to know, and a lot of the usual suspects are strangely quiet. GPT-5.1 was released yesterday. I won’t cover that today beyond noting it...

“The Pope Offers Wisdom” by Zvi 12.11.2025

The Pope is a remarkably wise and helpful man. He offered us some wisdom. Yes, he is generally playing on easy mode by saying straightforwardly true things, but that's meeting the world where it is. You have to start somewhere. Some rejected his teachings. Wisdom Is Offered Two thousand years after Jesus famously got nailed to a cross for suggesting we all be nice to each other for a change, Pope...

“Kimi K2 Thinking” by Zvi 11.11.2025

I previously covered Kimi K2, which now has a new thinking version. As I said at the time back in July, price in that the thinking version is coming. Is it the real deal? That depends on what level counts as the real deal. It's a good model, sir, by all accounts. But there have been fewer accounts than we would expect if it was a big deal, and it doesn’t fall into any of my use cases. Introducing...

“Variously Effective Altruism” by Zvi 10.11.2025

This post is a roundup of various things related to philanthropy, as you often find in the full monthly roundup. Preventing Value Drift Peter Thiel warned Elon Musk to ditch donating to The Giving Pledge because Bill Gates will give his wealth away ‘to left-wing nonprofits.’ As John Arnold points out, this seems highly confused. The Giving Pledge is a promise to give away your money, not a promise...

“On Sam Altman’s Second Conversation with Tyler Cowen” by Zvi 07.11.2025

Some podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level. This was very clearly one of those. So here we go. As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. If I am quoting directly I use quote marks, otherwise assume paraphrases. The entire conversation takes place with an unde...

“AI #141: Give Us The Money” by Zvi 06.11.2025

OpenAI does not waste time. On Friday I covered their announcement that they had ‘completed their recapitalization’ by converting into a PBC, including the potentially largest theft in human history. Then this week their CFO Sarah Friar went ahead and called for a Federal ‘backstop’ on their financing, also known as privatizing gains and socializing losses, also known as the worst form of socialis...

Słuchaj podcastu LessWrong posts by zvi w Replaio

Radio i podcasty w jednej aplikacji - za darmo, bez zakładania konta. Zainstaluj już dziś i nie przegap premiery

Pobierz z Google Play

Replaio nie jest wydawcą podcastów; nazwy audycji, okładki i audio należą do ich autorów i są rozpowszechniane przez publiczne kanały RSS