Center for AI Safety

AI Safety Newsletter

Narrations of the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This podcast also contains narrations of some of our publications. ABOUT USThe Center for AI Safety (CAIS) is a San Francisco-based research and field-building nonprofit. We believe that artificial intelligence has the potential to profoundly benefit the world, provided that we can develop and use it safely. However, in contrast to the dramatic progress in AI, many basic problems in AI safety have yet to be solved. Our mission is to reduce societal-...

Author

Center for AI Safety

Category

Technology

Podcast website

newsletter.safe.ai

Latest episode

Jul 6, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

AISN #29: Progress on the EU AI Act 04.01.2024

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. A Provisional Agreement on the EU AI Act On December 8th, the EU Parliament, Council, and Commission reached a provisional agreement on the EU AI Act. The agreement regulates the deployment of AI in high risk applications such as hiring and credit pricing,...

The Landscape of US AI Legislation 29.12.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This week we’re looking closely at AI legislative efforts in the United States, including: Senator Schumer's AI Insight Forum The Blumenthal-Hawley framework for AI governance Agencies proposed to govern digital platforms State and local laws against AI su...

AISN #28: Center for AI Safety 2023 Year in Review 21.12.2023

As 2023 comes to a close, we want to thank you for your continued support for AI safety. This has been a big year for AI and for the Center for AI Safety. In this special-edition newsletter, we highlight some of our most important projects from the year. Thank you for being part of our community and our work. Center for AI Safety's 2023 Year in Review The Center for AI Safety (CAIS) is on a missio...

AISN #27: Defensive Accelerationism 07.12.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Defensive Accelerationism  Vitalik Buterin, the creator of Ethereum, recently wrote an essay on the risks and opportunities of AI and other technologies. He responds to Marc Andreessen's manifesto on techno-optimism and the growth of the effective acc...

AISN #26: National Institutions for AI Safety 15.11.2023

Also, Results From the UK Summit, and New Releases From OpenAI and xAI. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This week's key stories include: The UK, US, and Singapore have announced national AI safety institutions. The UK AI Safety Summit concluded with a consensus statement, the creation of...

AISN #25: White House Executive Order on AI, UK AI Safety Summit, and Progress on Voluntary Evaluations of AI Risks. 31.10.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. White House Executive Order on AI While Congress has not voted on significant AI legislation this year, the White House has left their mark on AI policy. In June, they secured voluntary commitments on safety from leading AI companies. Now, the White House...

AISN #24: Kissinger Urges US-China Cooperation on AI, China’s New AI Law, US Export Controls, International Institutions, and Open Source AI. 18.10.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. China's New AI Law, US Export Controls, and Calls for Bilateral Cooperation China details how AI providers can fulfill their legal obligations. The Chinese government has passed several laws on AI. They’ve regulated recommendation algorithms and taken step...

AISN #23: New OpenAI Models, News from Anthropic, and Representation Engineering. 04.10.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. OpenAI releases GPT-4 with Vision and DALL·E-3, announces Red Teaming Network GPT-4 with vision and voice. When GPT-4 was initially announced in March, OpenAI demonstrated its ability to process and discuss images such as diagrams or photographs. This feat...

AISN #21: Google DeepMind’s GPT-4 Competitor, Military Investments in Autonomous Drones, The UK AI Safety Summit, and Case Studies in AI Policy. 05.09.2023

Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Google DeepMind’s GPT-4 Competitor Computational power is a key driver of AI progress, and a new report suggests that Google’s upcoming GPT-4 competitor will be trained on unprecedented amounts of compute. The model, currently named Gemini, may be trained...

AISN #20: LLM Proliferation, AI Deception, and Continuing Drivers of AI Capabilities. 29.08.2023

AI Deception: Examples, Risks, Solutions AI deception is the topic of a new paper from researchers at and affiliated with the Center for AI Safety. It surveys empirical examples of AI deception, then explores societal risks and potential solutions. The paper defines deception as “the systematic production of false beliefs in others as a means to accomplish some outcome other than the truth.” Impor...

[Paper] “An Overview of Catastrophic AI Risks” by Dan Hendrycks, Mantas Mazeika and Thomas Woodside 21.08.2023

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose catastrophic risks. Although numerous risks have been detailed separately, there is a pressing need for a systematic discussion and illustration of the potential dangers to better inform efforts to mitig...

[Paper] “Unsolved Problems in ML Safety” by Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt 21.08.2023

Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the tec...

[Paper] “X-Risk Analysis for AI Research” by Dan Hendrycks and Mantas Mazeika 21.08.2023

Artificial intelligence (AI) has the potential to greatly improve society, but as with any powerful technology, it comes with heightened risks and responsibilities. Current AI research lacks a systematic discussion of how to manage long-tail risks from AI systems, including speculative long-term risks. Keeping in mind the potential benefits of AI, there is some concern that building ever more inte...

AISN #19: US-China Competition on AI Chips, Measuring Language Agent Developments, Economic Analysis of Language Model Propaganda, and White House AI Cyber Challenge. 15.08.2023

US-China Competition on AI Chips Modern AI systems are trained on advanced computer chips which are designed and fabricated by only a handful of companies in the world. The US and China have been competing for access to these chips for years. Last October, the Biden administration partnered with international allies to severely limit China’s access to leading AI chips. Recently, there have been se...

AISN #18: Challenges of Reinforcement Learning from Human Feedback, Microsoft’s Security Breach, and Conceptual Research on AI Safety. 08.08.2023

Challenges of Reinforcement Learning from Human Feedback If you’ve used ChatGPT, you might’ve noticed the “thumbs up” and “thumbs down” buttons next to each of its answers. Pressing these buttons provides data that OpenAI uses to improve their models through a technique called reinforcement learning from human feedback (RLHF). RLHF is popular for teaching models about human preferences, but it fac...

AISN #17: Automatically Circumventing LLM Guardrails, the Frontier Model Forum, and Senate Hearing on AI Oversight. 01.08.2023

Automatically Circumventing LLM Guardrails Large language models (LLMs) can generate hazardous information, such as step-by-step instructions on how to create a pandemic pathogen. To combat the risk of malicious use, companies typically build safety guardrails intended to prevent LLMs from misbehaving. But these safety controls are almost useless against a new attack developed by researchers at Ca...

AISN #16: White House Secures Voluntary Commitments from Leading AI Labs, and Lessons from Oppenheimer . 25.07.2023

White House Unveils Voluntary Commitments to AI Safety from Leading AI Labs Last Friday, the White House announced a series of voluntary commitments from seven of the world's premier AI labs. Amazon, Anthropic, Google, Inflection, Meta, Microsoft, and OpenAI pledged to uphold these commitments, which are non-binding and pertain only to forthcoming "frontier models" superior to currently available...

AISN #15: China and the US take action to regulate AI, results from a tournament forecasting AI risk, updates on xAI’s plan, and Meta releases its open-source and commercially available Llama 2. 19.07.2023

Both China and the US take action to regulate AI Last week, regulators in both China and the US took aim at generative AI services. These actions show that China and the US are both concerned with AI safety. Hopefully, this is a sign they can eventually coordinate. China’s new generative AI rules On Thursday, China’s government released new rules governing generative AI. China’s new rules, which a...

AISN #14: OpenAI’s ‘Superalignment’ team, Musk’s xAI launches, and developments in military AI use . 12.07.2023

OpenAI announces a ‘superalignment’ team On July 5th, OpenAI announced the ‘Superalignment’ team: a new research team given the goal of aligning superintelligence, and armed with 20% of OpenAI’s compute. In this story, we’ll explain and discuss the team’s strategy. What is superintelligence? In their announcement, OpenAI distinguishes between ‘artificial general intelligence’ and ‘superintelligenc...

AISN #13: An interdisciplinary perspective on AI proxy failures, new competitors to ChatGPT, and prompting language models to misbehave. 05.07.2023

Interdisciplinary Perspective on AI Proxy Failures In this story, we discuss a recent paper on why proxy goals fail. First, we introduce proxy gaming, and then summarize the paper’s findings. Proxy gaming is a well-documented failure mode in AI safety. For example, social media platforms use AI systems to recommend content to users. These systems are sometimes built to maximize the amount of time...

AISN #12: Policy Proposals from NTIA’s Request for Comment, and Reconsidering Instrumental Convergence. 27.06.2023

Policy Proposals from NTIA’s Request for Comment The National Telecommunications and Information Administration publicly requested comments on the matter from academics, think tanks, industry leaders, and concerned citizens. They asked 34 questions and received more than 1,400 responses on how to govern AI for the public benefit. This week, we cover some of the most promising proposals found in th...

AISN #11: An Overview of Catastrophic AI Risks. 22.06.2023

An Overview of Catastrophic AI Risks Global leaders are concerned that artificial intelligence could pose catastrophic risks. 42% of CEOs polled at the Yale CEO Summit agree that AI could destroy humanity in five to ten years. The Secretary General of the United Nations said we “must take these warnings seriously.” Amid all these frightening polls and public statements, there’s a simple question t...

AISN #10: How AI could enable bioterrorism, and policymakers continue to focus on AI . 13.06.2023

How AI could enable bioterrorism Only a hundred years ago, no person could have single handedly destroyed humanity. Nuclear weapons changed this situation, giving the power of global annihilation to a small handful of nations with powerful militaries. Now, thanks to advances in biotechnology and AI, a much larger group of people could have the power to create a global catastrophe. This is the upsh...

AISN #9: Statement on Extinction Risks, Competitive Pressures, and When Will AI Reach Human-Level? . 06.06.2023

Top Scientists Warn of Extinction Risks from AI Last week, hundreds of AI scientists and notable public figures signed a public statement on AI risks written by the Center for AI Safety. The statement reads: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” The statement was signed by a broad, diverse coa...

AISN #8: Why AI could go rogue, how to screen for AI risks, and grants for research on democratic governance of AI. 30.05.2023

Yoshua Bengio makes the case for rogue AI AI systems pose a variety of different risks. Renowned AI scientist Yoshua Bengio recently argued for one particularly concerning possibility: that advanced AI agents could pursue goals in conflict with human values. Human intelligence has accomplished impressive feats, from flying to the moon to building nuclear weapons. But Bengio argues that across a ra...

Listen to the AI Safety Newsletter podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.