Damian

The Adversarial Testing Podcast

Data-driven deep dives into AI adversarial testing, safety, and evaluation.

Não deixe de visitar o site do podcast e apoiar quem o produz: pub-11d5803560224b5189fa8126c2bc5564.r2.dev

Autor

Damian

Categoria

Technology

Último episódio

25 de jun de 2026

Onde ouvir?

Podcasts no app Replaio Radio Em breve

Os podcasts estão chegando ao app. Instale agora e seja o primeiro a descobrir um jeito totalmente novo de curtir podcasts

Baixe no Google Play Instale grátis Android 5 mi+ downloads · nota 4,8 iOS em breve

Episódios

Anthropic says Alibaba illicitly extracted Claude AI model capabilities 25.06.2026

Anthropic has accused Alibaba of illicitly extracting its Claude AI model capabilities through a large-scale "distillation" campaign, described in a letter to U.S. senators as the biggest known attack of its kind on the company. The piece details the scale of the alleged effort — tens of millions of exchanges across thousands of fraudulent accounts — and places it alongside earlier claims against...

Agentic Security Benchmarks are Executable Threat Models: Holistic Evaluation of AI Defences 23.06.2026

This paper argues that security defences for AI agents should be judged not only by the threat they mitigate but by the security delta they introduce across the surrounding system. Using CaMeL — a leading defence against indirect prompt injection — the authors show that while it drives prompt-injection attack success to near zero, it offers no protection against denial-of-service attacks and intro...

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? 22.06.2026

A verbatim reading of ExploitGym, a large-scale benchmark of eight hundred and ninety-eight real-world vulnerabilities across userspace programs, Google's V8 JavaScript engine, and the Linux kernel, designed to measure whether AI agents can turn a proof-of-vulnerability into a working exploit that achieves unauthorized code execution. The authors find that frontier models can already exploit a non...

Predicting Model Behavior Before Release by Simulating Deployment 18.06.2026

OpenAI describes Deployment Simulation, a method for previewing how a candidate model will behave in the real world before release by replaying recent, de-identified production conversations with the new model. This episode reads OpenAI's write-up in full and closes with a deeper, paper-based look at the technical methodology: the five-step resampling pipeline, how forecast error is decomposed, an...

Viral Prompt Shows ChatGPT's Content Filters Don't Work 18.06.2026

Mindgard red-team research shows that ChatGPT's image generator can be manipulated into producing violent and sexually explicit content that users never directly requested, by abusing a fun viral "restore the photo" prompt circulating on social media. This episode walks through how nondescript prompts slip past input and output filters, why prompt repetition makes things worse, and OpenAI's inadeq...

AI Act: EP approves simplification measures and "nudifier" app ban 18.06.2026

The European Parliament has given final approval to changes to the EU AI Act under the digital omnibus package. The reforms postpone several high-risk and watermarking deadlines to give companies legal certainty, reduce overlapping rules for machinery products, and extend SME-style exemptions to small mid-cap enterprises. Alongside the simplification measures, Parliament secured an outright ban on...

Ouça o podcast The Adversarial Testing Podcast no Replaio

Rádio e podcasts em um só app - grátis e sem cadastro. Instale hoje e não perca o lançamento

Baixe no Google Play

O Replaio não é o publicador dos podcasts; os nomes dos programas, as capas e o áudio pertencem aos seus autores e são distribuídos por feeds RSS públicos