AI SafeGuard

AI Safety Breakthrough

Technology EN ↓ 9 episodes

The future of AI is in our hands. Join AI SafeGuard on "AI Safety Breakthrough" as we explore the frontiers of AI safety research and discuss how we can ensure a future where AI remains beneficial for everyone. We delve into the latest breakthroughs, uncover potential risks, and empower listeners to become informed participants in the conversation about AI's role in society. Subscribe now and become part of the solution! Intro about the author J, graduated from Carnegie Mellon University, School of Computer Science, 10+ years in Cybersecurity, Cyber Threat Intelligence, Risk, Compliance, priva...

Author

AI SafeGuard

Category

Technology

Podcast website

media.rss.com

Latest episode

Aug 13, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Navigating the New AI Security 13.08.2025

Welcome to Agentic AI Unlocked , your deep dive into the transformative world of Agentic AI —systems combining large language models with advanced reasoning and autonomous action. These intelligent agents promise to disrupt industries, yet introduce a fundamentally new threat surface . Risks like memory poisoning, tool misuse, prompt injection, and insider threats highlight the urgent need for rob...

DeepSeek: A Disruptive Force in AI 03.02.2025

This episode explores DeepSeek , a Chinese AI startup challenging the AI landscape with its free alternative to ChatGPT. We'll examine DeepSeek's innovative architecture, including Mixture-of-Experts (MoE) and Multi-head Latent Attention (MLA) , which optimize efficiency. The discussion will highlight DeepSeek's use of reinforcement learning (RL) and its impact on reasoning capabilities, as well a...

VLSBench: A Visual Leakless Multimodal Safety Benchmark 26.01.2025

Are current AI safety benchmarks for multimodal models flawed? This podcast explores the groundbreaking research behind VLSBench , a new benchmark designed to address a critical flaw in existing safety evaluations: visual safety information leakage (VSIL) We delve into how sensitive information in images is often unintentionally revealed in the accompanying text prompts, allowing models to identif...

Adaptive Stress Testing for Language Model Toxicity 20.01.2025

This episode explores ASTPrompter , a novel approach to automated red-teaming for large language models (LLMs). Unlike traditional methods that focus on simply triggering toxic outputs, ASTPrompter is designed to discover likely toxic prompts – those that could naturally emerge during regular language model use. The approach uses Adaptive Stress Testing (AST) , a technique that identifies likely f...

Global Responsible AI Maturity: A Survey of 1000 Organizations 16.01.2025

This episode dives into the critical topic of Responsible AI (RAI), exploring how organizations worldwide are grappling with the ethical and practical challenges of AI adoption. We'll be drawing insights from a comprehensive survey of 1000 organizations across 20 industries and 19 geographical regions

Ivy-VL: A Lightweight Multimodal Model for Everyday Devices 09.12.2024

In this episode, we dive into Ivy-VL , a groundbreaking lightweight multimodal AI model released by AI Safeguard in collaboration with Carnegie Mellon University (CMU) and Stanford University . With only 3 billion parameters, Ivy-VL processes both image and text inputs to generate text outputs, offering an optimal balance of performance, speed, and efficiency. Its compact design supports deploymen...

Agent Bench: Evaluating LLMs as Agents 27.11.2024

Large Language Models (LLMs) are rapidly evolving, but how do we assess their ability to act as agents in complex, real-world scenarios? Join Jenny as we explore Agent Bench, a new benchmark designed to evaluate LLMs in diverse environments, from operating systems to digital card games. We'll delve into the key findings, including the strengths and weaknesses of different LLMs and the challenges o...

Hacking AI for Good: Open AI’s Red Teaming Approach 24.11.2024

In this podcast, we delve into OpenAI's innovative approach to enhancing AI safety through red teaming—a structured process that uses both human expertise and automated systems to identify potential risks in AI models. We explore how OpenAI collaborates with external experts to test frontier models and employs automated methods to scale the discovery of model vulnerabilities. Join Jenny as we disc...

Surgical Precision: PKE’s Role in AI Safety 24.11.2024

Explore how Precision Knowledge Editing (PKE) refines AI for safety and ethical behavior in Surgical Precision: PKE’s Role in AI Safety . Join experts as we uncover the science, challenges, and breakthroughs shaping trustworthy AI. Perfect for tech enthusiasts and professionals alike, this podcast reveals how PKE ensures AI serves humanity responsibly.

Listen to the AI Safety Breakthrough podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.