Sigurd

AI Based Paper Discussions

A RSS feed with papers which are von interest in AI Safety Control and Agent Behavior

Vizitează neapărat site-ul podcastului și susține-i creatorul: rss.com

Autor

Sigurd

Categorie

Technology

Site-ul podcastului

rss.com

Cel mai nou episod

8 apr. 2026

Unde asculți?

Podcasturi în aplicație Replaio Radio În curând

Podcasturile ajung în curând în aplicație. Instaleaz-o acum și fii primul care descoperă o abordare complet nouă a podcasturilor

Descarcă din Google Play Instalează gratuit Android aproape 10 mil. descărcări · nota 4,8 iOS în curând

Episoade

Anthrophic Mythos 08.04.2026

Anthrophic Mythos

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems 12.03.2026

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Public Finance in the Age of AI 08.03.2026

Public Finance in the Age of AI

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships 08.03.2026

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships

The Generative AI Paradox 08.03.2026

The Generative AI Paradox

Training Agents to Self-Report Misbehavior 08.03.2026

Training Agents to Self-Report Misbehavior

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models 08.03.2026

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems 08.03.2026

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground 08.03.2026

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

RASP Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs 08.03.2026

Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs

Reducing Harmful Generative AI Outputs via Consensus Sampling 08.03.2026

Reducing Harmful Generative AI Outputs via Consensus Sampling

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers 08.03.2026

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Bridging Skill Gaps for the Future: New Jobs Creation in the AI Age 07.03.2026

The demand and supply of new skills—especially in IT and AI—are reshaping labor markets, impacting wages and hiring. About 1 in 10 job vacancies in advanced economies demands at least one new skill, often appearing first in the United States. The incidence is about half of that in emerging market economies. These skills boost average wages and employment but deepen polarization, mostly benefitting...

Reasoning Models Struggle to Control their Chains of Thought 07.03.2026

Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability — CoT controllability — we introduce the CoT-Control evaluation suite, which includes tasks that require models to...

Ascultă podcastul AI Based Paper Discussions în Replaio

Radio și podcasturi într-o singură aplicație - gratuit și fără înregistrare. Instaleaz-o azi și nu rata lansarea

Descarcă din Google Play

Replaio nu este editorul podcasturilor; numele emisiunilor, coperțile și materialul audio aparțin autorilor lor și sunt distribuite prin fluxuri RSS publice