Sigurd

AI Based Paper Discussions

A RSS feed with papers which are von interest in AI Safety Control and Agent Behavior

Husk at besøge podcastens hjemmeside og støtte skaberen: rss.com

Forfatter

Sigurd

Kategori

Technology

Podcastens hjemmeside

rss.com

Seneste episode

8. apr. 2026

Hvor kan du lytte?

Podcasts i appen Replaio Radio Kommer snart

Podcasts kommer snart til appen. Installer nu, og vær den første til at se en helt ny tilgang til podcasts

Hent den på Google Play Installer gratis Android næsten 10 mio. downloads · 4,8 i bedømmelse iOS snart

Episoder

Anthrophic Mythos 08.04.2026

Anthrophic Mythos

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems 12.03.2026

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Public Finance in the Age of AI 08.03.2026

Public Finance in the Age of AI

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships 08.03.2026

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships

The Generative AI Paradox 08.03.2026

The Generative AI Paradox

Training Agents to Self-Report Misbehavior 08.03.2026

Training Agents to Self-Report Misbehavior

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models 08.03.2026

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems 08.03.2026

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground 08.03.2026

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

RASP Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs 08.03.2026

Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs

Reducing Harmful Generative AI Outputs via Consensus Sampling 08.03.2026

Reducing Harmful Generative AI Outputs via Consensus Sampling

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers 08.03.2026

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Bridging Skill Gaps for the Future: New Jobs Creation in the AI Age 07.03.2026

The demand and supply of new skills—especially in IT and AI—are reshaping labor markets, impacting wages and hiring. About 1 in 10 job vacancies in advanced economies demands at least one new skill, often appearing first in the United States. The incidence is about half of that in emerging market economies. These skills boost average wages and employment but deepen polarization, mostly benefitting...

Reasoning Models Struggle to Control their Chains of Thought 07.03.2026

Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability — CoT controllability — we introduce the CoT-Control evaluation suite, which includes tasks that require models to...

Lyt til podcasten AI Based Paper Discussions i Replaio

Radio og podcasts i én app - gratis og uden tilmelding. Installer i dag, og gå ikke glip af lanceringen

Hent den på Google Play

Replaio er ikke podcastudgiver; programmernes navne, coverbilleder og lyd tilhører deres ophavsmænd og distribueres via offentlige RSS-feeds.