uk

Sigurd

AI Based Paper Discussions

Technology EN ↓ Епізодів: 16

A RSS feed with papers which are von interest in AI Safety Control and Agent Behavior

Обов'язково відвідайте сайт подкасту та підтримайте його автора: rss.com

Автор

Sigurd

Категорія

Technology

Сайт подкасту

rss.com

Останній епізод

8 кві 2026

Де слухати?

Подкасти в застосунку Replaio Radio Уже незабаром

Подкасти незабаром з'являться в застосунку. Встановіть уже зараз і першими побачте зовсім новий погляд на подкасти

Завантажити з Google Play Встановіть безкоштовно Android майже 10 млн завантажень · рейтинг 4,8 iOS незабаром

Епізоди

Anthrophic Mythos 08.04.2026

Anthrophic Mythos

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems 12.03.2026

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Public Finance in the Age of AI 08.03.2026

Public Finance in the Age of AI

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships 08.03.2026

Neural Steering Vectors Reveal Dose and Exposure-Dependent Impacts of Human-AI Relationships

The Generative AI Paradox 08.03.2026

The Generative AI Paradox

Training Agents to Self-Report Misbehavior 08.03.2026

Training Agents to Self-Report Misbehavior

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models 08.03.2026

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems 08.03.2026

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground 08.03.2026

The Generative AI Paradox: How Synthetic Realities Erode Shared Epistemic Ground

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

Evaluating Frontier Models for Stealth and Situational Awareness 08.03.2026

Evaluating Frontier Models for Stealth and Situational Awareness

RASP Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs 08.03.2026

Discovering Interpretable Algorithms by Decompiling Transformers into Human-Readable Programs

Reducing Harmful Generative AI Outputs via Consensus Sampling 08.03.2026

Reducing Harmful Generative AI Outputs via Consensus Sampling

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers 08.03.2026

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Bridging Skill Gaps for the Future: New Jobs Creation in the AI Age 07.03.2026

The demand and supply of new skills—especially in IT and AI—are reshaping labor markets, impacting wages and hiring. About 1 in 10 job vacancies in advanced economies demands at least one new skill, often appearing first in the United States. The incidence is about half of that in emerging market economies. These skills boost average wages and employment but deepen polarization, mostly benefitting...

Reasoning Models Struggle to Control their Chains of Thought 07.03.2026

Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability — CoT controllability — we introduce the CoT-Control evaluation suite, which includes tasks that require models to...

Слухайте подкаст AI Based Paper Discussions у Replaio

Радіо та подкасти в одному застосунку - безкоштовно й без реєстрації. Встановіть уже сьогодні та не пропустіть запуск

Завантажити з Google Play

Replaio не є видавцем подкастів; назви шоу, обкладинки та аудіо належать їхнім авторам і поширюються через публічні RSS-канали