Craig Spencer Smith
Eye on AI Weekly Research Watch
Weekly, digestible podcast explainers of significant research papers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages 23.06.2026 3:01
Code generation benchmarks have become central to how the AI community measures progress, but nearly all of them default to Python — a language that dominates training data and may be inflating model scores. Real software engineering, however, demands fluency across Rust, Go, Java, TypeScript, and many others. Multi-LCB extends the established LiveCodeBench framework to twelve languages while pres...
FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS 23.06.2026 2:22
Even the best text-to-speech systems stumble on proper nouns — a product name pronounced wrong in a voice assistant, or a person's name mangled by a navigation system, can undermine trust immediately. Retraining a full TTS model to fix these errors is expensive and slow. FlowEdit offers an elegant alternative: when a correction is provided, it stores a targeted adjustment in an associative memory...
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes 23.06.2026 3:08
Giving AI agents the ability to modify cloud infrastructure, databases, or deployment pipelines introduces a dangerous gap: a model that reasons incorrectly or gets manipulated could execute irreversible, high-impact actions. Existing security frameworks authorize identities, but they do not enforce that a specific certified action plan is what actually gets executed. The Sovereign Execution Broke...
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm 23.06.2026 2:50
Satellite radar imagery sees through clouds and darkness, making synthetic aperture radar (SAR) indispensable for disaster response, military surveillance, agricultural monitoring, and climate research. Yet multimodal AI research has largely been built on optical imagery because aligned, richly annotated SAR datasets have been scarce. SARLO-80 closes that gap, offering over 119,000 matched SAR-opt...
DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs 23.06.2026 3:19
Standard machine learning tells us what is likely given what we observe, but many real-world decisions demand something more: understanding what would have happened under different circumstances. Counterfactual reasoning is essential for fairness auditing, policy evaluation, and causal explanation. DeepSWIP extends DeepProbLog — a framework blending neural perception with logical reasoning — to su...
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents 23.06.2026 3:27
Customer service agents powered by language models must juggle multiple responsibilities simultaneously: tracking conversation state, calling external tools, and obeying domain-specific policies — all without losing their place. Current architectures bury all of this in a flat prompt, forcing the model to reconstruct context from scratch on every turn. LedgerAgent introduces a dedicated state ledg...
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech 23.06.2026 2:55
Voice interfaces are increasingly governed by natural language instructions — a user might request speech that sounds "warm and conversational" or "brisk and authoritative." But when a text-to-speech system fails to capture that nuance, diagnosing the problem is largely guesswork. This paper borrows the DAAM attribution framework from image generation and applies it to speech diffusion models for...
Toward Calibrated Mixture-of-Experts Under Distribution Shift 23.06.2026 3:22
When a model says it is 80% confident, it should be right about 80% of the time — that is calibration, and it matters enormously in high-stakes settings like medicine, finance, and autonomous systems. Mixture-of-experts architectures, which route inputs to specialized sub-models, have shown strong performance gains, but their calibration behavior under real-world distribution shift has been poorly...
Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation 23.06.2026 2:31
Recommendation systems quietly shape what billions of people watch, buy, and read. The latest frontier in this space is generative recommendation, which frames next-item prediction as a generation problem rather than a retrieval one. A core challenge is representing user behavior richly enough for a generative model to reason over it without drowning in noise or computational cost. G2Rec addresses...
How Transparent is DiffusionGemma? 23.06.2026 2:59
As AI systems take on more consequential roles, understanding how they reason has become as important as what they produce. Diffusion-based language models like DiffusionGemma represent a departure from traditional autoregressive generation, performing much of their computation in a continuous latent space rather than producing tokens step by step. This raises a pressing question: does that shift...
VISTA: View-Consistent Self-Verified Training for GUI Grounding 15.06.2026 2:38
Teaching AI to click the right button on a screen — GUI grounding — sounds simple but is surprisingly brittle. A core training problem is that reinforcement learning often collapses: on hard instances, every rollout fails, so there's no useful learning signal; on easy ones, every rollout succeeds, equally uninformative. VISTA solves this by generating multiple crops of the same GUI screenshot, com...
CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation 15.06.2026 2:23
High-throughput scientific experimentation — screening thousands of chemical compounds, for instance — is expensive and irreversible, making it a dangerous domain for unconstrained AI autonomy. CARE solves this by keeping a proven non-LLM optimizer as the default while allowing an LLM to propose challenger strategies, only authorizing the challenger when pre-outcome evidence actually supports the...
A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems 15.06.2026 2:38
Railway networks are extraordinarily complex — trains of different gauges share limited track, single-track sections require precise coordination, and unexpected disruptions cascade through entire timetables. Most optimization research stops at high-level scheduling, leaving the messy operational details — track switching, gauge compatibility, disruption response — to human operators under pressur...
Sensitivity Shaping for Latent Modeling 15.06.2026 2:50
Generative dynamics models let robots plan behavior in rich, uncertain environments — but safely deploying them requires reliably detecting when the robot is about to enter unfamiliar territory. Existing out-of-distribution detection methods bolt on detectors after the fact, and this paper shows why that fails: if the dynamics model is locally insensitive to different control inputs in critical re...
When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime 15.06.2026 2:33
Most AI failure research is theoretical or laboratory-based — this paper is a rare longitudinal postmortem of a real production LLM agent system running continuously since early 2026, with 22 documented incidents over eight weeks. The most dangerous failure class identified is "fail-plausible": the agent doesn't just fail to report an error, it transforms the error into fluent, convincing narrativ...
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models 15.06.2026 2:41
Audio AI models have gotten good at recognizing what they hear, but complex reasoning — understanding causation, context, and implication across sound, speech, and music — remains a frontier challenge. A key bottleneck is training data: existing datasets are highly redundant, meaning models see many acoustically similar samples that provide overlapping rather than additive learning signal. AudioDE...
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source 15.06.2026 2:46
AI agents can now autonomously plan changes, edit code, and submit pull requests — but open-source infrastructure was built around the assumption of a legally accountable human contributor who can attest to provenance and answer reviewers' questions. This paper systematically maps how six major open-source organizations (including Apache, Linux Foundation, and SymPy) have responded with contributi...
A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health 15.06.2026 2:42
Wearables generate a continuous stream of behavioral data — steps, screen time, sleep — that could power truly proactive health interventions, but it's been unclear which AI architectures best handle these signals across diverse populations and time horizons. This study benchmarks six deep learning models plus two foundation models across 800+ participants, tracking forecast accuracy out to eight...
Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts 15.06.2026 2:17
Predicting how long a patient will survive — and what risks they face — is one of medicine's most consequential tasks, yet most deep learning survival models treat all patients with a single shared representation that can obscure critical subgroup differences. AdaCSM addresses this with a Mixture-of-Experts framework that dynamically routes patients to specialized risk predictors while simultaneou...
Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms 15.06.2026 3:09
What if a musical masterpiece wasn't just art, but also an accidental blueprint for machine learning architectures? This paper argues — through computational analysis of entropy, dissonance, and self-similarity — that the three movements of Beethoven's Moonlight Sonata structurally instantiate streaming, recurrent, and positional encoding memory architectures respectively. The same pitch class acq...
When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks 15.06.2026 2:39
Self-improving AI — where a model uses a verifier to generate its own training feedback — sounds like a path to perpetual improvement, but this paper shows it can silently make models worse. The key problem is task specificity: a verifier that accurately scores math problems may perform near-randomly on multi-disciplinary reasoning, and when it does, it feeds the learner confidently wrong preferen...
From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing 15.06.2026 2:03
Voice synthesis technology has advanced to the point where synthetic speech is nearly indistinguishable from genuine recordings — a serious problem for voice authentication, call centers, and media verification. This paper transforms a self-supervised speech model into a Mixture-of-Experts architecture, where different specialist networks learn complementary acoustic cues for detecting spoofing. E...
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models 15.06.2026 3:05
Automatic speech recognition models like Whisper are impressively accurate, but when they fail — or when accountability matters — we rarely know why they made a particular decision. LEAF-X introduces a principled explainability framework that uses entropy patterns in attention heads to identify which audio frames most influenced a transcription. It produces sparser, more faithful attributions than...
Abstracting Cross-Domain Action Sequences into Interpretable Workflows 15.06.2026 2:48
Every click, tab switch, and file save is a data point — but raw interaction logs are too noisy and granular to reveal how people actually work. WorkflowView uses large language models to convert low-level behavioral logs into high-level activity descriptions, achieving strong semantic accuracy in a zero-shot setting. Tested across browser logs, online learning platforms, and Microsoft Word usage...
Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications 15.06.2026 2:40
Cameras aren't just optical devices — they're mechanical ones too, and sound can make them vibrate. This paper demonstrates that audible sound frequencies can resonate commercially available cameras, introducing artifacts that fool AI vision systems like YOLO into misclassifying objects, missing targets, or hallucinating things that aren't there. Unlike prior ultrasonic attacks limited to short ra...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.