Shaoqing Tan
Embodied AI 101
Stay in the loop on research in AI and physical intelligence.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
PhysBrain 1.0 VLA (TwinBrainVLA): Dual-Brain Vision-Language-Action with Physics-Grounded Learning 18.05.2026 25:59
Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.
MolmoAct2-LIBERO: An Open Vision-Language-Action Model for Robotics 17.05.2026 38:51
Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on manipulation tasks. Released with both checkpoint and dataset for VLA finetuning.
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Diffusion Transformers 17.05.2026 20:29
A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a Hybrid Linear Diffusion Transformer + Gated DeltaNet for long-context efficiency. Targets controllable physics simulation.
WildClawBench: A Real-World, Long-Horizon Benchmark for AI Agents 17.05.2026 32:09
New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluation protocols for cross-embodiment policies.
MCP-Cosmos: Bring Your Own World Model 17.05.2026 24:21
Introduces a latent-space world model framework that lets agents simulate state transitions and iteratively refine plans before real-world execution. Evaluated on 20+ MCP-Bench tasks with measurable gains in tool-use success.
OpenAI o1: Teaching LLMs to Think Slow and Deep 17.05.2026 14:07
Details OpenAI's reasoning-focused o1 model and its 'long thought' approach using test-time compute scaling. Explores how extended reasoning during inference can improve model performance on complex tasks.
The Llama 3 Herd of Models 17.05.2026 32:41
Comprehensive technical report on the Llama 3 family, covering architecture, training at scale, multimodal extensions, and real-world impact. Details the development of Meta's flagship open-source language model series.
LATENT: Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data 17.05.2026 31:29
Introduces a three-stage pipeline that extracts a latent action space from low-quality human tennis demonstrations, then trains a high-level policy in simulation via reinforcement learning. Enables dynamic whole-body humanoid tennis play with back-and-forth volleys at human level.
AnyFlow: Any-Step Video Diffusion for Predictive World Modeling 14.05.2026 13:33
First any-step video diffusion framework using flow maps, allowing a single model to adapt to arbitrary inference budgets for scalable high-quality video generation relevant to predictive world modeling.
# Robotics: The Endgame 14.05.2026 34:16
Technical roadmap mirroring LLM scaling: critiques VLAs, advocates video world models as second pretraining phase, introduces World Action Models (WAM), manipulation data flywheels, EgoScale with new Dexterity Scaling Law, and DreamDojo end-to-end neural physics engine for sim RL.
Claw-Eval: Toward Trustworthy and Transparent Evaluation of Autonomous Agents 08.04.2026 28:34
Benchmark with 2,159 rubric items across 300 tasks using trajectory-aware grading and 3-trial Pass^3 scoring to mitigate luck. Evaluates agent reliability in real-world robotics settings.
LIBERO-Para: Paraphrase Robustness in Robotic Manipulation 08.04.2026 32:19
Reveals paraphrase fragility in VLAs causing 22-52% success drops due to task misidentification. Introduces PRIDE metric weighting success by paraphrase difficulty on LIBERO benchmark manipulation tasks.
YOR: Your Own Mobile Manipulator for Generalizable Robotics 07.04.2026 27:08
Low-cost mobile manipulator design and training strategies for broad generalization in real-world tasks.
EgoSim: Egocentric World Simulator for Embodied Interaction Generation 07.04.2026 50:52
Closed-loop egocentric video simulator maintaining persistent 3D scene state for consistent interactions, enabling cross-embodiment transfer from human videos to robotic manipulation.
Accelerating Video World Models: From Generative Videos to Real-Time Simulators 07.04.2026 39:27
Comprehensive survey taxonomizing efficient architectures/algorithms for video world models as simulators, targeting compute bottlenecks in embodied AI, autonomous driving, and games with techniques like short-window attention for real-time long-horizon prediction.
From Tokens to Thoughts: Continuous Latent Reasoning in Large Models and Robot Control 07.04.2026 26:57
Curated collection of 100+ works surveying shift to continuous latent spaces in LLMs/VLMs/VLAs for improved reasoning over discrete tokens, with relevance to robotics action modeling.
CaP-X: Coding Agents for Physical eXecution 06.04.2026 13:48
CaP-X is an open-source agentic robotics framework where LLMs/VLMs generate code to call perception and control APIs for execution across diverse simulated and real robots in CaP-Gym's 187 manipulation tasks. The framework includes CaP-Bench for evaluating frontier models and CaP-RL, which boosts a 7B model's success from 20% to 72% with minimal sim-to-real gap.
DoRA: Weight-Decomposed Low-Rank Adaptation 06.04.2026 39:19
An upgrade over LoRA for parameter-efficient fine-tuning, enabling better performance in LLMs by decomposing weights into magnitude and direction components.
AI Model Collapse: What Happens When AI Trains on Its Own Outputs 06.04.2026 29:25
Seminal work showing how training on AI-generated data leads to 'model collapse' in neural networks, with urgent implications for future scaling.
PhAIL: Benchmarking Vision-Language-Action Models on Real-World Bin-Picking 05.04.2026 33:23
Real-world hardware evaluation of VLAs on blind bin-to-bin picking, achieving max 64 picks/hour across hundreds of runs, with full videos/data exposing gaps in production-scale robotic manipulation reliability.
Co-training Large Behavior Models: Data Modalities and Training Strategies for Robot Manipulation 05.04.2026 28:22
Comprehensive evaluation of 89 policies showing optimal co-training practices mixing real robot data with sim/egocentric human videos to boost diversity and performance in large robotics foundation models.
HyDRA: Hybrid Memory for Dynamic Video World Models 05.04.2026 21:53
Novel memory system preserving dynamic object identity and motion continuity across occlusions in video world models, addressing frozen/vanishing issues for improved predictive physics in embodied AI.
# WildWorld: Dynamic World Modeling with Actions and Explicit State 04.04.2026 32:44
Massive dataset enabling dynamic world models with explicit states and actions, supporting predictive modeling for cross-embodiment robotic control.
Omni-WorldBench: Evaluating Interactive 4D World Models 04.04.2026 39:56
New benchmark assessing world models on interaction tasks, pushing predictive physics and video modeling towards robotics applications with action-conditioned evaluation.
SIMART: From Static Meshes to Sim-Ready Articulated Models 04.04.2026 38:09
Unified MLLM framework with Sparse 3D VQ-VAE (70% token reduction) for part-level mesh decomposition and kinematic chain prediction, enabling physics-based robotic simulation from monolithic assets.
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.