Shaoqing Tan
Embodied AI 101
Stay in the loop on research in AI and physical intelligence.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
EgoSim: An Egocentric World Simulator for Embodied Interaction 04.04.2026 36:17
Closed-loop egocentric simulator persistently updating 3D scene state to generate spatially consistent interaction videos for continuous simulation, enabling cross-embodiment transfer from human videos to robotic manipulation tasks.
Digit's New Motor Cortex: Sim-to-Real RL for Whole-Body Control 03.04.2026 31:13
AI-trained capabilities for new whole-body motions using mocap/teleop data and sim-to-real reinforcement learning, deployable overnight on hardware.
EgoNav: Diffusion-Based Humanoid Navigation from Human Egocentric Video 03.04.2026 42:08
Diffusion-based humanoid navigation trained solely on 5 hours of human egocentric video data, enabling zero-shot deployment on Unitree G1 for complex behaviors like handling glass walls, crowds, and dynamic obstacles via 360° visual memory and hybrid trajectory sampling; upcoming release of dataset, models, and code.
CaP-X: A Code-as-Policy Framework for Robot Manipulation 03.04.2026 13:56
Comprehensive open-source agentic robotics framework treating VLMs/LLMs as code-generating APIs for perception (SAM3, Molmo) and control (IK, grasping), with CaP-Gym benchmark of 187 diverse manipulation tasks (tabletop, bimanual, mobile; sim/real) and CaP-Bench evaluating 12 frontier models; demonstrates rapid RL gains (7B model from 20% to 72% success) with strong sim-to-real transfer.
Embodied Intelligence Breakthrough: Generalist AI’s GEN-1 Robots 02.04.2026 15:27
We've created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where previous models achieve 64%, completes tasks roughly 3x faster than state of the art, and requires only 1 hour of robot data for each of these r...
CaP-X: LMs' First Physical Exam 02.04.2026 22:06
A novel benchmark that evaluates language models on physical examination tasks, testing their ability to understand and perform clinical physical exam procedures in simulated environments. This work introduces a comprehensive evaluation framework for AI systems in medical/clinical settings.
AI Model Collapse: The Danger of Training on AI-Generated Data 31.03.2026 31:31
Demonstrated that LLMs trained recursively on AI-generated data suffer model collapse, a degenerative process where they lose grasp of true data distributions. Sparked critical debates on data provenance and the importance of preserving human-generated training data.
High-Level Automated Reasoning with Qwen2.5-7B 31.03.2026 27:45
Qwen2.5-7B achieved 79.6% on MATH benchmark, surpassing GPT-4o, by employing atomic reasoning actions combined with Monte Carlo Tree Search. Demonstrated that strategic reasoning architectures can enable smaller models to outperform much larger ones.
Co-Training Large Behavior Models: Multimodal Data for Robot Manipulation 31.03.2026 33:09
Explores data modalities and co-training strategies to enhance large behavior models (foundation models) for improved performance in robot manipulation tasks, supporting end-to-end learning and cross-embodiment generalization.
HyDRA: Hybrid Memory for Dynamic Video World Models 30.03.2026 35:58
Memory architecture preserving identity and motion continuity for out-of-view dynamic subjects, addressing frozen/vanishing issues in video world models.
DexWM: Leveraging Human Videos for Dexterous Robot World Models 30.03.2026 31:23
Dataset of robot trajectories designed for training world models to learn dexterous hand-object interactions directly from human videos.
World Models in Robotics 29.03.2026 26:51
Technical survey categorizing world models into action-conditioned, video-inverse dynamics, and joint world-action models (WAMs), discussing their generalization, video data leverage, and trends for closing the robotics data gap.
SIMART: Decomposing Monolithic Meshes into Sim-Ready Articulated Assets 28.03.2026 45:11
Unified MLLM framework with Sparse 3D VQ-VAE that reduces tokens by 70% for efficient part-level decomposition and kinematic prediction in physics-based robotic simulations.
LeWorldModel: A Stable JEPA World Model from Pixels 28.03.2026 13:56
Stable end-to-end JEPA world model trained directly from pixels using simple MSE prediction loss and SIGReg anti-collapse regularization, enabling efficient latent planning under 1 second on 15M params with emergent spatial structure outperforming prior methods.
World Models for Robots: The Next Big Leap? 27.03.2026 20:21
Technical overview defining world models in robotics, their potential to solve diverse problems via video prediction, and key enablers like scale.
Harnessing Long-Running AI in Embodied Systems 27.03.2026 27:10
As AI moves from quick Q&A to marathon tasks, designers grapple with continuity. This episode explores how Anthropics harness design principles translate to embodied AI - robots that need to maintain context across long-running missions.
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations 26.03.2026 17:00
Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception directly from egocentric human demonstrations without teleoperation.
TurboQuant: Redefining AI Efficiency with Extreme Compression 26.03.2026 20:25
This episode explores TurboQuant, a revolutionary set of quantization algorithms from Google Research that redefines AI efficiency through extreme compression. We dive deep into how TurboQuant addresses one of AI's most pressing challenges: the memory bottleneck created by high-dimensional vectors in key-value caches. The research introduces theoretically grounded quantization methods that enable...
DexWM: Learning Dexterous Object Manipulation from Human Videos 25.03.2026 32:14
Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging Face.
FlashAttention-3: Fast & Accurate Attention with Asynchrony & Low-Precision 25.03.2026 17:26
Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.
When AI Trains on Its Own Output: The Model Collapse Problem 25.03.2026 24:41
Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influential papers.
MolmoBot: A Vision-Language Model for Zero-Shot Robot Manipulation 24.03.2026 38:06
Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on real-world tabletop tasks, outperforming π₀.₅ baseline at 39.2%.
LeWorldModel: Stable End-to-End JEPA from Pixels 24.03.2026 13:09
A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI systems.
EgoVerse: An Egocentric Data Ecosystem for Scaling Robot Learning 24.03.2026 41:57
Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via behavior cloning; includes cloud infrastructure, data viewer, and human-to-robot transfer algorithms to enable cross-embodiment learning without teleoperation.
HSImul3R: Physics-Driven Reconstruction of Human–Scene Interactions 24.03.2026 28:08
Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, deployable directly to humanoid robots for world modeling and manipulation.
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.