Shaoqing Tan

Embodied AI 101

Stay in the loop on research in AI and physical intelligence.

Author

Shaoqing Tan

Category

Technology

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation 25.06.2026

Proposes interleaved vision and language reasoning for robotic manipulation within a VLA framework. Aims to improve instruction following and task performance through integrated multimodal reasoning.

Playful Agentic Robot Learning 25.06.2026

Self-directed play combined with Code-as-Policy for reusable skill acquisition and downstream manipulation tasks.

Learning Unified Force and Position Control for Legged Loco-Manipulation 24.06.2026

A unified RL policy for quadrupeds and humanoids that jointly handles force and position control without force sensors, enabling compliant behaviors, force-aware imitation learning, and contact-rich tasks.

Robots that Collaborate: Sequential Asymmetric Imitation for Learning Coupled Robot Policies 24.06.2026

Explores imitation learning approaches for multi-robot systems, focusing on policy coupling through sequential asymmetric imitation to enable collaborative robot behaviors.

AstraBrain-WBC 0.5: A Humanoid Robot Cerebellum Foundation Model 24.06.2026

A humanoid robot 'cerebellum' foundation model trained on 20,000 hours of human motion data that demonstrates scaling laws for robot motion control and enables zero-shot execution of unseen motions on real humanoids.

SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping 24.06.2026

Combines the Spring-Loaded Inverted Pendulum (SLIP) model with reinforcement learning to achieve agile jumping behaviors in robotic systems.

DataClaw0: Agentic Tailoring for Raw Multimodal Streams 24.06.2026

A 9B model that filters noise from videos, GUI, and embodied data streams, reorganizing them into dense supervision via factual anchors and semantic synthesis; trained with SFT + GRPO across five domains with benchmarks.

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining 24.06.2026

Converts 6K+ hours of mixed human/robot egocentric video into robot pseudo-actions via camera-space alignment and reliability-aware loss, achieving 72.8% on RoboCasa and 91.1% on RoboTwin.

VERA: Video-to-Action World Model Policy 24.06.2026

A 14B-parameter video world model that converts predicted visual futures into embodiment-agnostic actions via Jacobian inverse-dynamics, enabling zero-shot cross-robot transfer across a Panda arm and 16-DoF hand with open-sourced weights and training code.

GEN-1: Scaled Dexterous Manipulation Foundation Model 24.06.2026

A dexterous manipulation foundation model trained on 500k hours of real-world bimanual data that handles deformable objects such as cardboard folding and screw packing, featuring online retry and adaptation capabilities.

Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation 22.06.2026

Develops an SE(3)-equivariant flow-based visuomotor policy leveraging spherical harmonics for efficient and geometrically consistent robot manipulation.

Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation 22.06.2026

Introduces a dual-stream transformer architecture inspired by cortical visual processing for learning robotic manipulation policies.

VisualClaw: A Self-Evolving Wearable Vision Agent 22.06.2026

An edge-filtered video streaming agent that evolves skills from memory and runs on smart glasses, reducing API costs by 98%, accompanied by the VisualClawArena benchmark dataset.

Kairos: A Native World Model Stack for Physical AI 21.06.2026

A 4B unified architecture for world understanding, generation, and action with hybrid linear attention enabling real-time edge inference across embodiments, outperforming 14B models on embodied benchmarks.

DragMesh-2: A Contact-Driven Framework for Dexterous Hand–Object Interaction 21.06.2026

A framework that trains a 51-DoF dexterous hand to open drawers and doors using only physical contact without requiring tactile sensors.

Guava: A Universal Harness for Robot Manipulation 21.06.2026

A 4B open-source VLA-style model trained on fewer than 2K simulation trajectories that matches closed frontier systems on real-world manipulation tasks with zero-shot generalization to novel objects and long-horizon behaviors, including failure-recovery demonstrations.

Geometric Action Model for Robot Policies 21.06.2026

A new geometric action model for robot manipulation policies that focuses on structured action representations to improve policy learning and generalization.

ENPIRE: Physical AutoResearch with a Fleet of 8 Robots 21.06.2026

ENPIRE demonstrates fully autonomous physical AutoResearch where Codex agents control a fleet of 8 robots overnight, self-improving through real hardware rollouts on tasks like zip-tie tying and GPU installation, while discovering physical scaling laws with built-in safety harnesses and frozen reward classifiers derived from demonstrations.

MolmoAct2: An Open Foundation Model for Real-World Robotics 21.06.2026

An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to enable community experiments on real robots for manipulation and generalist policies.

Hy-Embodied-0.5-VLA: A Massive Bimanual Teleoperation Dataset for Vision-Language-Action 15.06.2026

Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for multi-view egocentric teleop. The dataset and model are fully compatible with LeRobot v3.0.

Q-Guided Flow: Test-Time Gradient Guidance of Flow Policies 14.06.2026

New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.

Flow Reversal Steering: Guiding Diffusion-Based Robot Policies with High-Level Reasoning 14.06.2026

Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the diffusion noise space.

Test-Time Compute Scaling for Robot Policies (DIRECT) 14.06.2026

Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency trade-offs.

LabVLA: Bringing Vision-Language-Action to the Chemistry Lab 14.06.2026

RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and transfers to real Franka arms.

Humanoid-GPT: A Foundation Model for Zero-Shot Humanoid Control 13.06.2026

GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks like soccer, dancing, and digging. Requires no fine-tuning or task-specific adaptation.

Listen to the Embodied AI 101 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.