Shaoqing Tan

Embodied AI 101

Stay in the loop on research in AI and physical intelligence.

Author

Shaoqing Tan

Category

Technology

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Robometer and the Future of Robotic Reward Modeling 30.05.2026

New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.

Qwen-VLA: A Generalist Vision–Language–Action Robot Model 29.05.2026

A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.

EXPO-FT: Sample-Efficient Reinforcement Learning Fine-Tuning for Vision-Language-Action Models 29.05.2026

Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.

RoboMeter: Learning Dense Rewards from Successes and Failures 29.05.2026

RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.

MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents 27.05.2026

Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.

ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking 27.05.2026

Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.

TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes 26.05.2026

Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.

MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics 26.05.2026

Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation 25.05.2026

Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.

Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models 24.05.2026

New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.

FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents 24.05.2026

A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.

AgentFloor: A Benchmark for Long-Horizon Agent Planning 24.05.2026

A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.

AlexNet: The Deep Convolutional Network That Transformed Vision 23.05.2026

AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.

A Few Useful Things to Know About Machine Learning 23.05.2026

Practical insights into ML pitfalls and best practices for machine learning practitioners.

SimToolReal: A Universal Dexterous Tool-Use Policy 23.05.2026

Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.

Mimic-Video: Learning Physics Priors from Web-Scale Video for Robot Dexterity 23.05.2026

Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.

Deep Residual Learning for Image Recognition (ResNet) 23.05.2026

Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.

Attention Is All You Need – The Transformer Revolution 23.05.2026

Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.

NVIDIA Cosmos: World Foundation Models for Physical AI 20.05.2026

World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.

LATENT: Teaching a Humanoid to Play Tennis from Imperfect Data 19.05.2026

Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level performance on a humanoid robot.

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models 19.05.2026

Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.

World Action Models: The Next Frontier in Embodied AI 19.05.2026

First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.

Training a Whole-Body Control Foundation Model 18.05.2026

Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.

DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation 18.05.2026

Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.

MMSkills: Building Multimodal Skill Libraries for Visual Agents 18.05.2026

Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.

Listen to the Embodied AI 101 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.