Shaoqing Tan
Embodied AI 101
Stay in the loop on research in AI and physical intelligence.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Robometer and the Future of Robotic Reward Modeling 30.05.2026 43:19
New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.
Qwen-VLA: A Generalist Vision–Language–Action Robot Model 29.05.2026 35:31
A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.
EXPO-FT: Sample-Efficient Reinforcement Learning Fine-Tuning for Vision-Language-Action Models 29.05.2026 33:59
Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.
RoboMeter: Learning Dense Rewards from Successes and Failures 29.05.2026 37:24
RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.
MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents 27.05.2026 53:15
Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.
ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking 27.05.2026 19:53
Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.
TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes 26.05.2026 43:18
Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.
MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics 26.05.2026 28:45
Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation 25.05.2026 29:47
Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.
Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models 24.05.2026 27:07
New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.
FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents 24.05.2026 27:13
A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.
AgentFloor: A Benchmark for Long-Horizon Agent Planning 24.05.2026 35:08
A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.
AlexNet: The Deep Convolutional Network That Transformed Vision 23.05.2026 41:44
AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.
A Few Useful Things to Know About Machine Learning 23.05.2026 45:51
Practical insights into ML pitfalls and best practices for machine learning practitioners.
SimToolReal: A Universal Dexterous Tool-Use Policy 23.05.2026 28:31
Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.
Mimic-Video: Learning Physics Priors from Web-Scale Video for Robot Dexterity 23.05.2026 29:15
Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.
Deep Residual Learning for Image Recognition (ResNet) 23.05.2026 24:04
Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.
Attention Is All You Need – The Transformer Revolution 23.05.2026 26:47
Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.
NVIDIA Cosmos: World Foundation Models for Physical AI 20.05.2026 29:23
World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.
LATENT: Teaching a Humanoid to Play Tennis from Imperfect Data 19.05.2026 20:24
Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level performance on a humanoid robot.
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models 19.05.2026 42:10
Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.
World Action Models: The Next Frontier in Embodied AI 19.05.2026 36:18
First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.
Training a Whole-Body Control Foundation Model 18.05.2026 39:37
Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.
DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation 18.05.2026 43:38
Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.
MMSkills: Building Multimodal Skill Libraries for Visual Agents 18.05.2026 19:33
Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.