Shaoqing Tan

Embodied AI 101

Stay in the loop on research in AI and physical intelligence.

Author

Shaoqing Tan

Category

Technology

Latest episode

Jul 10, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model 13.06.2026

Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.

RISE: Self-Improving Robot Policy with Compositional World Model 13.06.2026

Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.

EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control 12.06.2026

Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning 12.06.2026

A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.

Robotic World Model: Learning to Simulate for Robust Robot Control 12.06.2026

Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.

AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 10.06.2026

Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.

ArtiFixer: Few-Step Diffusion for 3D Scene Reconstruction 10.06.2026

Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.

Deployment-Time Memorization in Foundation-Model Agents 10.06.2026

Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.

Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks 09.06.2026

Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.

SoCRATES: Evaluating LLM Mediators in Conflict Scenarios 09.06.2026

First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.

Unembedding Matrix as a Feature Lens: Unlocking Better Text Embeddings 09.06.2026

Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.

LeanMarathon: Autonomous Formalization of Math Proofs on Erdős Problems 08.06.2026

Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.

Deep Research Agents: Survey and Roadmap for Autonomous AI Research 08.06.2026

Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.

Cosmos 3: Omnimodal World Models for Physical AI 07.06.2026

Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.

Humanoid-GPT: GPT-Style Transformer for Zero-Shot Dynamic Humanoid Control 07.06.2026

GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotics.

Bending Paper, Shaping Dexterity: The Robotic Origami Challenge 05.06.2026

New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.

GraspGen-X: A Foundation Model for Zero-Shot 6-DoF Grasping 05.06.2026

First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.

When Does Deep RL Beat Calibrated Baselines? 04.06.2026

Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.

Training Deep Networks as Random Effects: An Optimization–Inference Duality 04.06.2026

Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.

Generative Depth Supervision for Embodied Vision-Language Models 02.06.2026

Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation 01.06.2026

Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 31.05.2026

NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit and predicts all coordinates in a single forward pass. This preserves intra-box geometric coherence while achieving 2.5x faster decoding throughp...

LT2: Linear-Time Looped Transformers 31.05.2026

Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.

One Learning Rate Doesn't Fit All: Layerwise Spectral Scheduling for Transformers 31.05.2026

Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.

SimToolReal: Procedural Tool Generation and a Universal Objective for Zero-Shot Tool Manipulation 30.05.2026

Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.

Listen to the Embodied AI 101 podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.