Shaoqing Tan
Embodied AI 101
Stay in the loop on research in AI and physical intelligence.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model 13.06.2026 37:18
Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.
RISE: Self-Improving Robot Policy with Compositional World Model 13.06.2026 39:30
Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.
EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control 12.06.2026 35:27
Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.
Robix: A Unified Model for Robot Interaction, Reasoning and Planning 12.06.2026 35:02
A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.
Robotic World Model: Learning to Simulate for Robust Robot Control 12.06.2026 19:55
Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.
AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 10.06.2026 23:13
Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.
ArtiFixer: Few-Step Diffusion for 3D Scene Reconstruction 10.06.2026 26:05
Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.
Deployment-Time Memorization in Foundation-Model Agents 10.06.2026 39:31
Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.
Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks 09.06.2026 34:47
Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.
SoCRATES: Evaluating LLM Mediators in Conflict Scenarios 09.06.2026 20:58
First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.
Unembedding Matrix as a Feature Lens: Unlocking Better Text Embeddings 09.06.2026 24:39
Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.
LeanMarathon: Autonomous Formalization of Math Proofs on Erdős Problems 08.06.2026 27:22
Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.
Deep Research Agents: Survey and Roadmap for Autonomous AI Research 08.06.2026 37:23
Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.
Cosmos 3: Omnimodal World Models for Physical AI 07.06.2026 28:02
Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.
Humanoid-GPT: GPT-Style Transformer for Zero-Shot Dynamic Humanoid Control 07.06.2026 22:58
GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotics.
Bending Paper, Shaping Dexterity: The Robotic Origami Challenge 05.06.2026 29:02
New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.
GraspGen-X: A Foundation Model for Zero-Shot 6-DoF Grasping 05.06.2026 35:13
First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.
When Does Deep RL Beat Calibrated Baselines? 04.06.2026 15:42
Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.
Training Deep Networks as Random Effects: An Optimization–Inference Duality 04.06.2026 21:34
Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.
Generative Depth Supervision for Embodied Vision-Language Models 02.06.2026 28:36
Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation 01.06.2026 30:50
Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 31.05.2026 32:30
NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit and predicts all coordinates in a single forward pass. This preserves intra-box geometric coherence while achieving 2.5x faster decoding throughp...
LT2: Linear-Time Looped Transformers 31.05.2026 42:37
Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.
One Learning Rate Doesn't Fit All: Layerwise Spectral Scheduling for Transformers 31.05.2026 26:35
Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.
SimToolReal: Procedural Tool Generation and a Universal Objective for Zero-Shot Tool Manipulation 30.05.2026 25:40
Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.