Shaoqing Tan

Embodied AI 101

Stay in the loop on research in AI and physical intelligence.

Autor

Shaoqing Tan

Kategoria

Technology

Ostatni odcinek

10 lip 2026

Gdzie słuchać?

Podcasty w aplikacji Replaio Radio Już wkrótce

Podcasty trafią do aplikacji już wkrótce. Zainstaluj teraz i jako pierwszy zobacz nowe podejście do podcastów

Pobierz z Google Play Zainstaluj za darmo Android 5 mln+ pobrań · ocena 4,8 iOS niedługo

Odcinki

DexWM: A Dexterous Manipulation World Model from Human Videos 29.06.2026

A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning 29.06.2026

Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.

Continual Robot Policy Learning via Variational Neural Dynamics 29.06.2026

Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.

PhysisForcing: Physics-Reinforced World Models for Robotic Manipulation 29.06.2026

Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.

Translation as a Bridging Action 29.06.2026

Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.

Play2Perfect: Dexterous Play Pretraining for Precise Assembly 28.06.2026

Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.

Dexora: Open-Source VLA for High-DoF Bimanual Dexterity 28.06.2026

First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.

WorldVLA: Towards Autoregressive Action World Model 28.06.2026

Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.

HumDex: Humanoid Dexterous Manipulation Made Easy 28.06.2026

HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.

ForceBand: Learning Forceful Manipulation with sEMG 27.06.2026

Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.

In-Context World Modeling for Robotic Control 27.06.2026

Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.

WOLF-VLA: Vision-Language-Action for Humanoid Walking 27.06.2026

Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.

Motion-Focused Latent Action for Cross-Embodiment VLA from Human Videos 27.06.2026

Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.

ManiFlow: Manipulation via Rectified Flow 27.06.2026

ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.

RL-100: Toward Highly Reliable Real-World Robot Reinforcement Learning 27.06.2026

RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation 26.06.2026

A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.

Bi-HIL: Bilateral Control-Based Multimodal Hierarchical Imitation Learning for Long-Horizon Contact-Rich Manipulation 26.06.2026

Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation 26.06.2026

Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.

ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning 26.06.2026

ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.

ConstrainedMimic: Safe Humanoid Robot Motion Tracking 26.06.2026

A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.

REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering 25.06.2026

Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.

HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching 25.06.2026

Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.

Reactive Diffusion Policy: Slow-Fast Visual-Tactile Learning for Contact-Rich Manipulation 25.06.2026

Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.

SARM2 + SPIRAL: Multi-Task Reward Models and RL Refinement for Long-Horizon Dexterous Manipulation 25.06.2026

Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.

Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm VLA Systems 25.06.2026

Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.

Słuchaj podcastu Embodied AI 101 w Replaio

Radio i podcasty w jednej aplikacji - za darmo, bez zakładania konta. Zainstaluj już dziś i nie przegap premiery

Pobierz z Google Play

Replaio nie jest wydawcą podcastów; nazwy audycji, okładki i audio należą do ich autorów i są rozpowszechniane przez publiczne kanały RSS