Shaoqing Tan
Embodied AI 101
Stay in the loop on research in AI and physical intelligence.
Gdzie słuchać?
Podcasty w aplikacji Replaio Radio Już wkrótcePodcasty trafią do aplikacji już wkrótce. Zainstaluj teraz i jako pierwszy zobacz nowe podejście do podcastów
Odcinki
DexWM: A Dexterous Manipulation World Model from Human Videos 29.06.2026 31:34
A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.
PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning 29.06.2026 13:00
Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.
Continual Robot Policy Learning via Variational Neural Dynamics 29.06.2026 32:32
Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.
PhysisForcing: Physics-Reinforced World Models for Robotic Manipulation 29.06.2026 23:32
Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.
Translation as a Bridging Action 29.06.2026 37:24
Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.
Play2Perfect: Dexterous Play Pretraining for Precise Assembly 28.06.2026 15:54
Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.
Dexora: Open-Source VLA for High-DoF Bimanual Dexterity 28.06.2026 38:22
First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.
WorldVLA: Towards Autoregressive Action World Model 28.06.2026 28:36
Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.
HumDex: Humanoid Dexterous Manipulation Made Easy 28.06.2026 33:09
HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.
ForceBand: Learning Forceful Manipulation with sEMG 27.06.2026 27:15
Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.
In-Context World Modeling for Robotic Control 27.06.2026 25:28
Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.
WOLF-VLA: Vision-Language-Action for Humanoid Walking 27.06.2026 29:26
Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.
Motion-Focused Latent Action for Cross-Embodiment VLA from Human Videos 27.06.2026 30:33
Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.
ManiFlow: Manipulation via Rectified Flow 27.06.2026 29:55
ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.
RL-100: Toward Highly Reliable Real-World Robot Reinforcement Learning 27.06.2026 32:28
RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation 26.06.2026 29:56
A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.
Bi-HIL: Bilateral Control-Based Multimodal Hierarchical Imitation Learning for Long-Horizon Contact-Rich Manipulation 26.06.2026 37:57
Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation 26.06.2026 13:19
Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.
ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning 26.06.2026 30:07
ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.
ConstrainedMimic: Safe Humanoid Robot Motion Tracking 26.06.2026 44:26
A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.
REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering 25.06.2026 22:58
Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching 25.06.2026 34:35
Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Learning for Contact-Rich Manipulation 25.06.2026 46:28
Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.
SARM2 + SPIRAL: Multi-Task Reward Models and RL Refinement for Long-Horizon Dexterous Manipulation 25.06.2026 40:43
Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm VLA Systems 25.06.2026 14:55
Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.
Podobne podcasty
Replaio nie jest wydawcą podcastów; nazwy audycji, okładki i audio należą do ich autorów i są rozpowszechniane przez publiczne kanały RSS