Chris Paxton and Michael Cho

RoboPapers

Chris Paxton & Michael Cho geek out over robotic papers with paper authors. robopapers.substack.com

Autor

Chris Paxton and Michael Cho

Kategorie

Technology

Podcast-Website

robopapers.substack.com

Neueste Folge

8. Jul 2026

Wo hören?

Podcasts in der App Replaio Radio Bald verfügbar

Podcasts kommen bald in die App. Installiere sie jetzt und erlebe als Erster einen ganz neuen Blick auf Podcasts

Bei Google Play herunterladen Kostenlos installieren Android 5 Mio.+ Downloads · Bewertung 4,8 iOS bald

Folgen

Ep#39: MolmoAct: An Action Reasoning Model that reasons in 3D space 28.10.2025

Reasoning models have massively expanded what LLMs are capable of, but this hasn’t necessarily applied to robotics. Perhaps this is in part because robots need to reason over space, not just words and symbols; so the robotics version of a reasoning model would need to think in 3D. That’s the idea behind MolmoAct, an “Action Reasoning Model” which generates spatial plans in order to predict precise...

Ep#38: Q Learning is Not Yet Scalable 24.10.2025

Offline reinforcement learning is crucial for robotics, but does it scale? We talk to Seohong, who discusses how for long-horizon manipulation problems the answer may be no — at least not yet. But there are tricks that you can use to make it work effectively. Watch episode #38 of RoboPapers with Michael Cho and Chris Paxton now! Abstract: In this work, we study the scalability of offline reinforce...

Ep#37: AMPLIFY: Actionless Motion Priors for Robot Learning from Videos 21.10.2025

Robots has a data problem, in that robotics data is rare. While human video is quite common, it’s not usually directly usable for robots for a variety of reasons, most significantly that it’s missing explicit, accurate robot actions. Instead, Jeremy proposes that we predict keypoint trajectories — basically, how any given point in an object will move as a robot performs a task. This lets us use ac...

Ep#36: Whole-Body Conditioned Egocentric Video Prediction 17.10.2025

Learning a true world model for a human body means taking high-dimensional actions representing the full body pose — the location of hands and feet, for example — and using it to predict the effects of each action. This would allow for an unprecedented level of simulation over the effects of each action on the world, but this level of information is usually not available. But, with a new dataset f...

Ep#35: Reinforcement Learning with Action Chunking 08.10.2025

Today, most robot learning from demonstration predicts action chunks , small robot action trajectory. Doing this is crucial for better performance, and has all kinds of advantages. But how can we apply these advantages to reinforcement learning ? We talked to Colin Li and Paul Zhou to find out more. Abstract: We present Q-chunking, a simple yet effective recipe for improving reinforcement learning...

Ep#34: RoboArena 03.10.2025

Evaluating robot policies is hard. Every lab has a different robot; reproducible evaluations are really challenging. This makes it hard for us to know which methods for learning robot policies are likely to perform the best in real-world scenarios. Taking a page from LLM evaluations like Chatbot Arena , RoboArena aims to address this problem through crowdsourcing evaluations with a network of diff...

Ep#33: A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search 26.09.2025

Learning policies via imitation is extremely potent, but making sure those policies will generalize to out of distribution settings is still very hard. SAILOR proposes a solution in learning to search via a learned world model, which outperforms existing imitation approaches. Gokul, Vibhakar, and Arnav tell us about their approach. Watch Episode #33 of RoboPapers, co-hosted by Michael Cho and Chri...

Ep #32: GMT: General Motion Tracking for Humanoid Whole-Body Control 21.09.2025

We’ve all seen videos of humanoid robots performing single tasks that are very impressive, like dancing or karate. But training humanoid robots to perform a wide range of complex motions is difficult. GMT is a general-purpose policy which can learn a wide range of robot motions. Watch Episode #32 of RoboPapers, with Zixuan Chen, co-hosted by Michael Cho and Chris Paxton, to learn more. Abstract: T...

Ep#31: Vision in Action: Learning Active Perception from Human Demonstrations 17.09.2025

Most robots are fixed in one location, with cameras at the correct location to solve whatever their task is going to be. This makes setting up the camera in the correct location a key part of task setup; it also makes the task unnecessarily difficult. Ideally, robots would move their camera around intelligently in order to gather all the information they need to perform a task. In “Vision in Actio...

Ep#30: R2S2: Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space 12.09.2025

How can we train humanoid mobile manipulators to perform complex manipulation tasks in the real world? Humanoids must be able to coordinate their arms and legs to perform complex tasks, but learning such skills via reinforcement learning is challenging. In R2S2, Li Yi tells us how we can learn a robot action space based on reinforcement learning which can transfer from simulation to the real world...

Ep#29: Exploring Low Cost Robots using LLM Agents 12.09.2025

One of the most exciting things about modern robot learning is that it’s so accessible. Ilia Larchenko talked to us about his journey into robotics: using low-cost robots like those from HuggingFace, building agents powered by Claude, and using these to put together robots that he can use to perform real, open-vocabulary tasks in the real world. Unlike most of our podcasts, this one isn’t really a...

Ep#28: DreamGen: Unlocking Generalization in Robot Learning through Video World Models 09.09.2025

Robotics has a data problem. Generative video models propose a solution: we can fine-tune these models with robot data to generate vast amounts of diverse task data, which in turn unlocks new behaviors and environmental generation. Abstract: We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through ne...

Ep#27: DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands 09.09.2025

How can we learn sim-to-real manipulation policies for grasping any object? Watch this episode with Ritvik Singh of NVIDIA to find out. Abstract: One of the most important yet challenging skills for robots is dexterous multi-fingered grasping of a diverse range of objects. Much of the prior work is limited by the speed, dexterity, or reliance on depth maps. In this paper, we introduce DextrAH-RGB,...

Ep#26: ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI 06.09.2025

Robotics as a field faces a large data gap, and one of the most promising directions for solving it is sim-to-real robot learning. We talked with Stone Tao about ManiSkill3: a powerful, easy-to-use framework for generalizable sim-to-real learning. In particular, we learned about how with Maniskill + HuggingFace LeRobot, you can easily train your own sim-to-real manipulation policies on a low-cost...

Ep#25: DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation 05.09.2025

Collecting dexterous humanoid robot data is difficult to scale. That's why Mengda Xu and Han Zhang built DexUMI: a tool for demonstrating how to control a dexterous robot hand, which allows you to quickly collect task data. Abstract: We present DexUMI - a data collection and policy learning framework that uses the human hand as the natural interface to transfer dexterous manipulation skills to var...

Ep#24: CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks 04.09.2025

How can we control humanoid robots to get them to perform a variety of challenging, real-world tasks while moving around in their environments? Yixuan Li and Siyuan Huang of BIGAI tell us how their complex system works. Abstract: Humanoid teleoperation plays a vital role in demonstrating and collecting data for complex humanoid-scene interactions. However, current teleoperation systems face critic...

Ep#23: FALCON 🦅- Learning Force-Adaptive Humanoid Loco-Manipulation 03.09.2025

One of the key advantages of humanoid robots is how they can handle heavier objects and exert forces in a more dynamic way, compared to wheeled robots. However, few research papers show this. We talked to Yuanhang Zhang about his really impressive work, FALCON, which shows humanoid robots pushing or pulling carts, lifting weights, and more. Abstract: Humanoid loco-manipulation holds transformative...

Ep#22: DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies 27.08.2025

How can we collect large-scale manipulation data in the real world? DexWild proposes a solution: an easy-to-use wearable that lets operators perform robotic tasks in a wide variety of environments easily. With the added data diversity, this results in robot policies which can operate in a variety of different environments. Watch the video to learn more! Abstract: Large-scale, diverse robot dataset...

Ep#21: TesserAct: Learning 4D Embodied World Models 24.08.2025

World models are an exciting area of research, wherein we predict video given robot actions and use these predictions for planning or training . But two dimensional image frames aren’t necessarily enough for robot planning; robotic tasks are inherently three dimensional, and so shouldn’t we be predicting how objects move around in 3d space? TesserAct does this, tuning a video generative model to p...

Ep#20: VideoMimic: Visual imitation enables contextual humanoid control 22.08.2025

Part of the advantage of humanoid robots is in their ability to interact with the environment: to climb stairs, to sit down, to navigate rough terrain, and so on. In VideoMimic, the authors produced a pipeline which let them: * Scan a scene into simulation using an iPhone * Train a whole-body control policy for manipulation and locomotion * Use this policy to control a humanoid robot as operated i...

Ep#19: Learning to Drive From a World Model 21.08.2025

Comma has been selling end-to-end driving kits for your car for years. Their Openpilot is available now, with source code on Github . Very interestingly, they train their lanekeeping behavior using generative world models. Watch to learn more. Abstract: Most self-driving systems rely on hand-coded perception outputs and engineered driving rules. Learning directly from human driving data with an en...

Ep#18 HuB: Learning Extreme Humanoid Balance 21.08.2025

This time, Tong and Boyuan tell us about how to achieve some of the most impressive humanoid robot balancing that we have ever seen. Take a look! Here’s the abstract: The human body demonstrates exceptional motor capabilities—such as standing steadily on one foot or performing a high kick with the leg raised over 1.5 meters—both requiring precise balance control. While recent research on humanoid...

Ep#17: EgoZero: Robot Learning from Smart Glasses 17.08.2025

Using human data is the key to unlocking the potential of general-purpose robotics: it’s much more plentiful and can potentially be collected quickly at scale for relatively little money. But how can we actually do this? Despite recent progress in general purpose robotics, robot policies still lag far behind basic human capabilities in the real world. Humans constantly interact with the physical w...

Ep#16: TWIST: Teleoperated Whole-Body Imitation System 17.08.2025

Learn how to do versatile, capable whole body teleoperation of humanoid robots — a key capability for unlocking the data we need to train general-purpose autonomous humanoids. Teleoperating humanoid robots in a whole-body manner marks a fundamental step toward developing general-purpose robotic intelligence, with human motion providing an ideal interface for controlling all degrees of freedom. Yet...

Ep#15: Navigation World Models 16.08.2025

World models are one of the hottest research areas. Amir’s work on navigation world models shows how we can use the world models to achieve new goals tin unfamiliar environments. Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future visual observations based on past observa...

Höre den Podcast RoboPapers in Replaio

Radio und Podcasts in einer App - kostenlos und ohne Anmeldung. Installiere sie noch heute und verpasse den Start nicht

Bei Google Play herunterladen

Replaio ist kein Herausgeber von Podcasts; die Namen der Sendungen, Cover und Audioinhalte gehören ihren Autoren und werden über öffentliche RSS-Feeds verbreitet