← → to navigate
Lecture 10 · Embodied AI with Web-Scale Video Data · October 1

Reinforcement Learning from Human Videos

A practitioner's guide
Abstract

Reinforcement learning promises robots that learn skills through practice, but two things hold it back: rewards nobody knows how to write, and robot data nobody can afford to collect. This lecture examines a way around both: using human motion (videos, mocap, references) as both the reward and the data for RL. Starting from DeepMimic's tracking-reward recipe, we look at how to define the reward and the RL task when learning from human videos, and how trajectory optimization can improve the references before we track them. We then follow the recipe through locomotion and dexterous manipulation, including the sim-to-real problems that stand between simulation and physical robots.

Speaker

Julen Urain is a Member of the Technical Staff at Amazon FAR, working on dexterous manipulation. He was previously a Research Scientist in the Robotics group at Meta FAIR, and a postdoctoral researcher at IAS and DFKI. He received his PhD summa cum laude from TU Darmstadt, advised by Jan Peters, and interned at NVIDIA's Seattle Robotics Lab. His research spans generative models, 3D robot learning, learning from video, and optimization, and has received multiple best-paper awards; he was a George Giralt PhD Award finalist and an RSS Pioneer in 2023.

Guest Lecture

Reinforcement Learning from Human Videos

A practitioner's guide

Using human motion (videos, mocap, references) as the reward and the data for training embodied agents.

Amazon FAR
📍 The Johns Hopkins University · 🕒 Lecture 10 · October 1
Motivation

Robots need skills.
Humans already have them.

01 — RL without references is awkward
Reward = forward velocity. It works, but the motions are unnatural. Richer skills need rewards nobody knows how to write. (Heess et al., DeepMind, 2017)
02 — Robot demonstrations don't scale
13 robots, 17 months, ~130k teleoperated demos, for a single task set. Fleets, operators and time make this very expensive. (RT-1, Brohan et al., RSS 2023)

Meanwhile, humans upload millions of hours of skilled behavior for free. Central question: can human motion serve as the reward signal for RL?

Setting Up the Problem

How should we define the reward?

Policy π(at | st) Simulator physics + robot action at state st+1 reward = ?

Every block is standard except the reward. For the skills we care about, nobody knows how to write it by hand.

Setting Up the Problem

Could the video itself be the reward?

A human doing the task

One YouTube video of a pianist. Can we use this video as the reward?

Classical RL Policy π(at | st) Simulator physics + robot action at state st+1 reward = ?

The video shows exactly what success looks like. Could it fill the reward block?

Setting Up the Problem

Reward as pixel comparison?

Policy π(at | st) Simulator physics + robot action at state st+1 reward
reward  =  −‖
human video frame human video · frame t
−
simulator render frame simulator render · frame t
‖²  ?

Comparing raw pixels is tough: different appearance, viewpoint and embodiment; noisy, prone to reward hacking (the policy exploits the metric), and it says little about how to move.

Setting Up the Problem

Or compare in a latent space?

Policy π(at | st) Simulator physics + robot action at state st+1 reward
reward  =  −‖ φ(
human video frame human video · frame t
) − φ(
simulator render frame simulator render · frame t
) ‖²

φ is a frozen encoder pre-trained on human video (recall Lecture 6), e.g. R3M (Ego4D). Inverse RL has long pursued this: learning the reward, or the features it is built on, from demonstrations. Rewards from XIRL and VIP point in this direction.

Invariant to appearance and viewpoint, and an active line of work. But the structure stays implicit: hard to interpret, easy to hack, silent about bodies, hands and objects. Can we make it explicit? Can we extract more structured signals from the video first?IRL: Ng & Russell, 2000; Abbeel & Ng, 2004 · R3M: Nair et al., 2022 · XIRL: Zakka et al., 2021 · VIP: Ma et al., 2023

Part I

Can we design more 3D, geometrically structured rewards?

Pixels and latents treat the video as appearance. But the video contains bodies, hands and objects: geometry moving through time. Part I turns video into references a robot can track.

The Opportunity

We can 4D-fy videos

As in Lecture 8 (HMR 2.0, HaMeR, WHAM, HaWoR), off-the-shelf models lift ordinary video into 3D geometry moving through time: bodies, hands and objects.

01 — 4D humans
RGB video → body mesh trajectory. Pose, shape, and motion of every joint, per frame. (HMR 2.0 / 4D-Humans, Goel et al., 2023)
02 — 4D hands
RGB video → hand mesh trajectory. Full articulation of the fingers, even in the wild. (HaMeR, Pavlakos et al., 2024)

A 4D reconstruction is a trajectory of states, and a trajectory of states can serve as a reward function: reward the robot for tracking it.

The Opportunity · This Month

Foundation models are 4D-fying for us

With GPT-6 Astra, real2sim (rebuilding a real scene as a simulation) is collapsing from a research problem into a prompt: video or photos in, physics-ready 3D & 4D scenes out.

01 — Video → world
Input video → reconstructed world that survives a gravity rollout, rendered along the recovered camera trajectory.
02 — Photos → articulated scene
A few photos → an articulated kitchen: cabinets, drawers and appliances that open and close.
03 — Hand video → robot pipeline
One human-hand video → full real2sim pipeline: tracking, inverse-kinematics (IK) retargeting to a 44-DoF hand, contact physics, all written by the model itself.

The references RL needs are becoming one prompt away: much of what follows is getting cheaper.

The Core Recipe

Imitation as a tracking reward

Instead of engineering task rewards, reward the agent for matching a reference motion inside a physics simulator.

rt  =  ωimitate · rimitatet  +  ωtask · rtaskt ω: scalar weights per term
the demonstration Hand-object demonstration: human hands opening an articulated object
The full task, as the human did it.
rimitate — from the body Motion reward: robot hand keypoints matched between demonstration and simulation
The robot's motion imitates the human's: hand landmarks for hands, body landmarks for humanoids. This term shapes how to move.
rtask — from the world Task reward: object pose matched between demonstration and simulation
The world changes the way it did for the human: object motion, piano key state, drawer angle. This term defines what counts as success.

Drop either term and RL finds the wrong optimum. Figures: DexMachina, Mandi et al., 2025.

The Embodiment Gap

The video is human, the robot is not

Left: the person. Middle: a fitted SMPL body. Right: the humanoid reference after retargeting. Proportions, joint limits and degrees of freedom all differ.

Human video → SMPL joints → robot reference
VideoMimic — Allshire et al., 2025
The Embodiment Gap

Kinematic retargeting

01 — Extract 3D landmarks / keypoints SMPL: photo to 2D keypoints to skeleton to body mesh

Fit a parametric model (SMPL for bodies, MANO for hands) and read out its joint keypoints: trajectories xi,t the robot must reproduce. (SMPL — Loper et al.)

02 — Kinematic retargeting / inverse kinematics (IK)
q*1:T = argminq Σt,i ‖FKi(qt) − s·xi,t‖² + λ‖qt − qt−1‖²
s.t.  qmin ≤ qt ≤ qmax

Joint angles qt that put the robot's links on the keypoints, solved frame by frame (differential IK), with λ tying each frame to the previous one. FKi: forward kinematics of link i; s: human-to-robot scale. (mink — Zakka, 2026)

This is pure kinematics: making q1:T dynamically feasible is RL's job.

Kinematic retargeting · Detail — IK with mink

Differential IK in practice

mink · Unitree G1 example
GMR · one human clip (LAFAN1) → 5 robots, real time on CPU
GMR · Unitree H1, ChaCha dance
GMR · PAL Talos, fighting

GMR scales the human motion to the robot, then solves frame-by-frame IK with mink (keypoint position + orientation tasks, joint limits). The output is kinematic only: it can still slide, float or self-penetrate. (mink — Zakka, 2026 · GMR — Araujo, Ze et al., 2025)

The Embodiment Gap

Human shape ≠ robot shape: reshape before retargeting

Landmarks from a human-proportioned body land in places the robot cannot reach. Fitting the human model to the robot's shape first makes the IK targets attainable.

Bodies · raw human  |  shape-fitted human  |  humanoid
Optimize the SMPL shape parameters β toward the robot's proportions, read the landmarks off the reshaped human, then solve the IK. H2O — He et al., 2024
Hands · morphometric optimization → residual RL → policy Morphometric Imitation, Fig. 1: human hand-object trajectories retargeted to Allegro, Dex3 and Sharpa hands, refined by residual RL, distilled into a real-world visuomotor policy
Same idea for hands: scale MANO's palm and each finger to the robot hand's URDF, then retarget while preserving the demonstrated contacts. Residual RL makes the result dynamically feasible. Morphometric Imitation — Sadjadpour et al., 2026
Grounding the References

Kinematic retargeting breaks physics

Kinematic retargeting failure modes: penetration and missed contacts
The IK never asked the simulator: the knee and shin pass through the box (1, 3) and the hand hovers without contact (2). Replayed with physics, the motion collapses. (OmniRetarget — Yang et al., 2025)
Grounding the References

Skating & collision costs

Grounding can happen inside the IK solver: add costs that pin contacts and forbid penetration.

01 — Foot-skating cost
Ltskating  =  Σfoot ∈ {L,R} ctfoot · ‖ptfoot − pt−1foot‖
           + ctfoot · ‖ptankle − pt−1ankle‖

ptfoot is the world position of the foot link at time t, and ctfoot ∈ {0,1} flags a detected foot contact. While in contact, foot & ankle displacement is penalized: a planted foot stays planted.

02 — Collision cost di > m NO COST di < 0 PENALIZED
Ltcollision  =  Σi ∈ links max(0, m − dit)²

Robot as capsules, scene as a heightmap: di is the capsule-to-surface distance, m a safety margin. The hinge activates only when a link comes closer than m; penetration (di < 0) costs most. Same form over link pairs for self-collision.

Both costs as used in VideoMimic, Allshire et al., 2025.

Grounding the References

OmniRetarget: beyond per-link costs

arXiv 2025
Yang, Huang, Wu, Kanazawa, Abbeel, Sferrazza, Liu, Duan, Shi — omniretarget.github.io

Skating & collision costs are blind to relations: where the hand sits relative to the box, or the body relative to the terrain. OmniRetarget preserves the whole interaction, not only the contacts.

The interaction mesh Interaction meshes: robot joints (blue) and object/environment points (green), target and source

Delaunay tetrahedralization over body joints (blue) and points sampled on the object & environment (green); target (robot) left, source (human) right.

Laplacian deformation energy
L(pt,i)  =  pt,i − Σj ∈ 𝒩(i) wij · pt,j
EL  =  Σi ‖ L(psourcet,i) − L(ptargett,i) ‖²

The Laplacian coordinate is a keypoint's offset from the average of its mesh neighbors (uniform wij = 1/|𝒩(i)|), so it encodes local relative geometry. Human keypoints are first rescaled by the robot-to-human height ratio; EL absorbs the remaining proportion mismatch. Minimize EL (+ a smoothness term) over the robot configuration qt, with collision avoidance, joint & velocity limits and no foot skating as hard constraints: skating and collisions become constraints, not costs.

Grounding the References

SPIDER: the simulator as the retargeter

IROS 2026
Pan, Wang, Qi, Liu, Bharadhwaj, Sharma, Wu, Shi, Malik, Hogan — jc-bao.github.io/spider-project · on your Lecture 17 reading list

IK costs only approximate physics. Instead: roll the motion out in a physics simulator and search for the controls that reproduce the demo. Whatever comes out is feasible by construction.

Sampling-based trajectory optimization

Each ghost is a parallel rollout of a sampled action sequence; thousands run per iteration in a GPU simulator.

Interactive · CEM on a toy hand trajectory
Part II

The reference is ready. Now, reinforcement learning.

4D-fied, retargeted and grounded, the reference tells the robot what to do. RL in simulation learns how: a closed-loop policy that tracks it under real physics, and doesn't fall when reality pushes back.

Why RL

Closed-loop, or it falls

Unitree H1 robots performing live at the 2025 Chinese New Year Gala: twirling handkerchiefs and dancing in sync on a real stage. An open-loop replay of the reference would fall on the first slip: these routines run on RL tracking policies that feel the state and correct at every step. CCTV via The Sun, 2025

DeepMimic

The blueprint: motion clips as RL rewards for physics-based characters.

DeepMimic

SIGGRAPH 2018
Peng, Abbeel, Levine, van de Panne — UC Berkeley / UBC
DeepMimic filmstrip: humanoid cartwheel and Atlas spin-kick

The tracking reward alone is not enough: RSI and ET are what make dynamic skills learnable.

DeepMimic · Detail — the shape of rconfig

Why a bell-curve reward?

r = exp(−k · (q̂ − q)²) 1 0 q = q̂ → r = 1 deviation q̂ − q (per joint) r ≈ 0, flat gentle near the top
Three properties that matter

Bounded in [0, 1], never negative. With early termination, a negative reward would make falling early the optimal way to stop accumulating penalty. Positive reward means surviving longer always pays.

Forgiving near the reference. The top of the bell is nearly flat, so small errors barely cost anything; the policy is not forced into brittle exact tracking.

Flat tails cut both ways. Far from the reference the reward is ≈0 whatever the agent does, so exploration gets no learning signal there. This is exactly the failure RSI fixes by starting episodes on the reference.

The temperature k sets the tolerance per term: tight for joint configuration, looser for velocities. Press ↑ to go back.

DeepMimic · Do RSI and ET matter?

The harder the skill, the more they matter

Backflip — a dynamic skill Backflip learning curves with and without RSI and ET

Without ET, falling floods the data and return plateaus: the policy settles for a partial flip. Without RSI, it never experiences the landing until it can already reach it. Only RSI + ET learns the full backflip.

Walk — an easy skill Walk learning curves with and without RSI and ET

For walking, every variant converges, though without ET (green) it needs several times more samples. The tricks matter far less when exploration is easy: every state is a few steps from a good one.

Learning curves from Fig. 11, Peng et al., 2018. The tracking reward defines the task; RSI & ET decide whether RL can solve it.

DeepMimic · Adding a task reward

Imitation shapes the throw, the task aims it

rt = ωimitate rimitatet + ωtask rtaskt   with rtask = hit the target
both terms: a baseball pitch, aimed at the target With imitation and task rewards: character performs a pitch and the ball hits the target
without the imitation term Without the imitation reward: the character lurches toward the target instead of throwing
The task is "solved" degenerately: the character stops throwing and carries the ball into the target.

Imitation-only policies hit the target in 5% of throws and 19% of strikes; with both terms, 75% and 99%. Drop either term and RL finds the wrong optimum. Figs. 7–8, Table 4, Peng et al., 2018.

SFV

Reinforcement Learning of Physical Skills from Videos

SFV: RL of Physical Skills from Videos

SIGGRAPH Asia 2018
Peng, Kanazawa, Malik, Abbeel, Levine — UC Berkeley
SFV pipeline: video, pose estimation, motion reconstruction, reference motion, motion imitation with RL
Tracking at Scale

From one reference to many

DeepMimic and SFV train one policy per clip. Two routes lead to a single controller that holds an entire motion library.

01 — Multi-reference RL ref 1 ref 2 ⋮ ref N one policy π single RL run any ref curriculum: practice what fails

The reference joins the observation; a sampling curriculum decides what to practice. PHC · SONIC

02 — Multiple teachers + distillation ref 1 ref 2 ⋮ ref N expert π₁ expert π₂ ⋮ expert πN RL per clip distill student π DAgger / BC

Per-clip RL keeps motion quality; the student inherits all of it. BeyondMimic · PianoMime (later in this lecture)

Both routes end in one network; they differ in where the RL happens.

Multi-reference RL · A representative case

SONIC: one policy, any reference

A single tracker trained on ~700 h of mocap follows whatever reference it is given, whatever produced it.

SONIC: Human video
Human video. Kung fu, copied live from a person via video pose estimation
SONIC: VR whole-body teleop
VR whole-body teleop. Operator's full-body motion streamed as the reference
SONIC: Text prompt
Text prompt. “Monkey movement”: text → motion generator → tracked
SONIC: Planner · style
Planner · style. Kinematic planner with a stealth-walking style
SONIC: Planner · low posture
Planner · low posture. Elbow crawling on the floor
SONIC: Planner · dynamic
Planner · dynamic. Boxing: timing and balance under fast motions

Same weights in every clip, real Unitree G1. Videos from nvlabs.github.io/GEAR-SONIC, Luo, Yuan, Wang, et al., NVIDIA, 2026.

Can we extend these ideas to dexterous manipulation?

Locomotion tracks the body. Manipulation must also track the world: objects, contacts, and fingers that never stop touching things.

Dexterous manipulation · Action space

Residual vs. non-residual control

Body tracking lets the policy output the whole action. Many dexterous-hand works instead learn a correction on top of the retargeted reference.

01 — Non-residual state reference policy π full action action
at = π(st, q̂t)   q̂t: retargeted reference

The policy finds every joint target itself. Fine when balance and dynamics dominate and RL must move far from the kinematic reference anyway. DeepMimic · SFV · SONIC · TCDM/PGDM

02 — Residual reference state retargeting / IK → base action residual π small, bounded + action
at = abaset + α·π(st, q̂t),  α small

From step 0 the hand already follows the human; RL only learns the corrections that make the grasp or key press work. PianoMime · ManipTrans · DexMachina (wrist)

TCDM / PGDM

Track the object, not the hand: DeepMimic for dexterous manipulation, with one pre-grasp as the exploration trick.

TCDM / PGDM: Exemplar Object Trajectories and Pre-Grasps

ICRA 2023
Dasari, Gupta, Kumar — CMU / Meta AI
TCDM: the goal object trajectory (a bunny flying through the air), then the learned Adroit-hand policy reproducing it
Goal trajectory, then learned policy. The reference is only the object's path; the hand motion is found by RL.
rt  =  rtaskt  (object pose tracking)
  • DeepMimic, applied to the object. Exponential tracking reward, early termination, time step in the state; but the reference is an object pose trajectory (from mocap, animators or expert policies), not a body motion.
  • TCDM benchmark. 50 tasks, 34 objects, 3 hands (Adroit, D'Hand, D'Manus). Reward, termination and hyper-parameters are identical across tasks: only the exemplar changes.
  • Non-residual. The policy outputs the full hand action directly (slide 30, left).

Object tracking is embodiment-agnostic: the same exemplar defines the task for any hand, the idea DexMachina builds on later. Dasari et al., 2023 · pregrasps.github.io

PGDM · How does RL explore?

One pre-grasp beats a full hand reference

01 — Start from a pre-grasp PGDM pipeline: a CEM planner moves the hand from its start state to a pre-grasp pose around a hammer, then the RL policy takes over

A planner moves the hand to one pre-grasp pose (from mocap via IK, teleop, hand labels or a grasp predictor); RL starts from there. The manipulation version of DeepMimic's reference state initialization.

02 — Final success, TCDM-30 PGDM Fig. 3: final success. With a pre-grasp every method reaches about 0.8 to 0.85; without it GRAFF, curriculum and DeepMimic stay near 0.2

Without a pre-grasp, even DeepMimic-style fingertip tracking stalls near 0.2; with it, every method reaches ≈0.8–0.85. Extra hand supervision adds nothing.

Fig. 3, Dasari et al., 2023. As with RSI and ET in DeepMimic: in dexterous manipulation, exploration is the bottleneck.

PianoMime

All the way to the internet: a generalist piano player from YouTube videos.

PianoMime

CoRL 2024
Qian, Urain, Zakka, Peters — TU Darmstadt / UC Berkeley
PianoMime: YouTube piano videos to an agent acting in MuJoCo, conditioned on a target song

YouTube as the demonstration source. Mine piano-performance videos + MIDI: fingertip trajectories from the video become the rimitate, pressed keys the rtask.

rt  =  ωimitate · rimitatet  +  ωtask · rtaskt

rimitate — distance between the robot's fingertips and the human pianist's fingertips extracted from the video: it tells the policy which finger plays which key.

rtask — the piano's state: reward for pressing the right keys, and only the right keys, read off the target MIDI.

PianoMime · Does the human reference matter?

Notes alone are not enough

PianoMime Fig 3: hand postures baseline vs ours, and F1 score per song for baseline, IK and ours

Baseline: task reward only, correct notes with no human reference. IK: kinematic replay of the human fingertips, no RL. Ours: IK warm start + residual RL (slide 30) with imitation and task rewards. Fig. 3, Qian et al., 2024.

PianoMime · Results · click a video to play it with sound (one at a time)

From YouTube to MuJoCo

Videos: pianomime.github.io · Qian et al., 2024
01 — Human video → robot
Human pianist (bottom), policy (top). The robot tracks the fingertips extracted from the YouTube video while pressing the notes in the MIDI.
02 — Unseen songs
One generalist policy. Distilled from many single-song experts, it plays songs it never trained on.

Bimanual Dexterous Manipulation

Two hands, one object: from human hand-object mocap to robot hands. Three recipes: imitate the hand and correct it (ManipTrans), track the object with a curriculum (DexMachina), match the contact wrenches (CHORD).

See also: ObjDex (Chen et al., 2024) · BiDexHD (Zhou et al., 2024) · DexMan (Hsieh et al., 2025)

ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning

CVPR 2025
Li, Li, Liu, Li, Huang — BIGAI
ManipTrans pipeline: a frozen hand imitator follows the reference hand trajectories; a trainable residual policy, given object trajectories, shape and contact info, adds a residual action
Stage 1: a hand imitator, frozen. Stage 2: a residual policy that also sees the object.
at = aIt + ΔaRt    rt = rhandt + robjectt + rcontactt
  • Imitate the hand first. A generalist hand-only imitator is trained with PPO on large mocap data (wrist and finger tracking), with no object in the scene.
  • Then a residual per task. A zero-initialised residual adds the corrections the object needs; its reward adds object following and contact force terms (slide 30, right).
  • Physics relaxation curriculum. Start with no gravity and high friction, then restore them gradually. Plus RSI and early termination with a shrinking object threshold.
  • DexManipNet. 3.3K episodes on the Inspire hand from FAVOR + OakInk-V2 (pen capping, bottle unscrewing), deployed on real hands.

Hand-first: the human hand motion is the base action, and RL only learns what the object adds. Li et al., 2025 · maniptrans.github.io

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

ICML 2026
Mandi, Hou, Fox, Narang, Mandlekar, Song — Stanford / NVIDIA
DexMachina overview (Fig. 2): a hand-object demonstration defines task, motion and contact rewards; virtual object controllers go from strong to weak to none over training

Object-first: tracking alone is not enough, and the decaying virtual-controller curriculum makes long-horizon dexterity learnable. Mandi et al., 2025

CHORD: Contact Wrench Guidance from Human Demonstration

arXiv 2026
Zhu, Liu, Jain, et al. — NVIDIA
CHORD method: human reference contacts (top) and robot contacts (bottom) are compared through the wrench spaces they span, left and right hand, over time; imitation reward plus contact wrench reward
Human (top) vs. robot (bottom) contacts, compared by the wrench spaces they span (middle, per hand) rather than by where the fingers touch.
rt = rtaskt + rimitt + rwrencht
  • Match what the contacts can do. Contact location alone does not fix the object's motion. Reward the robot when its contacts can push the object the way the human's contacts could; this transfers across hand shapes.
  • Combines both recipes. Residual on the retargeted motion (as ManipTrans) and annealed virtual object controllers (as DexMachina).
  • Scale. 4,739 bimanual tasks from mocap and in-house videos; 82% success on 1,831 tasks with one set of hyper-parameters. Real transfer on two Sharpa hands.
AUC, 60 tasksManipTransDexMachinaNo contact termContact position onlyCHORD (wrench)
Tracking score0.5060.7370.7740.8620.918

Contact-first: the human tells the robot what forces to apply, not where to put each finger. Zhu et al., 2026 · Tab. 1 & ablation · project page

CHORD · Rollouts

One recipe, simulation to hardware

Simulation · RL policies tracking human demos
Box (2×)
Mixer (2×)
Capsule machine (2×)
Tool use: spatula
Real world · bimanual, two Sharpa hands
Box
Mixer
Capsule machine
Waffle iron

Box, mixer and capsule machine run in both sim and real; the long-horizon sim clips are sped up 2×. Zhu et al., 2026 · videos from the project page

Sim2Real

From policies trained in simulation to real robots: crossing the reality gap.

Sim2Real · Main strategies

Randomize the simulator, or identify it

01 — Domain randomization

Train on many simulators. Randomize friction, mass, motors, delays: reality should look like one more sample.

+ no real data    − conservative policy, hand-tuned ranges

02 — System identification

Fit the simulator to the robot. Record real trajectories, tune the sim parameters until they match, then train.

+ precise, strong policy    − real data per robot, blind to what the sim cannot model

real robot   randomized sims   identified sim  ·  In practice: Sys-ID to centre, DR around it.

Also: online adaptation (RMA) · teacher–student / asymmetric critic · learned sim–real residuals (ASAP) · noise & latency models.  Tobin et al., 2017 · Peng et al., 2018 · Tan et al., 2018 · Hwangbo et al., 2019

Sim2Real · Interactive · a robot joint tracking a step command

Sys-ID matches the trajectory, DR covers it

Nominal sim matches real
Real data covered
Spread policy must handle

Toy PD-controlled joint (inset: one motor + link, dark = real, orange = sim, teal = randomized sims), I q̈ = Kp(qcmd − q) − (Kd + b) q̇, simulated live. DR samples inertia, friction and actuator delay. "Covered" = real measurements within ±0.04 rad of some simulated rollout.

Sim2Real · Dexterous manipulation · Automatic domain randomization

How wide? Let the policy earn every widening

ADR · one parameter over training parameter range (e.g. friction) nominal sim threshold success at the edge training →
01

Start with one simulator. The nominal one, no randomization.

02

Test at the edges. Now and then, set a parameter to the end of its range and measure success.

03

Good enough? Widen it. Failing? Shrink it back. Each parameter grows on its own.

04

A curriculum for free. The ranges grow as fast as the policy can keep up, and far wider than anyone would hand-tune.

With a memory (LSTM), the policy learns to identify the world it is in while acting: sys-ID happens inside the network. OpenAI et al., Rubik's Cube, 2019 · Handa et al., DeXtreme, ICRA 2023 (ADR on GPU)

Sim2Real · Humanoids in practice

Identify the actuators, randomize the rest

Randomize · body & world
Mass & centre of massCAD is only roughly right
Ground friction & terrainevery floor is different
Pusheslearn to recover from the unexpected
Sensor noise & delaysIMU, encoders, control lag
Calibration offsetsno two robots are assembled alike
friction · terrain mass · CoM IMU push motor
Identify · the actuators
Motor response Hwangbo 2019learn it from real excitation data
Joint friction BAM · Open Duckswing a pendulum on a test bench
Rotor inertia Berkeley Humanoidfrom the motor's CAD and gearing
Torque limits & backlash Disney BDXeach actuator alone on a torque bench
Motor heat Disney Olafa thermal model the policy can see

Still a gap? Learn a correction from a few real rollouts (ASAP).  For small hobby servos, actuator fidelity is most of the sim2real gap (Microduck).

Sim2Real · Across robots

Trained in simulation, walking in the real world

Disney BDXDisney Research
SIM
REAL
actuators measured on a bench
Disney OlafDisney Research
SIM
REAL
motor heat modelled in sim
Unitree G1 · ASAPCMU · NVIDIA
SIM
REAL
learned correction from real rollouts
Berkeley HumanoidUC Berkeley
SIM
REAL
rotor inertia & friction identified
ANYmalETH Zurich
SIMREAL
actuator network learned from data
Open Duck Miniopen source
SIM
REAL
servo friction fitted on a pendulum
MicroduckPollen Robotics
SIM
REAL
servo friction fitted on a pendulum

Every real clip is a policy that never trained on the real robot.

What changes from robot to robot is which part of the gap gets modelled, and it is almost always the actuators.

Grandia et al., RSS 2024 · Müller et al., 2025 · He et al., ASAP, RSS 2025 · Liao et al., 2024 · Hwangbo et al., Sci. Robotics 2019 · Open Duck Mini (A. Pirrone) · Microduck (Pollen Robotics; sim clip is a different policy)
Sim2Real · Detail — comparison across robots

Cheaper actuators need better actuator models

RobotSizeActuatorsSys-IDRandomizedSim · policy
MicroduckPollen Robotics, 2026
~25 cm
0.8 kg
Dynamixel XL330 hobby servos BAM “M6” model: voltage control, back-EMF, Stribeck + load-dep. friction Voltage & drop, delay, friction, armature, CoM, mass, encoder bias, IMU tilt, pushes, backlash mjlab (MuJoCo Warp) · PPO, 50 Hz
Open Duck Mini v2A. Pirrone, open source
~42 cm
<$400
Feetech STS3215 servos, 7.4 V BAM pendulum bench → damping, Kp, friction, armature into MJCF Friction 0.5–1.0, mass ×0.9–1.1, CoM ±5 cm, Kp, 0–60 ms delay, backlash MuJoCo Playground (MJX) · 50 Hz, BDX-style imitation reward
Disney BDXGrandia et al., RSS 2024
0.66 m
15.4 kg · 14 DoF
Unitree A1 / Go1 QDD + Dynamixel (head) Analytic actuator model from a torque bench: friction, torque–speed limit, backlash, armature, noise Actuator params, backlash 0.002–0.015 rad, armature +20%, encoder offset, mass, friction, terrain, pushes 90–150 N Isaac Gym · PPO, 50 Hz, 37.5 Hz low-pass
Disney OlafMüller et al., 2025
0.89 m
14.9 kg · 25 DoF
Unitree + Dynamixel Adds a thermal model (±1.9 °C); temperature fed to the policy, kept <80 °C Actuator & body DR (ranges not verified) Isaac Sim · 50 Hz
Berkeley HumanoidLiao et al., 2024
0.85 m
16 kg · 12 DoF
Custom 9:1 quasi-direct drive Light: armature from CAD rotor inertia, friction from simple tests Friction 0.2–1.25, mass ×0.9–1.1, joint friction, armature, encoder offset. No PD-gain or delay DR Isaac Lab · 50 Hz, 25 kHz PD
Unitree G1ASAP, He et al., RSS 2025
1.32 m
~35 kg · 23 DoF
PMSM joint motors Learned Δ-action on the ankles (100 real clips); a mass/CoM/KpKd search as baseline Standard legged-robot DR (ranges not verified) Isaac Gym · tested sim-to-sim in Isaac Sim, Genesis
Agility DigitRadosavovic et al., Sci. Robotics 2024
~1.6 m
45 kg · 16 act.
Electric joints + passive closed chains None reported; closed chains as stiff virtual springs Dynamics, control params, terrain, noise, delay Isaac Gym · transformer, 50 Hz, 1 kHz PD

heavy actuator modelling   light / learned   none  ·  Hobby servos (friction, backlash, voltage sag) need a bench model; transparent QDDs get by on CAD + DR. Bars: height to scale.

github.com/pollen-robotics/microduck · github.com/apirrone/Open_Duck_Mini · arXiv 2501.05204 · 2512.16705 · 2407.21781 · 2502.01143 · 2303.03381 · BAM: arXiv 2410.08650

Sim2Real · Dexterous manipulation · Results

Randomize in sim, then go zero-shot

SimToolRealKedia, Lum et al. · RSS 2026
One policy for any tool: trained on random handle-and-head shapes, it follows an object path taken from a human video.
Human video → goal poses in sim → real rollout
OmniResetYin, Westenbroek et al. · arXiv 2026
Diverse simulator resets instead of demos or curricula: one reward, PPO at 64K+ envs, then distilled to an RGB policy.
Leg twisting · sim vs. real
SIM
REAL
Peg reoriented against the hole (non-prehensile) · sim vs. real
SIM
REAL

Both train in simulation only and transfer zero-shot, with no real-world fine-tuning. SimToolReal randomizes the objects (procedural tools); OmniReset randomizes the starting states and, for the RGB student, the visuals. arXiv 2602.16863 · arXiv 2603.15789 · simtoolreal.github.io · OmniReset

Sim2Real · Observations · Teacher–student

Privileged teacher, deployable student

Teacher · sees the simulator's state object pose & velocity contact forces friction, mass, shape exact, noise-free, no camera π teacher actions RL in sim · fast imitate Student · sees what the robot will see joint history (proprioception) rendered camera: depth / RGB touch, if the hand has it noisy, delayed, partial π student actions supervised · deployed
Why split in two?

RL is hard; RL from pixels is harder and slower. The teacher solves the task on clean state; the student only has to copy, which is supervised learning.

Observations

What the teacher reads directly, the student must infer: friction and mass from its joint history, object pose from the camera.

Rendering

Only the student needs images, so only it pays for rendering, and it inherits a visual gap: randomize textures, lights and camera pose; add depth noise and holes; or feed point clouds and masks that look alike in sim and real.

The same split runs through this lecture: RMA and HORA infer the hidden physics; Dactyl and DeXtreme learn the pose from rendered images; RotateIt adds simulated touch. Lee et al., 2020 · Kumar et al., RMA, 2021 · Qi et al., HORA, 2022 · Chen et al., Visual Dexterity, 2023 · Qi et al., RotateIt, CoRL 2023

Sim2Real · Teacher–student · Distillation

Two ways to distill: whose states do we learn on?

01 — Teacher rollouts + behavior cloning teacher acts dataset student fits collect once, train offline teacher's states student drifts off them

+ simple, any model (e.g. diffusion), data reused   − small errors compound: the student reaches states the teacher never showed it

PianoMime, 2024 (song experts → one generalist) · Lin et al., CoRL 2025 (specialists → diffusion policy)

02 — DAgger: the student drives, the teacher labels student acts teacher labels retrain repeat, in the simulator mistakes get corrected

+ learns on its own mistakes, stays on track   − teacher and simulator must run during training

Ross et al., AISTATS 2011 · Lee et al., 2020 · HORA, 2022 · Visual Dexterity, 2023

DAgger only works because the teacher can be asked anywhere: in simulation the privileged state exists at every state the student visits. With a human teacher, that is expensive; with a sim teacher, it is free.

Guest Lecture

Thanks for listening

Reinforcement Learning from Human Videos: A practitioner's guide
Amazon FAR
📍 The Johns Hopkins University · 🕒 Lecture 10 · October 1