Lecture 10 · Embodied AI with Web-Scale Video Data · October 1
Reinforcement Learning from Human Videos
A practitioner's guide
Abstract
Reinforcement learning promises robots that learn skills through practice, but two things hold it back: rewards nobody knows how to write, and robot data nobody can afford to collect. This lecture examines a way around both: using human motion (videos, mocap, references) as both the reward and the data for RL. Starting from DeepMimic's tracking-reward recipe, we look at how to define the reward and the RL task when learning from human videos, and how trajectory optimization can improve the references before we track them. We then follow the recipe through locomotion and dexterous manipulation, including the sim-to-real problems that stand between simulation and physical robots.
Speaker
Julen Urain is a Member of the Technical Staff at Amazon FAR, working on dexterous manipulation. He was previously a Research Scientist in the Robotics group at Meta FAIR, and a postdoctoral researcher at IAS and DFKI. He received his PhD summa cum laude from TU Darmstadt, advised by Jan Peters, and interned at NVIDIA's Seattle Robotics Lab. His research spans generative models, 3D robot learning, learning from video, and optimization, and has received multiple best-paper awards; he was a George Giralt PhD Award finalist and an RSS Pioneer in 2023.
Guest Lecture
Reinforcement Learning from Human Videos
A practitioner's guide
Using human motion (videos, mocap, references) as the reward and the data for training embodied agents.
📍 The Johns Hopkins University · 🕒 Lecture 10 · October 1
Motivation
Robots need skills. Humans already have them.
01 — RL without references is awkwardReward = forward velocity. It works, but the motions are unnatural. Richer skills need rewards nobody knows how to write. (Heess et al., DeepMind, 2017)02 — Robot demonstrations don't scale13 robots, 17 months, ~130k teleoperated demos, for a single task set. Fleets, operators and time make this very expensive. (RT-1, Brohan et al., RSS 2023)
Meanwhile, humans upload millions of hours of skilled behavior for free. Central question: can human motion serve as the reward signal for RL?
Setting Up the Problem
How should we define the reward?
Every block is standard except the reward. For the skills we care about, nobody knows how to write it by hand.
Setting Up the Problem
Could the video itself be the reward?
A human doing the task
One YouTube video of a pianist. Can we use this video as the reward?
Classical RL
The video shows exactly what success looks like. Could it fill the reward block?
Setting Up the Problem
Reward as pixel comparison?
reward = −‖
human video · frame t
−
simulator render · frame t
‖² ?
Comparing raw pixels is tough: different appearance, viewpoint and embodiment; noisy, prone to reward hacking (the policy exploits the metric), and it says little about how to move.
Setting Up the Problem
Or compare in a latent space?
reward = −‖ φ(
human video · frame t
) − φ(
simulator render · frame t
) ‖²
φ is a frozen encoder pre-trained on human video (recall Lecture 6), e.g. R3M (Ego4D). Inverse RL has long pursued this: learning the reward, or the features it is built on, from demonstrations. Rewards from XIRL and VIP point in this direction.
Invariant to appearance and viewpoint, and an active line of work. But the structure stays implicit: hard to interpret, easy to hack, silent about bodies, hands and objects. Can we make it explicit? Can we extract more structured signals from the video first?IRL: Ng & Russell, 2000; Abbeel & Ng, 2004 · R3M: Nair et al., 2022 · XIRL: Zakka et al., 2021 · VIP: Ma et al., 2023
Part I
Can we design more 3D, geometrically structured rewards?
Pixels and latents treat the video as appearance. But the video contains bodies, hands and objects: geometry moving through time. Part I turns video into references a robot can track.
The Opportunity
We can 4D-fy videos
As in Lecture 8 (HMR 2.0, HaMeR, WHAM, HaWoR), off-the-shelf models lift ordinary video into 3D geometry moving through time: bodies, hands and objects.
01 — 4D humansRGB video → body mesh trajectory. Pose, shape, and motion of every joint, per frame. (HMR 2.0 / 4D-Humans, Goel et al., 2023)02 — 4D handsRGB video → hand mesh trajectory. Full articulation of the fingers, even in the wild. (HaMeR, Pavlakos et al., 2024)
A 4D reconstruction is a trajectory of states, and a trajectory of states can serve as a reward function: reward the robot for tracking it.
The Opportunity · This Month
Foundation models are 4D-fying for us
With GPT-6 Astra, real2sim (rebuilding a real scene as a simulation) is collapsing from a research problem into a prompt: video or photos in, physics-ready 3D & 4D scenes out.
01 — Video → worldInput video → reconstructed world that survives a gravity rollout, rendered along the recovered camera trajectory.02 — Photos → articulated sceneA few photos → an articulated kitchen: cabinets, drawers and appliances that open and close.03 — Hand video → robot pipelineOne human-hand video → full real2sim pipeline: tracking, inverse-kinematics (IK) retargeting to a 44-DoF hand, contact physics, all written by the model itself.
The references RL needs are becoming one prompt away: much of what follows is getting cheaper.
The Core Recipe
Imitation as a tracking reward
Instead of engineering task rewards, reward the agent for matching a reference motion inside a physics simulator.
rt = ωimitate · rimitatet + ωtask · rtasktω: scalar weights per term
the demonstrationThe full task, as the human did it.rimitate — from the bodyThe robot's motion imitates the human's: hand landmarks for hands, body landmarks for humanoids. This term shapes how to move.rtask — from the worldThe world changes the way it did for the human: object motion, piano key state, drawer angle. This term defines what counts as success.
Drop either term and RL finds the wrong optimum. Figures: DexMachina, Mandi et al., 2025.
The Embodiment Gap
The video is human, the robot is not
Left: the person. Middle: a fitted SMPL body. Right: the humanoid reference after retargeting. Proportions, joint limits and degrees of freedom all differ.
Fit a parametric model (SMPL for bodies, MANO for hands) and read out its joint keypoints: trajectories xi,t the robot must reproduce. (SMPL — Loper et al.)
Joint angles qt that put the robot's links on the keypoints, solved frame by frame (differential IK), with λ tying each frame to the previous one. FKi: forward kinematics of link i; s: human-to-robot scale. (mink — Zakka, 2026)
This is pure kinematics: making q1:Tdynamically feasible is RL's job.
Kinematic retargeting · Detail — IK with mink
Differential IK in practice
mink · Unitree G1 example
GMR · one human clip (LAFAN1) → 5 robots, real time on CPU
GMR · Unitree H1, ChaCha dance
GMR · PAL Talos, fighting
GMR scales the human motion to the robot, then solves frame-by-frame IK with mink (keypoint position + orientation tasks, joint limits). The output is kinematic only: it can still slide, float or self-penetrate. (mink — Zakka, 2026 · GMR — Araujo, Ze et al., 2025)
The Embodiment Gap
Human shape ≠ robot shape: reshape before retargeting
Landmarks from a human-proportioned body land in places the robot cannot reach. Fitting the human model to the robot's shape first makes the IK targets attainable.
Bodies · raw human | shape-fitted human | humanoidOptimize the SMPL shape parameters β toward the robot's proportions, read the landmarks off the reshaped human, then solve the IK. H2O — He et al., 2024
Hands · morphometric optimization → residual RL → policySame idea for hands: scale MANO's palm and each finger to the robot hand's URDF, then retarget while preserving the demonstrated contacts. Residual RL makes the result dynamically feasible. Morphometric Imitation — Sadjadpour et al., 2026
Grounding the References
Kinematic retargeting breaks physics
The IK never asked the simulator: the knee and shin pass through the box (1, 3) and the hand hovers without contact (2). Replayed with physics, the motion collapses. (OmniRetarget — Yang et al., 2025)
Grounding the References
Skating & collision costs
Grounding can happen inside the IK solver: add costs that pin contacts and forbid penetration.
ptfoot is the world position of the foot link at time t, and ctfoot ∈ {0,1} flags a detected foot contact. While in contact, foot & ankle displacement is penalized: a planted foot stays planted.
02 — Collision cost
Ltcollision = Σi ∈ links max(0, m − dit)²
Robot as capsules, scene as a heightmap: di is the capsule-to-surface distance, m a safety margin. The hinge activates only when a link comes closer than m; penetration (di < 0) costs most. Same form over link pairs for self-collision.
Skating & collision costs are blind to relations: where the hand sits relative to the box, or the body relative to the terrain. OmniRetarget preserves the whole interaction, not only the contacts.
The interaction mesh
Delaunay tetrahedralization over body joints (blue) and points sampled on the object & environment (green); target (robot) left, source (human) right.
Laplacian deformation energy
L(pt,i) = pt,i − Σj ∈ 𝒩(i) wij · pt,j
EL = Σi ‖ L(psourcet,i) − L(ptargett,i) ‖²
The Laplacian coordinate is a keypoint's offset from the average of its mesh neighbors (uniform wij = 1/|𝒩(i)|), so it encodes local relative geometry. Human keypoints are first rescaled by the robot-to-human height ratio; EL absorbs the remaining proportion mismatch. Minimize EL (+ a smoothness term) over the robot configuration qt, with collision avoidance, joint & velocity limits and no foot skating as hard constraints: skating and collisions become constraints, not costs.
Grounding the References
SPIDER: the simulator as the retargeter
IROS 2026
Pan, Wang, Qi, Liu, Bharadhwaj, Sharma, Wu, Shi, Malik, Hogan — jc-bao.github.io/spider-project· on your Lecture 17 reading list
IK costs only approximate physics. Instead: roll the motion out in a physics simulator and search for the controls that reproduce the demo. Whatever comes out is feasible by construction.
Sampling-based trajectory optimization
Each ghost is a parallel rollout of a sampled action sequence; thousands run per iteration in a GPU simulator.
Interactive · CEM on a toy hand trajectory
Part II
The reference is ready. Now, reinforcement learning.
4D-fied, retargeted and grounded, the reference tells the robot what to do. RL in simulation learns how: a closed-loop policy that tracks it under real physics, and doesn't fall when reality pushes back.
Why RL
Closed-loop, or it falls
Unitree H1 robots performing live at the 2025 Chinese New Year Gala: twirling handkerchiefs and dancing in sync on a real stage. An open-loop replay of the reference would fall on the first slip: these routines run on RL tracking policies that feel the state and correct at every step. CCTV via The Sun, 2025
DeepMimic
The blueprint: motion clips as RL rewards for physics-based characters.
DeepMimic
SIGGRAPH 2018
Peng, Abbeel, Levine, van de Panne — UC Berkeley / UBC
The tracking reward.Reward = how closely the simulated character's joint configuration & velocities match the reference clip at each timestep, plus an optional task term, trained with standard RL (PPO).
Reference State Initialization.Start episodes at random points along the clip, so the policy learns the landing of a backflip before it can reach it. This removes the exploration bottleneck.
Early Termination.End the episode on failure (e.g., the torso hits the ground), so hopeless states stop polluting the data and lying on the floor never becomes an optimum.
The tracking reward alone is not enough: RSI and ET are what make dynamic skills learnable.
DeepMimic · Detail — the shape of rconfig
Why a bell-curve reward?
r = exp(−k · (q̂ − q)²)
Three properties that matter
Bounded in [0, 1], never negative. With early termination, a negative reward would make falling early the optimal way to stop accumulating penalty. Positive reward means surviving longer always pays.
Forgiving near the reference. The top of the bell is nearly flat, so small errors barely cost anything; the policy is not forced into brittle exact tracking.
Flat tails cut both ways. Far from the reference the reward is ≈0 whatever the agent does, so exploration gets no learning signal there. This is exactly the failure RSI fixes by starting episodes on the reference.
The temperature k sets the tolerance per term: tight for joint configuration, looser for velocities. Press ↑ to go back.
DeepMimic · Do RSI and ET matter?
The harder the skill, the more they matter
Backflip — a dynamic skill
Without ET, falling floods the data and return plateaus: the policy settles for a partial flip. Without RSI, it never experiences the landing until it can already reach it. Only RSI + ET learns the full backflip.
Walk — an easy skill
For walking, every variant converges, though without ET(green) it needs several times more samples. The tricks matter far less when exploration is easy: every state is a few steps from a good one.
Learning curves from Fig. 11, Peng et al., 2018. The tracking reward defines the task; RSI & ET decide whether RL can solve it.
DeepMimic · Adding a task reward
Imitation shapes the throw, the task aims it
rt = ωimitate rimitatet + ωtask rtasktwith rtask = hit the target
both terms: a baseball pitch, aimed at the targetwithout the imitation termThe task is "solved" degenerately: the character stops throwing and carries the ball into the target.
Imitation-only policies hit the target in 5% of throws and 19% of strikes; with both terms, 75% and 99%. Drop either term and RL finds the wrong optimum. Figs. 7–8, Table 4, Peng et al., 2018.
SFV
Reinforcement Learning of Physical Skills from Videos
References from YouTube, not mocap.Monocular clips of acrobatics (backflips, cartwheels, dances) become DeepMimic-style tracking targets.
Pose estimation, then motion reconstruction.Per-frame 2D (OpenPose) and 3D (HMR, cf. Lecture 8) poses are jittery; an optimization turns them into one temporally consistent reference.
Motion imitation with RL.Plus adaptive state initialization (RSI that learns where to start) and motion completion: predicting the rest of a motion from a single frame.
Tracking at Scale
From one reference to many
DeepMimic and SFV train one policy per clip. Two routes lead to a single controller that holds an entire motion library.
01 — Multi-reference RL
The reference joins the observation; a sampling curriculum decides what to practice. PHC · SONIC
02 — Multiple teachers + distillation
Per-clip RL keeps motion quality; the student inherits all of it. BeyondMimic · PianoMime (later in this lecture)
Both routes end in one network; they differ in where the RL happens.
Multi-reference RL · A representative case
SONIC: one policy, any reference
A single tracker trained on ~700 h of mocap follows whatever reference it is given, whatever produced it.
Human video. Kung fu, copied live from a person via video pose estimationVR whole-body teleop. Operator's full-body motion streamed as the referenceText prompt. “Monkey movement”: text → motion generator → trackedPlanner · style. Kinematic planner with a stealth-walking stylePlanner · low posture. Elbow crawling on the floorPlanner · dynamic. Boxing: timing and balance under fast motions
Same weights in every clip, real Unitree G1. Videos from nvlabs.github.io/GEAR-SONIC, Luo, Yuan, Wang, et al., NVIDIA, 2026.
Can we extend these ideas to dexterous manipulation?
Locomotion tracks the body. Manipulation must also track the world: objects, contacts, and fingers that never stop touching things.
Dexterous manipulation · Action space
Residual vs. non-residual control
Body tracking lets the policy output the whole action. Many dexterous-hand works instead learn a correction on top of the retargeted reference.
01 — Non-residual
at = π(st, q̂t) q̂t: retargeted reference
The policy finds every joint target itself. Fine when balance and dynamics dominate and RL must move far from the kinematic reference anyway. DeepMimic · SFV · SONIC · TCDM/PGDM
02 — Residual
at = abaset + α·π(st, q̂t), α small
From step 0 the hand already follows the human; RL only learns the corrections that make the grasp or key press work. PianoMime · ManipTrans · DexMachina (wrist)
TCDM / PGDM
Track the object, not the hand: DeepMimic for dexterous manipulation, with one pre-grasp as the exploration trick.
TCDM / PGDM: Exemplar Object Trajectories and Pre-Grasps
ICRA 2023
Dasari, Gupta, Kumar — CMU / Meta AI
Goal trajectory, then learned policy. The reference is only the object's path; the hand motion is found by RL.
rt = rtaskt(object pose tracking)
DeepMimic, applied to the object.Exponential tracking reward, early termination, time step in the state; but the reference is an object pose trajectory (from mocap, animators or expert policies), not a body motion.
TCDM benchmark.50 tasks, 34 objects, 3 hands (Adroit, D'Hand, D'Manus). Reward, termination and hyper-parameters are identical across tasks: only the exemplar changes.
Non-residual.The policy outputs the full hand action directly (slide 30, left).
Object tracking is embodiment-agnostic: the same exemplar defines the task for any hand, the idea DexMachina builds on later. Dasari et al., 2023 · pregrasps.github.io
PGDM · How does RL explore?
One pre-grasp beats a full hand reference
01 — Start from a pre-grasp
A planner moves the hand to one pre-grasp pose (from mocap via IK, teleop, hand labels or a grasp predictor); RL starts from there. The manipulation version of DeepMimic's reference state initialization.
02 — Final success, TCDM-30
Without a pre-grasp, even DeepMimic-style fingertip tracking stalls near 0.2; with it, every method reaches ≈0.8–0.85. Extra hand supervision adds nothing.
Fig. 3, Dasari et al., 2023. As with RSI and ET in DeepMimic: in dexterous manipulation, exploration is the bottleneck.
PianoMime
All the way to the internet: a generalist piano player from YouTube videos.
PianoMime
CoRL 2024
Qian, Urain, Zakka, Peters — TU Darmstadt / UC Berkeley
YouTube as the demonstration source. Mine piano-performance videos + MIDI: fingertip trajectories from the video become the rimitate, pressed keys the rtask.
rt = ωimitate · rimitatet + ωtask · rtaskt
rimitate — distance between the robot's fingertips and the human pianist's fingertips extracted from the video: it tells the policy which finger plays which key.
rtask — the piano's state: reward for pressing the right keys, and only the right keys, read off the target MIDI.
PianoMime · Does the human reference matter?
Notes alone are not enough
Baseline: task reward only, correct notes with no human reference. IK: kinematic replay of the human fingertips, no RL. Ours: IK warm start + residual RL (slide 30) with imitation and task rewards. Fig. 3, Qian et al., 2024.
PianoMime · Results · click a video to play it with sound (one at a time)
01 — Human video → robotHuman pianist (bottom), policy (top). The robot tracks the fingertips extracted from the YouTube video while pressing the notes in the MIDI.02 — Unseen songsOne generalist policy. Distilled from many single-song experts, it plays songs it never trained on.
Bimanual Dexterous Manipulation
Two hands, one object: from human hand-object mocap to robot hands. Three recipes: imitate the hand and correct it (ManipTrans), track the object with a curriculum (DexMachina), match the contact wrenches (CHORD).
See also: ObjDex (Chen et al., 2024) · BiDexHD (Zhou et al., 2024) · DexMan (Hsieh et al., 2025)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
CVPR 2025
Li, Li, Liu, Li, Huang — BIGAI
Stage 1: a hand imitator, frozen. Stage 2: a residual policy that also sees the object.
at = aIt + ΔaRt rt = rhandt + robjectt + rcontactt
Imitate the hand first.A generalist hand-only imitator is trained with PPO on large mocap data (wrist and finger tracking), with no object in the scene.
Then a residual per task.A zero-initialised residual adds the corrections the object needs; its reward adds object following and contact force terms (slide 30, right).
Physics relaxation curriculum.Start with no gravity and high friction, then restore them gradually. Plus RSI and early termination with a shrinking object threshold.
DexManipNet.3.3K episodes on the Inspire hand from FAVOR + OakInk-V2 (pen capping, bottle unscrewing), deployed on real hands.
Hand-first: the human hand motion is the base action, and RL only learns what the object adds. Li et al., 2025 · maniptrans.github.io
DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
ICML 2026
Mandi, Hou, Fox, Narang, Mandlekar, Song — Stanford / NVIDIA
Functional retargeting.From one human hand-object demo (from ARCTIC, a Lecture 8 reading), learn a policy whose main reward is reproducing the object's trajectory (as in TCDM), with auxiliary hand-motion and contact terms from the retargeted demo: long-horizon, bimanual tasks with articulated objects.
Virtual object controllers with decaying strength.At first the object is driven toward its targets "by magic", and the policy only needs to follow along; the assist decays on a curriculum until real contacts carry the task.
Cross-embodiment by design.The object trajectory is embodiment-agnostic, so the same demo trains Allegro, Inspire, XHand and Schunk hands, and the benchmark doubles as a functional comparison of hand hardware.
Object-first: tracking alone is not enough, and the decaying virtual-controller curriculum makes long-horizon dexterity learnable. Mandi et al., 2025
CHORD: Contact Wrench Guidance from Human Demonstration
arXiv 2026
Zhu, Liu, Jain, et al. — NVIDIA
Human (top) vs. robot (bottom) contacts, compared by the wrench spaces they span (middle, per hand) rather than by where the fingers touch.
rt = rtaskt + rimitt + rwrencht
Match what the contacts can do.Contact location alone does not fix the object's motion. Reward the robot when its contacts can push the object the way the human's contacts could; this transfers across hand shapes.
Combines both recipes.Residual on the retargeted motion (as ManipTrans) and annealed virtual object controllers (as DexMachina).
Scale.4,739 bimanual tasks from mocap and in-house videos; 82% success on 1,831 tasks with one set of hyper-parameters. Real transfer on two Sharpa hands.
AUC, 60 tasks
ManipTrans
DexMachina
No contact term
Contact position only
CHORD (wrench)
Tracking score
0.506
0.737
0.774
0.862
0.918
Contact-first: the human tells the robot what forces to apply, not where to put each finger. Zhu et al., 2026 · Tab. 1 & ablation · project page
Box, mixer and capsule machine run in both sim and real; the long-horizon sim clips are sped up 2×. Zhu et al., 2026 · videos from the project page
Sim2Real
From policies trained in simulation to real robots: crossing the reality gap.
Sim2Real · Main strategies
Randomize the simulator, or identify it
01 — Domain randomization
Train on many simulators. Randomize friction, mass, motors, delays: reality should look like one more sample.
+ no real data − conservative policy, hand-tuned ranges
02 — System identification
Fit the simulator to the robot. Record real trajectories, tune the sim parameters until they match, then train.
+ precise, strong policy − real data per robot, blind to what the sim cannot model
real robot randomized sims identified sim · In practice: Sys-ID to centre, DR around it.
Also: online adaptation (RMA) · teacher–student / asymmetric critic · learned sim–real residuals (ASAP) · noise & latency models. Tobin et al., 2017 · Peng et al., 2018 · Tan et al., 2018 · Hwangbo et al., 2019
Sim2Real · Interactive · a robot joint tracking a step command
Sys-ID matches the trajectory, DR covers it
Nominal sim matches real
Real data covered
Spread policy must handle
Toy PD-controlled joint (inset: one motor + link, dark = real, orange = sim, teal = randomized sims), I q̈ = Kp(qcmd − q) − (Kd + b) q̇, simulated live. DR samples inertia, friction and actuator delay. "Covered" = real measurements within ±0.04 rad of some simulated rollout.
Start with one simulator. The nominal one, no randomization.
02
Test at the edges. Now and then, set a parameter to the end of its range and measure success.
03
Good enough? Widen it. Failing? Shrink it back. Each parameter grows on its own.
04
A curriculum for free. The ranges grow as fast as the policy can keep up, and far wider than anyone would hand-tune.
With a memory (LSTM), the policy learns to identify the world it is in while acting: sys-ID happens inside the network. OpenAI et al., Rubik's Cube, 2019 · Handa et al., DeXtreme, ICRA 2023 (ADR on GPU)
Sim2Real · Humanoids in practice
Identify the actuators, randomize the rest
Randomize · body & world
Mass & centre of massCAD is only roughly right
Ground friction & terrainevery floor is different
Pusheslearn to recover from the unexpected
Sensor noise & delaysIMU, encoders, control lag
Calibration offsetsno two robots are assembled alike
Identify · the actuators
Motor response Hwangbo 2019learn it from real excitation data
Joint friction BAM · Open Duckswing a pendulum on a test bench
Rotor inertia Berkeley Humanoidfrom the motor's CAD and gearing
Torque limits & backlash Disney BDXeach actuator alone on a torque bench
Motor heat Disney Olafa thermal model the policy can see
Still a gap? Learn a correction from a few real rollouts (ASAP). For small hobby servos, actuator fidelity is most of the sim2real gap (Microduck).
Sim2Real · Across robots
Trained in simulation, walking in the real world
Disney BDXDisney Research
SIM
REAL
actuators measured on a bench
Disney OlafDisney Research
SIM
REAL
motor heat modelled in sim
Unitree G1 · ASAPCMU · NVIDIA
SIM
REAL
learned correction from real rollouts
Berkeley HumanoidUC Berkeley
SIM
REAL
rotor inertia & friction identified
ANYmalETH Zurich
SIMREAL
actuator network learned from data
Open Duck Miniopen source
SIM
REAL
servo friction fitted on a pendulum
MicroduckPollen Robotics
SIM
REAL
servo friction fitted on a pendulum
Every real clip is a policy that never trained on the real robot.
What changes from robot to robot is which part of the gap gets modelled, and it is almost always the actuators.
Grandia et al., RSS 2024 · Müller et al., 2025 · He et al., ASAP, RSS 2025 · Liao et al., 2024 · Hwangbo et al., Sci. Robotics 2019 · Open Duck Mini (A. Pirrone) · Microduck (Pollen Robotics; sim clip is a different policy)
Sim2Real · Detail — comparison across robots
Cheaper actuators need better actuator models
Robot
Size
Actuators
Sys-ID
Randomized
Sim · policy
MicroduckPollen Robotics, 2026
~25 cm 0.8 kg
Dynamixel XL330 hobby servos
BAM “M6” model: voltage control, back-EMF, Stribeck + load-dep. friction
Adds a thermal model (±1.9 °C); temperature fed to the policy, kept <80 °C
Actuator & body DR (ranges not verified)
Isaac Sim · 50 Hz
Berkeley HumanoidLiao et al., 2024
0.85 m 16 kg · 12 DoF
Custom 9:1 quasi-direct drive
Light: armature from CAD rotor inertia, friction from simple tests
Friction 0.2–1.25, mass ×0.9–1.1, joint friction, armature, encoder offset. No PD-gain or delay DR
Isaac Lab · 50 Hz, 25 kHz PD
Unitree G1ASAP, He et al., RSS 2025
1.32 m ~35 kg · 23 DoF
PMSM joint motors
Learned Δ-action on the ankles (100 real clips); a mass/CoM/KpKd search as baseline
Standard legged-robot DR (ranges not verified)
Isaac Gym · tested sim-to-sim in Isaac Sim, Genesis
Agility DigitRadosavovic et al., Sci. Robotics 2024
~1.6 m 45 kg · 16 act.
Electric joints + passive closed chains
None reported; closed chains as stiff virtual springs
Dynamics, control params, terrain, noise, delay
Isaac Gym · transformer, 50 Hz, 1 kHz PD
heavy actuator modelling light / learned none · Hobby servos (friction, backlash, voltage sag) need a bench model; transparent QDDs get by on CAD + DR. Bars: height to scale.
One policy for any tool: trained on random handle-and-head shapes, it follows an object path taken from a human video.
Human video → goal poses in sim → real rollout
OmniResetYin, Westenbroek et al. · arXiv 2026
Diverse simulator resets instead of demos or curricula: one reward, PPO at 64K+ envs, then distilled to an RGB policy.
Leg twisting · sim vs. real
SIM
REAL
Peg reoriented against the hole (non-prehensile) · sim vs. real
SIM
REAL
Both train in simulation only and transfer zero-shot, with no real-world fine-tuning. SimToolReal randomizes the objects (procedural tools); OmniReset randomizes the starting states and, for the RGB student, the visuals. arXiv 2602.16863 · arXiv 2603.15789 · simtoolreal.github.io · OmniReset
Sim2Real · Observations · Teacher–student
Privileged teacher, deployable student
Why split in two?
RL is hard; RL from pixels is harder and slower. The teacher solves the task on clean state; the student only has to copy, which is supervised learning.
Observations
What the teacher reads directly, the student must infer: friction and mass from its joint history, object pose from the camera.
Rendering
Only the student needs images, so only it pays for rendering, and it inherits a visual gap: randomize textures, lights and camera pose; add depth noise and holes; or feed point clouds and masks that look alike in sim and real.
The same split runs through this lecture: RMA and HORA infer the hidden physics; Dactyl and DeXtreme learn the pose from rendered images; RotateIt adds simulated touch. Lee et al., 2020 · Kumar et al., RMA, 2021 · Qi et al., HORA, 2022 · Chen et al., Visual Dexterity, 2023 · Qi et al., RotateIt, CoRL 2023
Sim2Real · Teacher–student · Distillation
Two ways to distill: whose states do we learn on?
01 — Teacher rollouts + behavior cloning
+ simple, any model (e.g. diffusion), data reused − small errors compound: the student reaches states the teacher never showed it
PianoMime, 2024 (song experts → one generalist) · Lin et al., CoRL 2025 (specialists → diffusion policy)
02 — DAgger: the student drives, the teacher labels
+ learns on its own mistakes, stays on track − teacher and simulator must run during training
Ross et al., AISTATS 2011 · Lee et al., 2020 · HORA, 2022 · Visual Dexterity, 2023
DAgger only works because the teacher can be asked anywhere: in simulation the privileged state exists at every state the student visits. With a human teacher, that is expensive; with a sim teacher, it is free.
Guest Lecture
Thanks for listening
Reinforcement Learning from Human Videos: A practitioner's guide