Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co‑Training

Samuel Liu1,2,*, Youngsun Kim2, Martin Matak2, Gilwoo Lee2

1University of Cambridge  ·  2Industrial Next
*Corresponding author: syl63@cam.ac.uk

Overview: simulated data varied along world and behavior grounding axes, co-trained with real teleoperation demonstrations, and a schematic latent space of rollout trajectories.

We vary simulated data along two axes, world grounding and behavior grounding, giving four simulated datasets. Each is combined with 100 real teleoperation demonstrations to co-train policies from scratch or post-train foundation models. Hypothesized mechanism (schematic): a deployed policy imitates the real demonstrations in states they cover, relies on simulated behavior elsewhere, and returns to real-like states once it can.

Abstract

Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. We distinguish world grounding, which aligns simulation with the real system, and behavior grounding, which aligns simulated trajectories with human motion. We build a real2sim2real pipeline that varies these axes independently to generate data for co-training. On a dynamic dexterous pick-and-sort task, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10. Deployed policies behave like a mixture of real-derived and simulation-derived policies, imitating real demonstrations in covered states and relying on simulated behavior elsewhere, which we examine through latent-space analysis. Together, these results suggest complementary roles: world grounding lets policies use simulated experience beyond real-data coverage, while behavior grounding matters mainly when world grounding is imperfect. Grounded simulation remains beneficial when co-training foundation models.

Video

Simulated data along two grounding axes

We generate about 1,500 simulated trajectories for each of four configurations, setting each grounding axis independently. World grounding aligns the simulator with the real system through camera calibration, scene registration and robot system identification. Behavior grounding adapts the human demonstrations to new object poses and belt speeds with MimicGen; ungrounded behavior is synthesized by grasp proposals and a motion planner.

Motion-planned behavior
ungrounded
Human-like behavior
grounded
Grounded world
World grounded
All 3 camera views
Fully grounded
All 3 camera views
Ungrounded world
Ungrounded
All 3 camera views
Behavior grounded
All 3 camera views

Head-camera view at 1× speed, one example trajectory per dataset. All four datasets use the same renderer (Isaac Sim RTX); world grounding changes calibration and dynamics (camera intrinsics and blur, scene and robot poses, controller response), not the rendering style. Differences in world grounding are therefore hard to see in these clips, while differences in behavior show in the motion.

Grounded co-training raises success from 52% to 86%

Each co-trained policy uses the same 100 real demonstrations plus about 1,500 simulated trajectories, differing only in how the simulated data is grounded. Each policy is evaluated over 50 real-world trials: 24 nominal and 26 under shifts in lighting, belt speed or object that the real demonstrations do not cover. Choose a condition to compare.

Success rate, 500M flow-matching policy trained from scratch

Example rollouts

One real-world rollout per policy and condition, seen from the head camera at 1× speed. Each clip shows the policy’s typical outcome: a success where it succeeded in at least half of its trials, a failure otherwise. Choose a condition above to switch clips; “All trials” shows a random condition for each policy.

Real only

Ungrounded

Behavior grounded

World grounded

Fully grounded

Grounded simulation still helps foundation models

Post-training pre-trained models shows the same pattern. Going from 100 to 1,000 real demonstrations helps a lot, so real data is still a bottleneck. With 100 real demonstrations plus fully grounded simulation, our 2B model reaches 90%, at least matching 1,000 real demonstrations.

Success rate, post-trained foundation models

Switching between real-like and sim-like behaviors

Co-trained policies behave like a mixture of real-derived and simulation-derived policies: they imitate the real demonstrations in states those cover, and fall back on simulated behavior elsewhere. Rollouts in our main conditions rarely leave the real data, so to make the switch easier to see we also trained a diagnostic policy with only 10 real demonstrations and 1,500 fully ungrounded simulated trajectories. It grasps with a two-finger pinch and moves like the motion-planned simulation data (left). The latent-space view (right) follows a separate real rollout, showing how a policy’s observations can move from the real region toward the simulated one.

100 real demos + ungrounded sim 10 real demos + ungrounded sim
Real-world rollouts at 1× speed. With 10 real demonstrations the policy picks with a two-finger pinch, as in the motion-planned simulation data.
A separate real rollout (not the one shown on the left): observation embeddings (PCA) along its path (black, orange dot = current frame) against real demonstrations (blue) and simulated demonstrations (red), with the nearest real and simulated frames.

Better world grounding reduces the need for behavior grounding

Behavior grounding helped by a similar amount under both world-grounding configurations, suggesting our calibration-based world grounding was not yet accurate enough. We therefore improved its visual fidelity with CRAFT-inspired video translation. Translated simulated observations overlap the real ones in latent space, while the RTX-rendered observations form a separate cluster. Co-trained with 25 real and 75 simulated trajectories, policies using behavior-grounded and behavior-ungrounded data both succeeded in 17 of 24 trials, on par with our policy trained on 100 real trajectories (16/24).

Observation-embedding projection: Grounded-CRAFT points overlap the Real points, while Grounded-RTX Render points form a separate cluster.
Observation-embedding projection of real, RTX-rendered and CRAFT-translated observations.
Translated simulation (top row) and real camera views (bottom row) for the head, left-wrist and right-wrist cameras.

BibTeX

@misc{liu2026getting,
  title         = {Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training},
  author        = {Liu, Samuel and Kim, Youngsun and Matak, Martin and Lee, Gilwoo},
  year          = {2026},
  eprint        = {2610.00821},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2610.00821}
}