RH20T Dataset
Contact-rich multimodal manipulation paired with human demonstrations
A large multimodal dataset pairing robot trajectories with human demonstrations, force, audio, depth, and tactile signals.
Contact-rich policy learning
Commercial-use rights differ by subset; privacy-sensitive human recordings are included.
Start with an official sample or one task; do not download the full release before schema validation.
Has observation and action/state proxy, but weak task-phase context.
fit 70 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
WAM-required observation, intent, and action alignment is not fully verified.
fit 50 · confidence 15
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · depth · A large multimodal dataset pairing robot trajectories with human demonstrations, force, audio, depth, and tactile signals.
Action / hand pose / robot state
actions · robot state · Contact-rich multimodal manipulation paired with human demonstrations
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
No decision-grade evidence captured yet.
Feedback / correction / failure
No decision-grade evidence captured yet.
Sim-real pairing
depth · A large multimodal dataset pairing robot trajectories with human demonstrations, force, audio, depth, and tactile signals.
License / format / access
License required · Mixed CC BY-SA 4.0 / CC BY-NC 4.0 subsets · MP4 · NumPy
Model and task fit · OpenBot inference
Has observation and action/state proxy, but weak task-phase context.
fit 70 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
WAM-required observation, intent, and action alignment is not fully verified.
fit 50 · confidence 15
Failure, correction, intervention, and recovery annotations have not been verified.
fit 50 · confidence 15
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Language intent / task phaseunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Feedback / correction / failureunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Contact-rich policy learning
- Human-to-robot transfer
- Multimodal world models
Related models and papers
Model references linked to similar loop signals.
Helix 02
A hierarchical whole-body VLA for continuous humanoid locomotion, manipulation, balance, tactile feedback, and long-horizon autonomy.
LingBot-Depth
A masked depth model for refining incomplete sensor depth into metric geometry for reconstruction, tracking, and manipulation.
LingBot-Vision
A family of vision foundation encoders pretrained for dense spatial perception and downstream embodied understanding.
LingBot-Map
A feed-forward 3D foundation model for streaming scene reconstruction, camera-pose estimation, and long-sequence point-cloud generation.
SpatialVLA
A spatially enhanced 4B VLA pretrained on 1.1 million real-robot episodes with explicit geometry-aware representations.
Integration notes
- Commercial-use rights differ by subset; privacy-sensitive human recordings are included.
