Xperience-10M Dataset
Multimodal human experience for embodied AI
A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
World model pretraining
Very large and controlled-access; the practical OpenBot path is metadata indexing plus targeted subset pulls.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · depth · camera pose · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
Action / hand pose / robot state
camera pose · hand pose · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
Feedback / correction / failure
No decision-grade evidence captured yet.
Sim-real pairing
depth · camera pose · Real-to-sim and sim-to-real data alignment · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
License / format / access
Gated · Apache-2.0 · Hugging Face dataset · multimodal episode files
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Failure, correction, intervention, and recovery annotations have not been verified.
fit 50 · confidence 15
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Feedback / correction / failureunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- World model pretraining
- Real-to-sim and sim-to-real data alignment
- Multimodal episode quality checks
Related models and papers
Model references linked to similar loop signals.
GR00T N1
Open foundation model for generalist humanoid robots, focused on whole-body and manipulation behavior from multimodal robot data.
NVIDIA Cosmos
World foundation model platform for physical AI, built around predictive world modeling and data processing workflows.
UniSim
Interactive real-world simulator research that models how visual scenes change under actions and interaction.
Genie 2
Large-scale foundation world model for generating action-controllable interactive environments from visual prompts.
Looped World Models
World-model architecture using looped transformer refinement over latent states for iterative environment understanding.
Integration notes
- Very large and controlled-access; the practical OpenBot path is metadata indexing plus targeted subset pulls.
- Useful as a reference for the signals OpenBot Data should preserve.
