OpenBot
Back to datasets
Egocentric datasetLicense requiredReadiness 83 · confidence 51

Ego-Exo4D Dataset

Synchronized first-person and third-person skilled activity

1,422hours

A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.

Best for

Learning from skilled human demonstrations

Not for / blocker

Especially relevant when a task needs both wearable camera context and external validation views.

Download decision

Inspect schema and run a bounded sample audit before committing to the full release.

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Verified facts and provenance

Claims, metadata verification, and sample verification are shown separately.

curated official source
Official claim · signals
VideoExternal CameraAudioPoseObject masksLanguage
Metadata verified · schema / annotations

Unknown — no machine-readable schema facts have been captured.

Sample / pipeline verification

Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.

Declared loop signal coverage

Signals inferred from official metadata; Data pipeline verification is still pending.

6/7 categories present or partial

Observation / ego video

video · external camera · Especially relevant when a task needs both wearable camera context and external validation views. · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.

present

Action / hand pose / robot state

pose · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.

partial

Gaze / attention

No decision-grade evidence captured yet.

unknown

Language intent / task phase

language · JSON annotations · Converting human keysteps into robot task graphs · Especially relevant when a task needs both wearable camera context and external validation views.

present

Feedback / correction / failure

Cross-view reconstruction and evaluation · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.

present

Sim-real pairing

Cross-view reconstruction and evaluation

partial

License / format / access

License required · Ego-Exo4D License Agreement · Ego-Exo4D CLI · MP4

present

Model and task fit · OpenBot inference

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Failure miningUseful

Has failure/evaluation-style labels with action or manipulation context.

fit 82 · confidence 55

Good tasks

video-language reasoning3D / sim-real alignment

Blockers and unresolved evidence

  • Gaze / attentionunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.

Raw dataset signals

VideoExternal CameraAudioPoseObject masksLanguage

OpenBot fit

  • Learning from skilled human demonstrations
  • Cross-view reconstruction and evaluation
  • Converting human keysteps into robot task graphs

Related models and papers

Model references linked to similar loop signals.

Integration notes

  • Especially relevant when a task needs both wearable camera context and external validation views.
  • The expert commentary and keysteps are useful for Data curation labels.

Related by signals