OpenBot
Back to datasets
Egocentric datasetLicense requiredReadiness 95 · confidence 60

Ego4D Dataset

Large-scale daily-life egocentric video

3,670+hours

A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.

Best for

Pretraining egocentric perception

Not for / blocker

Best treated as a large source corpus rather than a direct robot-action dataset.

Download decision

Inspect schema and run a bounded sample audit before committing to the full release.

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Verified facts and provenance

Claims, metadata verification, and sample verification are shown separately.

curated official source
Official claim · signals
VideoAudioNarrationGaze3D scanStereo
Metadata verified · schema / annotations

Unknown — no machine-readable schema facts have been captured.

Sample / pipeline verification

Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.

Declared loop signal coverage

Signals inferred from official metadata; Data pipeline verification is still pending.

7/7 categories present or partial

Observation / ego video

video · stereo · Large-scale daily-life egocentric video · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.

present

Action / hand pose / robot state

Mining object interaction priors · Best treated as a large source corpus rather than a direct robot-action dataset. · Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data. · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.

partial

Gaze / attention

gaze

present

Language intent / task phase

narration · JSON annotations · Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data. · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.

present

Feedback / correction / failure

Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data.

present

Sim-real pairing

3d scan

present

License / format / access

License required · Ego4D License Agreement · Ego4D CLI · MP4

present

Model and task fit · OpenBot inference

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Failure miningUseful

Has failure/evaluation-style labels with action or manipulation context.

fit 82 · confidence 55

Good tasks

3D / sim-real alignmentlong-horizon planningfailure and recovery mining

Blockers and unresolved evidence

No major loop category is completely missing. Check quality and alignment before training.

Raw dataset signals

VideoAudioNarrationGaze3D scanStereo

OpenBot fit

  • Pretraining egocentric perception
  • Long-horizon activity understanding
  • Mining object interaction priors

Related models and papers

Model references linked to similar loop signals.

Integration notes

  • Best treated as a large source corpus rather than a direct robot-action dataset.
  • Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data.

Related by signals