OpenBot
Back to datasets
Egocentric datasetLicense requiredReadiness 81 · confidence 43

EgoMM Dataset

Video + audio + IMU egocentric clips

463hours

A tri-modal egocentric dataset built from EgoLife and Ego-Exo4D sources, packaged into fixed-duration clips with optional narrations.

Best for

Wearable-assistant evaluation

Not for / blocker

A practical bridge between video-only egocentric data and robot-ready sensor streams.

Download decision

Inspect schema and run a bounded sample audit before committing to the full release.

Policy learningUnknown

Observation/action alignment has not been verified.

fit 50 · confidence 15

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUnknown

WAM-required observation, intent, and action alignment is not fully verified.

fit 50 · confidence 15

Verified facts and provenance

Claims, metadata verification, and sample verification are shown separately.

curated official source
Official claim · signals
VideoAudioIMUIMUNarration
Metadata verified · schema / annotations

Unknown — no machine-readable schema facts have been captured.

Sample / pipeline verification

Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.

Declared loop signal coverage

Signals inferred from official metadata; Data pipeline verification is still pending.

5/7 categories present or partial

Observation / ego video

video · A practical bridge between video-only egocentric data and robot-ready sensor streams. · Video + audio + IMU egocentric clips

present

Action / hand pose / robot state

No decision-grade evidence captured yet.

unknown

Gaze / attention

No decision-grade evidence captured yet.

unknown

Language intent / task phase

narration · A tri-modal egocentric dataset built from EgoLife and Ego-Exo4D sources, packaged into fixed-duration clips with optional narrations.

present

Feedback / correction / failure

Wearable-assistant evaluation

present

Sim-real pairing

Long-sequence reconstruction from clips

partial

License / format / access

License required · Apache-2.0 · MP4 · MP3

present

Model and task fit · OpenBot inference

Policy learningUnknown

Observation/action alignment has not been verified.

fit 50 · confidence 15

World modelUseful

Has visual observations plus geometry, calibration, depth, or reconstruction cues.

fit 82 · confidence 55

WAMUnknown

WAM-required observation, intent, and action alignment is not fully verified.

fit 50 · confidence 15

Failure miningUseful

Has failure/evaluation-style labels, but limited action trace detail.

fit 68 · confidence 55

Good tasks

3D / sim-real alignment

Blockers and unresolved evidence

  • Action / hand pose / robot stateunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.
  • Gaze / attentionunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.

Raw dataset signals

VideoAudioIMUIMUNarration

OpenBot fit

  • Wearable-assistant evaluation
  • IMU-aware activity recognition
  • Long-sequence reconstruction from clips

Integration notes

  • A practical bridge between video-only egocentric data and robot-ready sensor streams.
  • Good for testing OpenBot Data's multimodal schema adapters.

Related by signals