Ego4D Dataset
Large-scale daily-life egocentric video
A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.
Pretraining egocentric perception
Best treated as a large source corpus rather than a direct robot-action dataset.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · stereo · Large-scale daily-life egocentric video · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.
Action / hand pose / robot state
Mining object interaction priors · Best treated as a large source corpus rather than a direct robot-action dataset. · Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data. · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.
Gaze / attention
gaze
Language intent / task phase
narration · JSON annotations · Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data. · A broad first-person video benchmark for daily activities, long-horizon understanding, narration, object interaction, and temporal memory.
Feedback / correction / failure
Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data.
Sim-real pairing
3d scan
License / format / access
License required · Ego4D License Agreement · Ego4D CLI · MP4
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
No major loop category is completely missing. Check quality and alignment before training.
Raw dataset signals
OpenBot fit
- Pretraining egocentric perception
- Long-horizon activity understanding
- Mining object interaction priors
Related models and papers
Model references linked to similar loop signals.
Gemini Robotics
Embodied reasoning and action model family that brings Gemini capabilities into robotics through vision, language, and physical action.
Looped World Models
World-model architecture using looped transformer refinement over latent states for iterative environment understanding.
LingBot-World
An open interactive world simulator derived from video generation, with camera- and action-conditioned variants and long-horizon generation.
LingBot-Map
A feed-forward 3D foundation model for streaming scene reconstruction, camera-pose estimation, and long-sequence point-cloud generation.
Integration notes
- Best treated as a large source corpus rather than a direct robot-action dataset.
- Useful for building retrieval, narration, hand-object, and failure-mining tasks around OpenBot Data.
