EgoSchema Dataset
Long-form egocentric video QA
A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.
Video-language model evaluation
Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · Ego4D video references · Video-language model evaluation · Underlying video access follows Ego4D licensing.
Action / hand pose / robot state
Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · temporal reasoning labels · Video-language model evaluation · A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.
Feedback / correction / failure
Video-language model evaluation
Sim-real pairing
No decision-grade evidence captured yet.
License / format / access
License required · MIT · JSON · Ego4D video references
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Sim-real pairingunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Video-language model evaluation
- Long-horizon plan checking
- Narrative consistency tests for agents
Integration notes
- Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
- Underlying video access follows Ego4D licensing.
