Ego-Exo4D Dataset
Synchronized first-person and third-person skilled activity
A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.
Learning from skilled human demonstrations
Especially relevant when a task needs both wearable camera context and external validation views.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · external camera · Especially relevant when a task needs both wearable camera context and external validation views. · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.
Action / hand pose / robot state
pose · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · JSON annotations · Converting human keysteps into robot task graphs · Especially relevant when a task needs both wearable camera context and external validation views.
Feedback / correction / failure
Cross-view reconstruction and evaluation · A multimodal skilled-activity dataset with time-synced egocentric and exocentric cameras plus language, pose, masks, and proficiency annotations.
Sim-real pairing
Cross-view reconstruction and evaluation
License / format / access
License required · Ego-Exo4D License Agreement · Ego-Exo4D CLI · MP4
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Learning from skilled human demonstrations
- Cross-view reconstruction and evaluation
- Converting human keysteps into robot task graphs
Related models and papers
Model references linked to similar loop signals.
RT-2
Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.
GR00T N1
Open foundation model for generalist humanoid robots, focused on whole-body and manipulation behavior from multimodal robot data.
Gemini Robotics
Embodied reasoning and action model family that brings Gemini capabilities into robotics through vision, language, and physical action.
SayCan
Language-grounded robotics approach that combines language-model planning with affordance scores from robot skills.
NVIDIA Cosmos
World foundation model platform for physical AI, built around predictive world modeling and data processing workflows.
Looped World Models
World-model architecture using looped transformer refinement over latent states for iterative environment understanding.
Integration notes
- Especially relevant when a task needs both wearable camera context and external validation views.
- The expert commentary and keysteps are useful for Data curation labels.
