Wearable AI Dataset
Egocentric video QA for wearable assistants
A benchmark for long-form QA, conversational QA, and proactive assistant behavior over first-person wearable-camera videos.
Agent evaluation over streaming first-person video
Less robot-action focused, but useful for evaluating an agent's temporal grounding.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · Agent evaluation over streaming first-person video · Conversation-grounded video retrieval · Egocentric video QA for wearable assistants
Action / hand pose / robot state
Less robot-action focused, but useful for evaluating an agent's temporal grounding.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · task phase
Feedback / correction / failure
Agent evaluation over streaming first-person video
Sim-real pairing
No decision-grade evidence captured yet.
License / format / access
Gated · MIT · JSONL · MP4
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Sim-real pairingunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Agent evaluation over streaming first-person video
- Conversation-grounded video retrieval
- Proactive assistant benchmarks
Integration notes
- Less robot-action focused, but useful for evaluating an agent's temporal grounding.
- Access requires accepting dataset terms on Hugging Face.
