OpenBot
Back to datasets
Egocentric datasetGatedReadiness 81 · confidence 43

Wearable AI Dataset

Egocentric video QA for wearable assistants

2,100clips

A benchmark for long-form QA, conversational QA, and proactive assistant behavior over first-person wearable-camera videos.

Best for

Agent evaluation over streaming first-person video

Not for / blocker

Less robot-action focused, but useful for evaluating an agent's temporal grounding.

Download decision

Inspect schema and run a bounded sample audit before committing to the full release.

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has rich observation and semantic context, but limited geometry/sim-real alignment.

fit 68 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Verified facts and provenance

Claims, metadata verification, and sample verification are shown separately.

curated official source
Official claim · signals
VideoLanguageLanguageLanguageTask phase
Metadata verified · schema / annotations

Unknown — no machine-readable schema facts have been captured.

Sample / pipeline verification

Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.

Declared loop signal coverage

Signals inferred from official metadata; Data pipeline verification is still pending.

5/7 categories present or partial

Observation / ego video

video · Agent evaluation over streaming first-person video · Conversation-grounded video retrieval · Egocentric video QA for wearable assistants

present

Action / hand pose / robot state

Less robot-action focused, but useful for evaluating an agent's temporal grounding.

partial

Gaze / attention

No decision-grade evidence captured yet.

unknown

Language intent / task phase

language · task phase

present

Feedback / correction / failure

Agent evaluation over streaming first-person video

present

Sim-real pairing

No decision-grade evidence captured yet.

unknown

License / format / access

Gated · MIT · JSONL · MP4

present

Model and task fit · OpenBot inference

Policy learningUseful

Has observation, action/state proxy, and task or language context.

fit 85 · confidence 55

World modelUseful

Has rich observation and semantic context, but limited geometry/sim-real alignment.

fit 68 · confidence 55

WAMUseful

Contains observation, intent, action/state, and feedback-like supervision.

fit 88 · confidence 55

Failure miningUseful

Has failure/evaluation-style labels with action or manipulation context.

fit 82 · confidence 55

Good tasks

video-language reasoning

Blockers and unresolved evidence

  • Gaze / attentionunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.
  • Sim-real pairingunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.

Raw dataset signals

VideoLanguageLanguageLanguageTask phase

OpenBot fit

  • Agent evaluation over streaming first-person video
  • Conversation-grounded video retrieval
  • Proactive assistant benchmarks

Integration notes

  • Less robot-action focused, but useful for evaluating an agent's temporal grounding.
  • Access requires accepting dataset terms on Hugging Face.

Related by signals