MicroAGI01 Dataset
Household manipulation with pose annotations
An egocentric RGB-D household manipulation dataset with camera pose, 3D hand landmarks, task segmentation, and MCAP recordings.
Household skill segmentation
One of the more directly robotics-relevant egocentric datasets in this catalog.
Start with an official sample or one task; do not download the full release before schema validation.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · depth · camera pose · An egocentric RGB-D household manipulation dataset with camera pose, 3D hand landmarks, task segmentation, and MCAP recordings.
Action / hand pose / robot state
camera pose · hand pose · Foxglove layout · Hand-pose action proxy extraction
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
task phase · Household manipulation with pose annotations · An egocentric RGB-D household manipulation dataset with camera pose, 3D hand landmarks, task segmentation, and MCAP recordings.
Feedback / correction / failure
No decision-grade evidence captured yet.
Sim-real pairing
depth · camera pose · An egocentric RGB-D household manipulation dataset with camera pose, 3D hand landmarks, task segmentation, and MCAP recordings.
License / format / access
License required · CC-BY-4.0 · MCAP · CSV
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Failure, correction, intervention, and recovery annotations have not been verified.
fit 50 · confidence 15
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Feedback / correction / failureunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Household skill segmentation
- Hand-pose action proxy extraction
- Foxglove-to-OpenBot episode conversion
Related models and papers
Model references linked to similar loop signals.
RT-2
Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.
RT-1
Robotics Transformer policy trained on large-scale real-world robot demonstrations for language-conditioned manipulation.
RT-X
Cross-embodiment RT model family trained from the Open X-Embodiment mixture to study transfer across robots and tasks.
RDT-1B
Diffusion foundation model for bimanual manipulation that uses large-scale robot data to generate action trajectories.
ACT
Action Chunking with Transformers predicts short action sequences for efficient imitation learning in manipulation tasks.
RoboCat
Self-improving generalist robotic agent that collects new demonstrations to improve its own manipulation capabilities.
Integration notes
- One of the more directly robotics-relevant egocentric datasets in this catalog.
- Good target for an OpenBot Data MCAP adapter.
