# Annotation examples

## HoloAssist egocentric example

Video, hand tracking, camera calibration and action annotations: **HoloAssist**, Xin Wang, Taein Kwon et al., ICCV 2023. [Official dataset](https://holoassist.github.io/). Episode `R007-7July-DSLR`, a participant handling a DSLR camera and lens with bare hands. No wrist-mounted marker rig is visible in the reviewed source and sequence samples. No endorsement implied.

The official dataset release uses [Community Data License Agreement, Permissive, Version 2.0](https://cdla.dev/permissive-2-0/). It permits use, modification and sharing, with the full license text made available alongside redistributed data. The complete text accompanies these files as [CDLA-Permissive-2.0.txt](CDLA-Permissive-2.0.txt).

`holoassist-source.mp4` and `holoassist-overlay.mp4` contain the same **1,098 original frames**, 896 × 504, at **680108/22669 fps** (approximately 30.001676 fps), for **36.5979550306716 seconds**. Source frames 369–1466 cover 12.299312756–48.897267787 seconds. Both videos are silent H.264 encodes without cropping, mirroring, temporal interpolation or speed changes. The two posters show the same source frame, 420.

The overlay projects **26 supplied joints per hand** through the supplied camera pose and RGB intrinsics. Anatomical left is sage `#8bc39d`; right is warm ivory `#eee8d8`. These are dataset-provided HoloLens hand-tracking results, not RoboLens pose inference or human-corrected ground truth. Only active, valid, tracked joints with finite projections in front of the camera are drawn. No smoothing, inferred points or filled tracking gaps were added. Supplied tracked joints can remain behind an occluding object. Both hands briefly lose drawable tracking at clip 30.365–31.065 seconds, with additional per-hand gaps retained in the evidence report.

[holoassist-actions.json](holoassist-actions.json) preserves five source coarse actions and sixteen fine actions, original event IDs and attributes, source times and clip-relative times. Display labels shorten the original source sentences. The five task steps are Attach lens, Detach lens cover, Attach lens cover, Turn DSLR on and Turn DSLR off. The `start` and `end` fields are exact aliases of the retained `clip_start_seconds` and `clip_end_seconds`; all source gaps and the final 15.268 ms unlabeled tail remain. Event 30's original noun is inconsistent with its own sentence, so both original fields are preserved. [The untouched episode action record](holoassist-original-actions.json) is also available.

[holoassist-hand-points.json](holoassist-hand-points.json) is a separate download containing all 1,098 frames, original unrounded 3D joint translations, calibrated 2D projections, camera transforms, timestamps, joint ordering, validity/tracking flags and source SHA-256 hashes. The projection convention follows the dataset author's [reference implementation](https://github.com/taeinkwon/PyHoloAssist).

Prior verification checked every decoded frame for raw/overlay timing alignment and every supplied 3D value, quality flag and 2D projection against source data. Contact sheets sampled the complete encoded sequence and action boundaries (75 sequence samples and 27 boundary samples). That visual review is sampled AI review, not a measured hand-accuracy benchmark. Current publication checks verify that the copied media and point data match the previously reviewed hashes. Provenance, review reports, contact sheets and reproduction instructions are in `artifacts/annotation/holoassist-evidence/`; reproduction scripts are in `scripts/` and require explicit source/output paths.

## DROID robot example

Footage: DROID Dataset Team, [DROID](https://droid-dataset.github.io/), licensed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/); the data license is stated in the [original dataset paper](https://www.jiajunwu.com/papers/droid_rss.pdf). No endorsement implied.

Episode: `PennPAL+acda9df3+2023-06-23-19h-56m-41s`, camera `ext2_left_eye:27085680`. The 25-second excerpt uses normalized source time 17–42 seconds, with 375 frames at 15 fps and 960 × 540 resolution. `droid-actions.json` retains the original resource URL, timing and hashes.

Action proposals were generated using Gemini 3.8 Flash with samples at 10 fps, then reviewed and corrected by an AI agent against all 375 displayed frames. These are AI annotations, not supplied DROID action labels or human ground truth. See [droid-actions.json](droid-actions.json) for evidence and limits.
