# Findings from the five 6thsense examples

The footage is a useful starting point for **hierarchical task understanding and
temporal annotation**. Large state changes and task phases are visible across
all five recordings. Detailed labels need different treatment for large-object
manipulation, repeated transfers, and fine contact work.

The most useful next step is to review this annotation draft and agree on action
granularity before comparing automated annotators. The exploration has produced
an evidence-backed specification and review artifact, not an automated semantic
pipeline or a measured model benchmark.

## Delivered

| Clip | Duration | Subtasks | Action spans | Interpretation from visible evidence |
|---|---:|---:|---:|---|
| Bagging | 41.966 s | 5 | 15 | Remove a filled waste-bin liner, manipulate its handles, and fit a new liner |
| Fitting board | 63.598 s | 5 | 15 | Align an electronic assembly and housing, press parts together, and apply a driver |
| Folding | 64.364 s | 8 | 21 | Arrange/fold a black garment and light trousers, then place both in a suitcase |
| Bin picking | 84.664 s | 5 | 29 | Transfer green modules, rearrange trays, then transfer black housings |
| Soldering | 133.553 s | 8 | 30 | Enter/prepare a workstation, work a lead, apply an iron, and manipulate a board assembly |
| **Total** | **388.145 s / 6m 28s** | **31** | **110** | |

These counts describe the draft's segmentation choices. They are not physical
event counts or proof of exhaustive annotation. In particular, several bin-picking
spans group retrieval, orientation and placement; they do not establish an exact
part count. One soldering support span intentionally overlaps iron application.

## What the recordings establish

- All five files are H.264 videos at 1920 × 600, with no audio stream.
- Visual inspection shows two side-by-side views. Each half is 960 × 600. Treating
  the full image as one wide camera view would duplicate scene content.
- The baseline review uses the left half. Full paired frames remain available;
  no calibrated stereo reconstruction or systematic right-view comparison was run.
- The stream frame rate is 3125/104, approximately **30.048 fps**, not exactly 30.
  Every-30-frame samples are separated by **0.9984 seconds**. All annotation
  evidence references use original presentation timestamps, not `frame/30`.
- A total of **11,663 frames** were inspected by the decoding/timestamp tools.
  The media has regular, strictly increasing PTS and no reported FFmpeg decoding
  errors. This does not verify capture-clock accuracy, camera synchronization,
  or absence of dropped frames before encoding.
- The five source files total **459,617,520 bytes**. Source SHA-256 hashes are
  recorded and verified to confirm the original files were preserved.

## Findings by task

**Bagging.** The filename alone is underspecified: this looks like bin-liner
replacement, not packing a manufactured item. The sequence has a strong visible
before/after structure. A closure/tying motion is visible, but exact knot topology,
closure security, and leakage are not established. A useful annotation says
"cross and pull handles" with a tying interpretation recorded as uncertain.

**Fitting board.** Coarse assembly steps are visible, and the hands/tool can be
assigned useful roles. The small fitting is partly covered. "Apply driver at
housing" is supportable; "tighten screw to specification" is not. The clip also
contains a long ending with the driver held and no additional assembly visible.
That should remain a labeled pause, not be silently removed.

**Folding.** This is the clearest starting example for agreeing on annotation
detail. Two distinguishable garments move from an unfolded condition to compact
bundles in luggage. Reorientation and repeated spreading are observable, but do
not automatically mean failed attempts. Define whether each adjustment is an
action or part of an "arrange garment" span before measuring segmentation error.

**Bin picking.** The important structure is two material-transfer phases with a
tray change between them. It supports testing repetitive-cycle segmentation,
object groups, and source/destination relations. A one-second sample may miss an
entire grasp or release. Individual identity is not established for similar
modules. Use group IDs until tracking or denser review supports more detail.

**Soldering.** The video includes doorway navigation, bench preparation, wire
work, equipment handling, and board work. Labeling the entire recording as
"soldering" would hide useful structure. The initial doorway is dark; in the
base samples, mean grayscale intensity falls below 50/255 between approximately
1 and 5 seconds. That is a descriptive pixel statistic, not a calibrated reject
criterion. Gloves obscure the connection region during later iron application.
Neither joint quality nor electrical continuity is visible. Keep apparent
soldering and verified success separate.

## What denser review changed

The first pass used **395 samples** covering all clips, every 30 frames plus the
final frame. Five short windows added **88 samples** at an eight-frame stride
(approximately 3.76 samples/s). Three overlapped base samples, leaving **480 unique
source frames**, approximately 4.12% of the decoded frames. Contact sheets made
these frames reviewable; this is not a claim of continuous framewise inspection.

Seven transitions were revised using the dense images. Examples:

| Transition | Initial draft boundary | Revised midpoint | Observed bracket |
|---|---:|---:|---:|
| Bag opening → lowering replacement liner | 35.500 s | 36.209 s | 36.076–36.342 s |
| First green-module orientation → placement motion | 4.500 s | 4.293 s | 4.160–4.426 s |
| Carry black garment → lower it inside suitcase | 35.800 s | 35.277 s | 35.144–35.410 s |
| Driver alignment → apparent application | 32.000 s | 32.082 s | 31.949–32.215 s |

These are **revisions within one assistant's review**, not measured improvements
against independent ground truth. A narrower image bracket does not prove that
the chosen action definition is correct. In soldering, extra images did not
resolve the hidden contact well enough to narrow its semantic onset confidently.

This suggests a practical policy to test: broad pass for task structure, detailed
pass for action spans, and targeted extra frames for boundaries that matter.
Local measurements can help choose candidate intervals; they cannot decide what
the person intended or whether a hidden connection succeeded.

## Annotation product to aim for first

Retain episode summaries and inferred goals separately from the assigned task
instruction, which is unknown here. Represent subtasks as a complete timeline,
including preparation and pauses. Let action tracks overlap for bimanual work.
Use stable IDs for distinguishable objects and group IDs for indistinguishable
parts. Record before/after states only when visible. Attach frame evidence,
boundary uncertainty and review status to every span.

Do not include force, torque, calibrated trajectories, temperatures or verified
functional success in this initial semantic product. The videos do not provide
those measurements. Wrist markers and gloves are visible, but no associated
tracking measurements were supplied with these five files.

## Recommended next experiment

1. **Review folding first.** Correct the draft and decide whether unfolding,
   regripping and smoothing belong in separate actions. Then review one assembly
   and one repetitive-transfer example using the same definitions.
2. **Create independent reference labels.** Preserve reviewer decisions, boundary
   ranges and disagreements. These assistant labels must not be the sole basis
   for judging an assistant/model annotator's accuracy.
3. **Run three evidence conditions once a local visual model is available:** a
   coarse contextual pass; a roughly one-second sequence; and that sequence with
   denser targeted windows. Use the same model, schema and clips so the comparison
   isolates evidence sampling. Do not assume a generic model's timestamps are
   accurate without checking them against the source.
4. **Measure separate errors:** task/subtask correctness, missed actions, invented
   actions, hand/object assignment, and boundary error at an agreed granularity.
   Include human correction time. Keep mechanical/functional outcomes unknown
   unless separate evidence exists.
5. **Reserve new sessions for generalization.** Bagging, fitting board and folding
   informed the initial protocol; bin picking and soldering were inspected after
   it was drafted. All five are now explored. They are not an untouched held-out
   test, and several appear to share a workspace and capture setup.

Avoid expanding to an elaborate multi-model pipeline until this test exposes
specific errors that a detector, tracker or additional view could address.

## Execution and verification limits

Only existing local filesystem, Python, FFmpeg, image-viewing and Node tools were
used in this run. No additional service was called, no paid model request was
made, no dependency/model was downloaded, and no permission prompt was needed.
Semantic labels were authored by the assistant in this session from extracted
images; there was **no locally installed semantic-model inference run**. Therefore
this report provides no local-model throughput, accuracy or inference-cost claim.

The supplied integrity checker verifies source hashes, source timestamps,
evidence-file existence, full subtask coverage, parent containment, and uncertainty
ranges. It also rejects five deliberately invalid variants. Twelve core Node tests
cover timeline overlap, frame selection, step navigation, editing, invalid input
and embedded data. Ten further tests exercise the app's event handlers in a
minimal DOM harness, including live labels, bounded replay and preserving edits.
JavaScript syntax and local HTML links are checked. These tests do not exercise
actual browser rendering or video decoding.

Browser rendering, interactive video playback, local-storage persistence and
download behavior have not been tested in a browser in this restricted session.
The available interactive-browser workflow requires a tool that is not enabled;
it was not used. A static HTML review is supplied as an additional access path.

All annotation statuses remain **pending human review**. No accuracy percentage
or per-label acceptance rate would be justified by this exploration alone.
