Curatrix
Blog

Egocentric video for manipulation learning: head, helmet or chest camera?

Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.

1 min readBy Curatrix Research

Egocentric video is the most practical way to record skilled manual work at scale: it is cheap, it captures the worker’s attention as well as their hands, and it does not require instrumenting the site. But “egocentric” hides a real design decision. A camera on the forehead, on a helmet and on the chest produce materially different datasets.

Head-mounted

A head-mounted action camera follows gaze. It captures what the worker is looking at, which is a strong signal for intent and for which object matters next. The downside is motion: head turns produce blur and rapid viewpoint changes, and the hands leave the frame whenever the worker glances away. Best for tasks where attention shifts are part of the skill, such as inspecting a joint before fastening.

Helmet-mounted

On a construction site the helmet is mandatory anyway, so a helmet mount adds no equipment burden and sits slightly higher than the forehead, which widens the view of the workspace. Motion characteristics are similar to head-mounted capture. Helmet mounts are the default for our construction-site and industrial recordings.

Chest-mounted

A chest camera is stable and keeps both hands in frame almost continuously, which makes it the better source for hand-object interaction and for learning grasp and force cues. It loses the gaze signal and is occasionally blocked by the worker’s forearms. For fine assembly at a bench, chest capture usually produces the cleaner episodes.

Use more than one

In practice the strongest datasets combine two views, synchronised at capture time. Each release in our catalogue lists its sensor rig explicitly, so you can filter for the viewpoint your model needs, and the methodology page describes how streams are aligned and segmented into episodes.

Related articles

All articles
2 min read

Why Physical AI needs field data, not lab benchmarks

Robots trained on tidy lab demonstrations fail on real job sites. A practical look at what field-captured manipulation data adds and how to judge it.

  • physical-ai
  • data-quality
  • field-capture
1 min read

Five metrics for judging a robotics dataset before you buy it

Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.

  • data-quality
  • evaluation
  • physical-ai
1 min read

From raw footage to training episodes: inside the Curatrix pipeline

Segmentation, annotation schema, manifests and formats. What happens between a tradesperson pressing record and a versioned dataset landing in your training run.

  • pipeline
  • formats
  • annotation
Next step

Start with a scoped pilot

Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.

Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.