Egocentric video for manipulation learning: head, helmet or chest camera?
Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.
Egocentric video is the most practical way to record skilled manual work at scale: it is cheap, it captures the worker’s attention as well as their hands, and it does not require instrumenting the site. But “egocentric” hides a real design decision. A camera on the forehead, on a helmet and on the chest produce materially different datasets.
Head-mounted
A head-mounted action camera follows gaze. It captures what the worker is looking at, which is a strong signal for intent and for which object matters next. The downside is motion: head turns produce blur and rapid viewpoint changes, and the hands leave the frame whenever the worker glances away. Best for tasks where attention shifts are part of the skill, such as inspecting a joint before fastening.
Helmet-mounted
On a construction site the helmet is mandatory anyway, so a helmet mount adds no equipment burden and sits slightly higher than the forehead, which widens the view of the workspace. Motion characteristics are similar to head-mounted capture. Helmet mounts are the default for our construction-site and industrial recordings.
Chest-mounted
A chest camera is stable and keeps both hands in frame almost continuously, which makes it the better source for hand-object interaction and for learning grasp and force cues. It loses the gaze signal and is occasionally blocked by the worker’s forearms. For fine assembly at a bench, chest capture usually produces the cleaner episodes.
Use more than one
In practice the strongest datasets combine two views, synchronised at capture time. Each release in our catalogue lists its sensor rig explicitly, so you can filter for the viewpoint your model needs, and the methodology page describes how streams are aligned and segmented into episodes.