Why Physical AI needs field data, not lab benchmarks
Robots trained on tidy lab demonstrations fail on real job sites. A practical look at what field-captured manipulation data adds and how to judge it.
Most manipulation datasets are recorded where it is convenient to record: on a bench, under fixed lighting, with the same few objects laid out the same way. That is fine for a benchmark. It is a poor proxy for an electrician routing cable through a half-finished building, where the light changes, the workpiece is never square and the next tool is wherever it was dropped.
What the field adds
Field data contributes three things a lab cannot reproduce at scale. First, variation that is correlated the way it is in reality: worn tools, partial occlusion by the worker’s own hands, clutter that changes as the task progresses. Second, skill. A tradesperson with years of experience performs a fastening sequence differently from a research assistant following a script, and the difference shows up in grip choice, force and recovery from small errors. Third, long-horizon structure: real tasks are chains of subtasks with decisions between them, not isolated pick-and-place episodes.
The cost of training on the wrong distribution
Policies trained on narrow data are not wrong on average; they are brittle at the edges, and the edges are where deployment happens. Teams often discover this only after a pilot, when a model that scored well on held-out lab episodes stalls the first time a cable drum is positioned differently. Collecting field data up front is cheaper than discovering the gap in production.
How to judge a field dataset
Ask four questions. How many distinct contributors and sites are represented, not just how many hours? Is the task segmented into episodes with labelled phases, or is it raw footage you will have to cut yourself? Are the usage rights explicit for every clip, including the people and places in frame? And can you reproduce a given release later, with the same manifest and the same annotations?
Those are the criteria we build Curatrix releases around. Our dataset catalogue lists contributor counts, environments and licence scope for each release, and the methodology page documents how every recording becomes a versioned dataset.