Curatrix
Blog

Why Physical AI needs field data, not lab benchmarks

Robots trained on tidy lab demonstrations fail on real job sites. A practical look at what field-captured manipulation data adds and how to judge it.

2 min readBy Curatrix Research

Most manipulation datasets are recorded where it is convenient to record: on a bench, under fixed lighting, with the same few objects laid out the same way. That is fine for a benchmark. It is a poor proxy for an electrician routing cable through a half-finished building, where the light changes, the workpiece is never square and the next tool is wherever it was dropped.

What the field adds

Field data contributes three things a lab cannot reproduce at scale. First, variation that is correlated the way it is in reality: worn tools, partial occlusion by the worker’s own hands, clutter that changes as the task progresses. Second, skill. A tradesperson with years of experience performs a fastening sequence differently from a research assistant following a script, and the difference shows up in grip choice, force and recovery from small errors. Third, long-horizon structure: real tasks are chains of subtasks with decisions between them, not isolated pick-and-place episodes.

The cost of training on the wrong distribution

Policies trained on narrow data are not wrong on average; they are brittle at the edges, and the edges are where deployment happens. Teams often discover this only after a pilot, when a model that scored well on held-out lab episodes stalls the first time a cable drum is positioned differently. Collecting field data up front is cheaper than discovering the gap in production.

How to judge a field dataset

Ask four questions. How many distinct contributors and sites are represented, not just how many hours? Is the task segmented into episodes with labelled phases, or is it raw footage you will have to cut yourself? Are the usage rights explicit for every clip, including the people and places in frame? And can you reproduce a given release later, with the same manifest and the same annotations?

Those are the criteria we build Curatrix releases around. Our dataset catalogue lists contributor counts, environments and licence scope for each release, and the methodology page documents how every recording becomes a versioned dataset.

Related articles

All articles
1 min read

Five metrics for judging a robotics dataset before you buy it

Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.

  • data-quality
  • evaluation
  • physical-ai
1 min read

Egocentric video for manipulation learning: head, helmet or chest camera?

Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.

  • field-capture
  • sensors
  • physical-ai
2 min read

What “rights-cleared” training data actually means

Consent, location releases, GDPR and licence scope: the four layers that separate usable robotics data from footage you cannot ship a product on.

  • licensing
  • compliance
  • data-quality
Next step

Start with a scoped pilot

Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.

Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.