Curatrix
Blog

Notes on real-world training data for Physical AI

Field capture, rights, annotation and evaluation — practical writing from the team building robotics datasets with German tradespeople.

1 min read

Five metrics for judging a robotics dataset before you buy it

Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.

  • data-quality
  • evaluation
  • physical-ai
1 min read

From raw footage to training episodes: inside the Curatrix pipeline

Segmentation, annotation schema, manifests and formats. What happens between a tradesperson pressing record and a versioned dataset landing in your training run.

  • pipeline
  • formats
  • annotation
1 min read

Egocentric video for manipulation learning: head, helmet or chest camera?

Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.

  • field-capture
  • sensors
  • physical-ai
2 min read

What “rights-cleared” training data actually means

Consent, location releases, GDPR and licence scope: the four layers that separate usable robotics data from footage you cannot ship a product on.

  • licensing
  • compliance
  • data-quality
2 min read

Why Physical AI needs field data, not lab benchmarks

Robots trained on tidy lab demonstrations fail on real job sites. A practical look at what field-captured manipulation data adds and how to judge it.

  • physical-ai
  • data-quality
  • field-capture

Looking for a quick answer instead? The FAQ covers capture, rights, formats and pricing. FAQ →

Next step

Start with a scoped pilot

Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.

Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.