Notes on real-world training data for Physical AI
Field capture, rights, annotation and evaluation — practical writing from the team building robotics datasets with German tradespeople.
Five metrics for judging a robotics dataset before you buy it
Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.
- data-quality
- evaluation
- physical-ai
From raw footage to training episodes: inside the Curatrix pipeline
Segmentation, annotation schema, manifests and formats. What happens between a tradesperson pressing record and a versioned dataset landing in your training run.
- pipeline
- formats
- annotation
Egocentric video for manipulation learning: head, helmet or chest camera?
Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.
- field-capture
- sensors
- physical-ai
What “rights-cleared” training data actually means
Consent, location releases, GDPR and licence scope: the four layers that separate usable robotics data from footage you cannot ship a product on.
- licensing
- compliance
- data-quality
Why Physical AI needs field data, not lab benchmarks
Robots trained on tidy lab demonstrations fail on real job sites. A practical look at what field-captured manipulation data adds and how to judge it.
- physical-ai
- data-quality
- field-capture
Looking for a quick answer instead? The FAQ covers capture, rights, formats and pricing. FAQ →
Start with a scoped pilot
Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.
Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.