Five metrics for judging a robotics dataset before you buy it
Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.
Dataset marketing leads with hours. Hours are easy to count and easy to inflate: one contributor repeating one task in one room for a week produces a large number and a narrow dataset. Before committing budget to a release, look at five other numbers.
1. Contributor diversity
How many distinct people performed the task? Skilled workers differ in handedness, pace, grip and error-recovery strategy. A policy that has seen twenty electricians generalises to the twenty-first far better than one that has memorised a single person’s habits.
2. Environment spread
Workshop, construction site, residential, industrial: each has its own lighting, clutter and background motion. Check how many sites and environment types are represented and whether that matches where your robot will actually operate.
3. Episode completeness
What fraction of recorded hours became accepted episodes with a clear start and end state? A low ratio is not necessarily bad, since strict acceptance criteria reject a lot, but you should know the ratio and the criteria behind it.
4. Annotation density
Per episode, how many phases, tools, objects and outcome states are labelled, and by whom? Annotations by people who know the trade are worth more than crowd-sourced bounding boxes, and the difference is visible in label consistency across contributors.
5. Rights coverage
Finally: for what share of episodes can the provider show a contributor release, a site authorisation and a privacy review record? Anything below one hundred percent is a liability you inherit. Every Curatrix release publishes these figures in its datasheet; the catalogue shows the headline numbers, and the FAQ explains how they are calculated. If you want to discuss a specific release, get in touch.