Curatrix
Blog

From raw footage to training episodes: inside the Curatrix pipeline

Segmentation, annotation schema, manifests and formats. What happens between a tradesperson pressing record and a versioned dataset landing in your training run.

1 min readBy Curatrix Research

A day of recording on site yields hours of continuous video that is, on its own, nearly useless for training. The value is created afterwards, in a pipeline that turns footage into episodes with structure a model can learn from. This post walks through that pipeline as it runs for every Curatrix release.

Segmentation into episodes

Footage is first cut into task episodes: one complete execution of the specified task, from a defined start state to a defined end state. Idle time, tool fetching and interruptions are trimmed or labelled as such. Episode boundaries follow the written task specification agreed with the customer, so the same task recorded by ten contributors produces comparable episodes.

Annotation schema

Each episode is annotated with subtasks and phases, the tools and objects in use, and outcome states such as “fastened”, “misaligned” or “retried”. The schema is fixed per engagement and documented alongside the release. Annotators are trained on the trade in question; labelling a torque sequence correctly requires knowing what a torque sequence is.

Manifest and provenance

The manifest is the index of the release. It maps every episode to its video files, sensor streams, annotation records, contributor and site releases and licence scope. Releases are immutable and versioned; a later release adds episodes or corrections without rewriting the earlier one, which keeps experiments reproducible.

Formats

Video ships as MP4, tabular annotations and metadata as Parquet, the manifest as JSON, and synchronised multi-stream recordings as MCAP. Loaders for LeRobot and ROS are on the roadmap. The docs describe each format in detail, and the dataset catalogue shows which formats a given release includes today.

Related articles

All articles
1 min read

Five metrics for judging a robotics dataset before you buy it

Hours of video is the least informative number on a datasheet. Contributor diversity, environment spread, episode completeness, annotation density and rights coverage tell you more.

  • data-quality
  • evaluation
  • physical-ai
1 min read

Egocentric video for manipulation learning: head, helmet or chest camera?

Where the camera sits changes what a model can learn. Trade-offs between head-mounted, helmet-mounted and chest-mounted capture for skilled manual work.

  • field-capture
  • sensors
  • physical-ai
2 min read

What “rights-cleared” training data actually means

Consent, location releases, GDPR and licence scope: the four layers that separate usable robotics data from footage you cannot ship a product on.

  • licensing
  • compliance
  • data-quality
Next step

Start with a scoped pilot

Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.

Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.