Curatrix
Methodology

How a task demonstration becomes a robotics dataset

Every Curatrix dataset follows the same documented route: a defined task, an authorized capture, privacy processing, structured annotation, a rights review, and a versioned release you can train against and reproduce.

End-to-end pipeline

Eight stages, one protocol

The pipeline is fixed. What changes per engagement is the task specification, the capture setup and the annotation schema.

  1. 01

    Customer defines a task

    You describe the manipulation task, the target environment and how the model will be evaluated. We turn that into a written data specification.

  2. 02

    Authorized tradespeople demonstrate it

    Qualified tradespeople perform the task at approved sites, under an agreed protocol, wearing head, helmet or chest-mounted cameras.

  3. 03

    Footage securely uploaded

    Recordings are transferred over an encrypted channel into access-restricted EU-region storage and registered against the capture session.

  4. 04

    Privacy and confidentiality processing

    Automated detection removes personal and confidential material, followed by a human privacy review of every clip.

  5. 05

    Task segmentation and annotation

    Footage is cut into task episodes and annotated with subtasks, phases, tools, objects and outcome states.

  6. 06

    Quality and rights review

    Episodes are checked against the acceptance criteria, and contributor releases, site authorizations and licence scope are confirmed.

  7. 07

    Versioned robotics dataset

    An immutable release is built with manifests, annotations, provenance records, licence terms and documentation.

  8. 08

    Customer training and evaluation

    You train and evaluate against the agreed benchmark. Gaps feed the next targeted collection round.

Capture

Capture protocols and controlled work environments

Nothing is recorded opportunistically. Each dataset starts from a written protocol and an authorized site.

  • Written capture protocol

    Camera placement, framing, lighting conditions, tool set, episode start and stop conditions, and what counts as a valid attempt are fixed before the first recording.

  • Controlled work environments

    Capture takes place in workshops, industrial facilities and sites where the responsible party has approved recording in writing and the scene can be cleared of unrelated material.

  • Worker authorization

    Tradespeople opt in per project, know what is recorded and what it is used for, can stop a recording at any time, and are paid for their time.

  • Rig and calibration

    Head, helmet and chest mounts are calibrated per contributor and session. Intrinsics and mount geometry are recorded alongside the footage.

Processing

Privacy and confidentiality processing

Personal and confidential material is removed before any clip becomes eligible for a release.

  • Secure ingest

    Footage is uploaded over an encrypted channel into EU-region storage. Raw material is access-restricted from the moment it lands.

  • Automated detection

    Faces, bystanders, screens, documents, licence plates, badges and identifiable signage are detected automatically across every frame.

  • Human privacy review

    A reviewer checks each clip after automated processing. Clips that cannot be cleared are removed rather than shipped with a caveat.

  • Site confidentiality

    Where a site owner marks material as confidential, the affected segments are excluded from every release regardless of privacy status.

Structure

Task segmentation and annotation

An episode is a complete attempt at one thing, described in enough structure to train against and to debug against.

  • Episode segmentation

    Continuous footage is cut into task episodes with explicit start and end conditions, each one a single attempt at the specified task.

  • Subtask and phase labels

    Episodes are broken into subtasks and phases — approach, grasp, align, actuate, verify, release — with frame-accurate boundaries.

  • Tool and object metadata

    Tools, workpieces, fixtures and consumables in view are labelled by type and, where it matters, by size or specification.

  • Success, failure and recovery

    Episodes are labelled by outcome. Failure and recovery episodes are kept deliberately: they are the hardest demonstrations to collect and the most useful to learn from.

Quality assurance

Acceptance criteria, not best effort

Material that does not meet the criteria is dropped rather than delivered with a disclaimer.

  • Explicit acceptance criteria

    Every episode is checked against the protocol for framing, occlusion, motion blur, completeness of the attempt and label agreement.

  • Independent re-annotation

    A sample of every batch is annotated a second time by a different annotator and compared. Disagreement above the agreed threshold sends the batch back.

  • Rejected material is not delivered

    Hours that fail review are excluded from the release and are not counted in its accepted-hours figure.

  • Schema validation

    Manifests, annotations, calibration and provenance records are validated against the published schema before a release is built.

Versioning and provenance

Every release is reproducible and traceable

A training run should always be attributable to an exact dataset state, and every episode to a documented capture.

  • Immutable versioned releases

    A published release never changes. Corrections and additions ship as a new version with a changelog and a stable identifier.

  • Per-episode provenance

    Each episode carries its capture date range, environment class, protocol version, processing pipeline version and annotation schema version.

  • Rights review

    Contributor releases, site authorizations and licence scope are reviewed and recorded against every episode before a release is built.

  • Disclosure logging

    Every delivery is logged with recipient, scope, dataset version and time.

Evaluation

Judged on model behaviour, not hours delivered

A dataset is only useful if it moves a benchmark, so the benchmark is agreed before collection begins.

  • Baseline agreed up front

    The evaluation task and the current baseline are fixed before capture starts, so the result is measurable on your terms.

  • Contributor-disjoint splits

    Releases ship with suggested train and evaluation splits that keep contributors and sites separate, so evaluation does not measure memorised scenes.

  • Targeted follow-up collection

    Where evaluation exposes a gap, the next round targets it: a specific subtask, a specific failure mode, a specific environment.

Designed for GDPR-conscious physical-data collection. Rights, privacy, and provenance are built into the data pipeline.

This describes how the pipeline is designed and operated. It is not a claim of certification, and no dataset is presented as fully anonymous.

Next step

Start with a scoped pilot

Bring a task. On the call we scope the capture protocol, the episode volume, the annotation schema and the delivery format.

Pilots typically start from approximately €25,000. This is an indicative figure only—final scope and pricing are set after a technical qualification call.