Skip to the main content

Checking buyer capability gates…

← Data-product catalog
Reference specification — not availablepreview-spec-0.1

Reference specification: Human Repair & Tool-Use Episodes

A proposed episode format for real-world repair and tool-use media enriched with the initial state, objective, action sequence, mistakes, interventions, evidence, and immediate and delayed outcomes that a model cannot safely recover from pixels alone.

Overview

Intended uses

  • Multimodal model evaluation
  • Robotic task planning research
  • Failure diagnosis
  • Procedural reasoning research

Task coverage target

Household repairWoodworkingAssemblyMechanical troubleshootingRestoration

Technical shape

Episode fields

  • Initial state and objective
  • Environment, constraints, tools, and materials
  • Ordered actions and human decisions
  • Mistakes, diagnosis, and corrective interventions
  • Immediate and delayed outcomes
  • Human explanation, assertions, and evidence anchors
  • Provenance, rights profile, and sensitive-data flags

Formats

  • JSONL manifest
  • Parquet metadata
  • JSON-LD context
  • Original or approved derivative media

Delivery targets

  • Buyer-owned object storage
  • Read-only managed access
  • Checksummed pilot package

Planned machine interoperability: Schema.org Dataset JSON-LD on the canonical page, Croissant metadata in each ML-ready revision, and a versioned API contract for procurement agents.

Quality evidence

No metric is reported until a candidate revision is measured. A target explains the gate; it is not presented as an achieved result.

Reference quality metrics and current values
MetricMeasured valueQualification targetWhy it matters
Qualified episodesNot measuredSet by a funded buyer specificationCount only episodes that pass the published technical, context, and rights gates.
Human verificationNot measuredReported by field and reviewer tierMachine suggestions and contributor-verified facts must remain distinguishable.
Failure and correction coverageNot measuredBalanced to the evaluation purposeSuccess-only media is insufficient for many reasoning and robotics tasks.
Delayed outcome coverageNot measuredWindow defined before collectionFollow-up intervals must be disclosed; missing follow-up cannot be treated as success.
Commercial and model-use clearanceNot measured100% for the rights actually licensedEach included episode must pass the exact final purpose and use-scope checks.
Duplicate and split leakageNot measuredMeasured before finalizationExact and near duplicates, contributors, and linked episodes must be partitioned safely.

Rights, risk, and limitations

Prohibited in this reference profile

  • Identifying or re-identifying a person
  • Biometric recognition or surveillance
  • Employment, credit, insurance, or eligibility decisions
  • Raw-data redistribution or public display
  • Any use not written into the final license

Known limitations

  • Inventory, distributions, and quality measurements do not exist until a candidate revision is assembled.
  • Contributor descriptions and delayed outcomes may be self-reported unless independently verified.
  • Consumer devices and home environments can introduce substantial capture variation.
  • A rights declaration is not the same as independent ownership or subject-consent evidence.
  • The reference specification excludes minors and high-risk sensitive categories at launch.

Commercial status

This reference specification is not inventory.

Its status does not change when any platform feature gate changes. It cannot be sampled, licensed, purchased, entitled, or delivered.

Price
Not established
Evaluation sample
Not available
License
No offer or entitlement