Reference specification: Human Repair & Tool-Use Episodes
A proposed episode format for real-world repair and tool-use media enriched with the initial state, objective, action sequence, mistakes, interventions, evidence, and immediate and delayed outcomes that a model cannot safely recover from pixels alone.
Overview
Intended uses
- Multimodal model evaluation
- Robotic task planning research
- Failure diagnosis
- Procedural reasoning research
Task coverage target
Technical shape
Episode fields
- Initial state and objective
- Environment, constraints, tools, and materials
- Ordered actions and human decisions
- Mistakes, diagnosis, and corrective interventions
- Immediate and delayed outcomes
- Human explanation, assertions, and evidence anchors
- Provenance, rights profile, and sensitive-data flags
Formats
- JSONL manifest
- Parquet metadata
- JSON-LD context
- Original or approved derivative media
Delivery targets
- Buyer-owned object storage
- Read-only managed access
- Checksummed pilot package
Planned machine interoperability: Schema.org Dataset JSON-LD on the canonical page, Croissant metadata in each ML-ready revision, and a versioned API contract for procurement agents.
Quality evidence
No metric is reported until a candidate revision is measured. A target explains the gate; it is not presented as an achieved result.
| Metric | Measured value | Qualification target | Why it matters |
|---|---|---|---|
| Qualified episodes | Not measured | Set by a funded buyer specification | Count only episodes that pass the published technical, context, and rights gates. |
| Human verification | Not measured | Reported by field and reviewer tier | Machine suggestions and contributor-verified facts must remain distinguishable. |
| Failure and correction coverage | Not measured | Balanced to the evaluation purpose | Success-only media is insufficient for many reasoning and robotics tasks. |
| Delayed outcome coverage | Not measured | Window defined before collection | Follow-up intervals must be disclosed; missing follow-up cannot be treated as success. |
| Commercial and model-use clearance | Not measured | 100% for the rights actually licensed | Each included episode must pass the exact final purpose and use-scope checks. |
| Duplicate and split leakage | Not measured | Measured before finalization | Exact and near duplicates, contributors, and linked episodes must be partitioned safely. |
Rights, risk, and limitations
Prohibited in this reference profile
- Identifying or re-identifying a person
- Biometric recognition or surveillance
- Employment, credit, insurance, or eligibility decisions
- Raw-data redistribution or public display
- Any use not written into the final license
Known limitations
- Inventory, distributions, and quality measurements do not exist until a candidate revision is assembled.
- Contributor descriptions and delayed outcomes may be self-reported unless independently verified.
- Consumer devices and home environments can introduce substantial capture variation.
- A rights declaration is not the same as independent ownership or subject-consent evidence.
- The reference specification excludes minors and high-risk sensitive categories at launch.
Commercial status
This reference specification is not inventory.
Its status does not change when any platform feature gate changes. It cannot be sampled, licensed, purchased, entitled, or delivered.
- Price
- Not established
- Evaluation sample
- Not available
- License
- No offer or entitlement