ZL Corpus

A reusable data product for Physical AI.

Package synchronized observations, motion, actions, task context, provenance, and reviewed labels into versioned corpora built for model training and evaluation.

Data product

The delivery layer of the Zerolaw stack.

ZL Corpus turns processed and reviewed Physical AI episodes into a consistent product that model teams can reuse.

Each release keeps the data, schema, calibration context, provenance, task semantics, and quality state connected. Teams receive a corpus that can be filtered, versioned, split, and traced instead of a directory of unrelated recordings.

Inside a corpus release

The context required to use the data.

A ZL Corpus release carries the relationships that would otherwise have to be reconstructed downstream.

Episodes

Synchronized experience

Observations, motion, actions, task boundaries, and outcomes aligned on a shared timeline.

Schema

Consistent interfaces

Explicit signal definitions, coordinate systems, timestamps, and dataset structure.

Provenance

Traceable origin

Capture device, calibration, environment, processing version, and review history.

Quality

Reviewed status

Validation results, confidence, known exceptions, and human-reviewed task semantics.

Product architecture

Built through the complete Zerolaw pipeline.

  1. 01
    ZL Capture

    Records task-relevant vision, inertial, depth, pose, and custom multimodal signals.

  2. 02
    ZL Core

    Cleans, slices, synchronizes, calibrates, and fuses captured streams.

  3. 03
    ZL Annotation

    Adds reviewed task semantics and feeds quality findings back into platform rules.

  4. 04
    ZL Corpus

    Organizes accepted episodes, schemas, metadata, and splits as a versioned data product.

Delivery contract

A corpus release is explicit about what it contains.

LayerExample contentsWhy it matters
Episode dataVision, IMU, pose, depth, actions, task statesPreserves the observation-action sequence
SchemaSignal definitions, timestamps, coordinate framesKeeps training interfaces consistent
MetadataEnvironment, task, device, subject, outcomeSupports filtering and coverage analysis
QualityValidation status, confidence, exceptions, reviewSupports acceptance and sampling decisions
VersioningRelease identifier, processing history, split definitionMakes experiments reproducible

Model development

Designed around repeatable training decisions.

Corpus exploration

Filter episodes by task, environment, signal availability, outcome, and quality state.

Training splits

Keep explicit train, validation, and evaluation definitions tied to a corpus version.

Failure analysis

Trace model behavior back to capture context, calibration, labels, and processing history.

Dataset evolution

Add coverage and corrections without losing the provenance of earlier experiments.

ZL Corpus FAQ

Product questions.

What is ZL Corpus?

ZL Corpus is Zerolaw's Physical AI data product. It packages synchronized episodes, schemas, provenance, reviewed labels, quality information, and dataset splits into a reusable release.

How is it different from raw recordings?

Raw recordings are sensor files. ZL Corpus preserves the relationships between observations, motion, actions, task context, calibration, labels, and outcomes in a consistent versioned structure.

How does it connect to the platform?

ZL Capture records the required signals, ZL Core processes and synchronizes them, ZL Annotation adds reviewed semantics and quality feedback, and ZL Corpus organizes the resulting episodes for model development.

Build with Zerolaw

Define the corpus around your model and task.

Discuss data requirements