Data models and semantics
Agree on identity, state, timestamps and units before mapping data between systems. Record which system owns each field and how conflicting values are handled.
Sources and scopeSource record 25 August 2026
Technical source record: 25 August 2026. Check the linked documentation for current product requirements.
Technology neutral guidance; protocol specific schemas remain authoritative.
On this page
Overview#
Most cross system defects aren't byte level problems. They are disagreements about identity, state, time, units, or authority hidden behind syntactically valid XML or JSON.
Model the source before normalising it#
Preserve three representations where feasible:
- Raw source: immutable message or a privacy safe hash/reference plus transport metadata.
- Parsed source model: fields retaining the publisher's names and types.
- Canonical model: intentionally mapped entities and meanings used by consumers.
This makes lossy mappings visible and allows reprocessing when the canonical schema evolves. Never discard unknown fields merely because the current consumer ignores them; retain them within bounded, privacy controlled storage when policy permits.
Stable identity#
Don't use a display name, IP address, array index, or mutable location as the primary identity. A useful identifier record can include:
source_system: pacs-a
source_entity_type: door
source_entity_id: "8c18..."
canonical_entity_id: site-a/building-1/door-004
observed_name: East Lobby
mapping_version: 3
Distinguish person, identity record, credential, token instance, reader, access point, door, lock, input, controller, camera, stream, recording track, zone, and site. Their relationships change independently.
Types and units#
- Define integer width/sign and overflow behaviour for binary protocols.
- Treat identifiers that happen to contain digits as strings unless arithmetic is intended.
- Record unit, scale, precision, and valid range with numeric telemetry.
- Preserve
null, absent, empty, zero, false, unknown, and unsupported as distinct where the source does. - Define enum behaviour for new/unknown values; don't route an unknown alarm priority into the lowest class.
- Identify text encoding and normalisation; cap lengths before logging or rendering.
JSON's grammar is standardised by RFC 8259, but it doesn't define application semantics, duplicate name handling across all implementations, numeric precision, or a schema. XML likewise needs the applicable namespace/schema and application rules; the W3C publishes the XML specifications.
Provenance envelope#
For an observation, retain:
- source system and entity;
- source event/message identifier and sequence where available;
- source, receive, normalise, and publish times separately;
- source clock quality or uncertainty if known;
- protocol/profile/schema and product version;
- mapping version;
- authenticated peer/workload identity;
- trace/correlation identifiers;
- raw evidence reference and integrity metadata;
- privacy/security classification.
Canonical schemas#
Keep canonical fields small and stable; put source specific detail under a namespaced extension. Version the schema and mapping independently. Compatibility rules should state:
- whether consumers ignore unknown fields;
- which fields may become required;
- default versus absent behaviour;
- enum extension handling;
- timestamp and identifier format;
- retention and redaction requirements;
- how mapping corrections are replayed.
Semantic mapping record#
Every non trivial mapping should document source value, target value, conditions, information loss, authority, and fallback. For example, don't map all of forced-open, held-open, contact-open, and unlock-commanded to door_open without preserving the distinction.
Validation boundary#
Validate structure before allocating large objects; validate semantics before changing state; validate authorisation immediately before actuation. A schema valid command can still refer to the wrong site, stale entity, unsupported unit, or unauthorised operation.