Author preprint record

An Evidence-Preserving Evaluation Protocol for Camera-First Planar Digital Twins

A six-class measurement protocol that keeps planar mapping, latency, continuity, replay equality, overload behavior, and backend agreement as separate evidence-backed claims. The worked audit uses Metriplane as the open-source reference implementation.

This is a public author-preprint record, not an IEEE publication. The manuscript has not been peer reviewed and should not be cited as an accepted journal article.

Abstract

Six claims, six measurement boundaries.

Camera-first planar digital twins turn fiducial detections into metric planar state. Their evaluations are hard to compare because localization, timing, continuity, replay, overload, and backend results are often reported over different inputs and system boundaries. This paper separates those claims into six measurement classes. For each class, the protocol records the measurand, stimulus, sampling unit, analysis, required outputs, evidence chain, and a conformance decision.

The protocol was tested on an open-source implementation. The M1 campaign retained 153 hash-linked observations from three complete marker replacements at 17 assigned locations under three separately estimated homographies. Relative to the assigned coordinates, mean radial deviation was 4.677 mm, the fixed-cell bootstrap 95% interval was 4.466-4.890 mm, the 95th percentile was 7.458 mm, the maximum was 11.244 mm, and radial root-mean-square deviation was 4.993 mm. The pooled marker-placement-capture within-cell vector component was 2.421 mm.

These values describe the complete apparatus and placement procedure, not instrument-only accuracy or a complete uncertainty statement. Earlier artifacts support narrower statements about timing, continuity, replay, overload, and CPU/GPU agreement. The audit keeps those software results separate from physical localization and exposes unsupported generalizations.

M1  mapping
M2  latency
M3  continuity
M4  replay equality
M5  overload behavior
M6  backend agreement

non-substitution rule:
software equality != physical accuracy

Protocol structure

Each result must answer one defined question.

M1 - Mapping

Planar coordinate deviation against independently established reference values, with residuals and an uncertainty budget.

M2 - Latency

A unique camera frame traced through declared start and end boundaries, reported as direct end-to-end and per-stage distributions.

M3 - Continuity

Expected-marker presence in a timestamped frame stream, including coverage, gaps, maximum gap duration, and acquisition-side loss boundaries.

M4 - Replay equality

Complete recorded input replayed under a fixed clock, with missing and extra records separated from common-field comparisons.

M5 - Overload

An offered-load sweep around service capacity, including queue policy, accepted and dropped work, latency, retention, and recovery.

M6 - Backend agreement

Identical inputs processed by each backend with synchronization, explicit numerical tolerance, repeated timings, and an equivalence margin when claimed.

Worked audit

The numerical result is deliberately narrower than an accuracy claim.

The revised M1 campaign retained 153 hash-linked observations from 17 assigned locations, three remove-and-replace placements, and three separately estimated homographies. Relative to the frozen assigned coordinates, the mean radial deviation was 4.677 mm, the 95th percentile was 7.458 mm, the observed maximum was 11.244 mm, and radial RMSE was 4.993 mm.

These values describe one complete apparatus and placement procedure. The assigned coordinates were not independently surveyed with calibrated equipment, so the paper does not relabel them as metrologically traceable absolute accuracy or expanded uncertainty.

Evidence boundary

The audit keeps useful results even when full conformance fails.

No class reached full protocol conformance in the worked audit. M1 lacks independently established coordinate values and a complete uncertainty budget. The retained M2-M6 artifacts support narrower statements about run-specific stage timing, logged-row continuity, intersection-scoped replay equality, one bounded-queue overload run, and CPU/GPU agreement at stored precision. They do not substitute for physical localization accuracy.

Version and provenance boundary

The article's evidence remains tied to v0.1.3, not the newest package.

The evaluated release is Metriplane v0.1.3 at commit ef26b07c40d76c6dd476f80bc2ad2b4504b2ef40. The later v0.2.0 and v0.2.1 releases did not produce any acquisition, processing result, analysis, figure, or table reported in this preprint. The current package must therefore not be used as a substitute identifier for the article's experimental evidence.

v0.1.3: evaluated

Canonical source and derived evidence named by this paper. Use this release when discussing the protocol audit and reported campaign.

Open evaluated release

v0.2.0: later research artifact

A separate DOI-archived incident-evidence release. It is relevant to the SoftwareX manuscript, not to the measurements reported here.

Open later release

v0.2.1: current package

The PyPI packaging and delivery release. Use it for ordinary installation, but not as the experimental version for this preprint.

Open package details

Citation

Use this stable public record until a journal record exists.

The landing-page URL is intended to remain stable. The record will be updated with the final journal citation and DOI if the manuscript is published.

Parkkinen, M. (2026). An Evidence-Preserving Evaluation Protocol for Camera-First Planar Digital Twins. Author preprint record.
https://www.metriplane.com/research/evidence-preserving-evaluation-protocol/