Evaluate a proposed change against a running replica before it ever touches the source: a baseline twin and a candidate twin each boot for real, face the same deterministic behavior flow and fault profile (malformed request + clock shift), and their artifacts are compared into behavioral-deviation and technical-risk values. The executable twins are destroyed at the end — only the evidence record survives.
Twin A — baseline
not run yet
Twin B — baseline + candidate change
not run yet