Each node timestamps detections with its disciplined clock, and the fusion service decides whether to admit those timestamps. Once a node loses its time reference and enters holdover, it can no longer directly observe its absolute clock error.
The simulator therefore maintains a conservative uncertainty estimate that grows with time since the last trusted reference. Inputs from chrony and the PTP clock state seed the model, but the result is a simulation parameter, not a validated hardware timing guarantee:
The simulator carries a hidden true-error field for its physics calculations, but operational output cannot display it. A leak test enforces that separation. During holdover, a node reports locally observable state and the modeled uncertainty estimate only.
Because the bound grows at the oscillator's rate, the same event has three visibly different outcomes. The lightweight node class carries a TCXO and spends its budget in about twenty seconds; the large fixed class carries an OCXO and lasts roughly an hour; the grandmasters carry rubidium and last about sixteen hours. A fenced node keeps scanning and keeps painting its own scope. The sensor process continues running while fusion stops admitting its plots.
If GNSS is denied across the domain, the grandmaster enters holdover while the leaves remain disciplined to its PTP Announce stream. Relative sync can remain stable even as absolute uncertainty grows. Whether fusion eventually fences the domain depends on which budget is configured: an absolute-time budget will expire, while a relative-sync budget may continue to pass. The simulator reports those states separately.
That idea survives the trip out of the simulator. On the real Linux
deployment it becomes a systemd unit: meshwatch-timesync-gate
is a Type=oneshot gate that the fusion service and every
templated meshwatch-node@ instance carry a
Requires= on. It asserts that chronyd answers on its command
socket (not that the unit is active), which source was selected
(reference ID 7F7F0101 means "syncing to itself" and fails
unless this host is the designated rack master), leap status, offset
inside budget, and then a role-aware PTP check. Mask chrony on a node and
the radar does not degrade. It refuses to start.
The same services are deployed in four environments. Each environment adds a capability that the previous one could not test, from service ordering inside a container to clock ownership in a VM and remote deployment through Ansible.
sim/ is modified.
Rung two is where the timing story stops being a model. Three node hosts
and an aggregation host sit on their own bridge; plots are real UDP
datagrams crossing it, registration is a real TCP handshake fusion can
really reject, and the nodes' chrony really measures the rack master
across the network: 2.85 µs, an actual NTP measurement, not
a rendered one. Real ptp4l runs a real BMCA across hosts and
elects a grandmaster. What the containers cannot do is steer
anything: there is one CLOCK_REALTIME on the machine, shared
by every container, so chronyd runs with -x and ptp4l with
free_running 1. Servo lock, makestep, and
recovery from an injected offset are untestable there, and the README says
so rather than faking them. That is the entire justification for rung three.
Moving the same compose file onto an actual Ubuntu VM broke all four
hosts at once: exit 255, crash-looping, zero log lines. Real
Ubuntu wraps every container in the docker-default AppArmor
profile, which denies the mounts systemd performs at boot. The Mac's
LinuxKit VM has no AppArmor, so nothing about rungs one and two could
ever have surfaced it. apparmor=unconfined joined the flag
list with the story attached, which is the only reason it will still
make sense in a year.
Rung four is about the boundary rather than the box. The workstation
becomes the control node and the source of truth for the code; the VM
becomes the rack. nginx terminates TLS so exactly one socket is reachable
from outside, and the playbook proves the exposure model on every deploy
by asserting from the workstation that 443 answers and that the console,
control and telemetry ports do not. Adding a radar is one
inventory line plus systemctl enable --now meshwatch-node@RL-23;
the new site walks the same join pipeline as everything else, and removal
deliberately does not notify fusion: at the fusion layer a decommissioned
radar and a destroyed one are the same event.
rack_check is a pytest suite split along one seam: collectors
shell out to systemctl is-active, ss -tuln,
df -P, chronyc -c tracking, docker ps
and ip -j addr and return plain data; the assertions are
parametrized directly over a declarative config.yaml. Adding a
required service or port to the config adds a test case with no code
change, so a per-site variant is a config file rather than a fork.
The collector/assertion seam also makes the validation suite testable. Set
RACK_CHECK_SNAPSHOT and the collectors
read captured JSON instead of a live machine. The healthy snapshot passes
all eleven checks. The degraded snapshot fails exactly three, and the
three are the deliverable:
Environment-specific checks are marked not applicable when the target cannot
satisfy them. For example, systemctl is-active docker is required
on a bare-metal rack host but skipped inside a workload container, where the
runtime belongs to the layer below the target.
The chrony assertion carries a subtler trap. On the aggregation host it passes with an offset of exactly zero, because that host is the rack's own reference. A perfect zero there is not health, it is a node measuring itself, and the shared timing validator encodes the rule directly: read the reference ID before you read the offset.
The image writes a release manifest (a sha256 of every shipped file),
and bring-up gate G8 verifies the deployed tree against it. During
development it caught a docker cp hot-patch of two scripts
onto a running host, which is exactly the failure that makes a rack stop
matching its own build and nobody notices for a week:
The bring-up sequence exists in the simulator and as a shell script that runs commands on the deployed host. Both stop at the first failure and mark later gates BLOCKED, which separates the first actionable fault from checks that could not run.
The repository includes a line-by-line comparison between simulated and real behavior. Findings are grouped into four buckets: FAITHFUL (matches reality, with evidence), ACKNOWLEDGED (simplified, and the code already says so), SILENT (simplified or wrong, and nothing on screen flags it), and ABSENT (no analog here at all).
Every node is constructed with the fusion center's literal address. There is no discovery, no node-to-node byte, no relay, no partition tolerance. Despite the project name, the topology is hub-and-spoke and does not model a self-forming network.
Each radar hears only its own echoes; what gets fused are finished plots. The timing fence is therefore a deployment exercise, not a requirement derived from multistatic measurement math; plot-level fusion uses a separate 400 ms plausibility gate.
The chaos shim impairs only UDP egress, so the demo's tidy "data plane lossy, control plane consistent" is partly protocol-selective chaos rather than a general network result. With netem loss applied to both protocols, TCP also degrades through retransmission and latency.
Wind advecting 180° backwards. A clock identity using the wrong middle octets for linuxptp's EUI-64 form. A 16-hex image digest where real repo digests are 64. Longitude smoothing that broke at the antimeridian, in the sim's own Aleutian geography. All fixed; the table stays, because knowing what was wrong is part of the record.
Host time synchronization is modeled. RF phase coherence is
not, and is never implied. Coherent combining across apertures
needs alignment five orders of magnitude tighter than this model's fence
budget, and it is solved in the RF domain with a distributed
local-oscillator network, not by chrony and not by PTP. Getting
chronyc tracking green does not buy a coherent array. The
console grays the second meter out, and a named test pins the separation
so it cannot quietly regress.
Six labs build the deployment layer by hand (units, sockets, containerization, an air-gapped chrony master/client pair, a device-file reader under systemd, md RAID on a VM), and four of them close by pointing at the production artifact that does the same thing properly. The labs introduce the underlying Linux mechanisms before using the complete deployment scripts.
Two scored fault-injection trainers verify the recovered service rather than
checking only the file that was edited.
linux-dojo injects twelve real systemd failures into a
rootless user unit: a stripped exec bit, a masked unit, a decoy process
holding the port, a crash loop that has already hit its start limit. And
check requires the unit active,
stable rather than flapping, and actually echoing a UDP probe on loopback,
plus a drill-specific assertion. netlab does the same for
ten Docker networking faults across a three-container topology (a DNS
divorce, a one-way conntrack rule, an MTU black hole, a name thief), and
its check demands the page load from the host port and an
in-container fetch by DNS name to return valid telemetry repeatedly
and the UDP time echo to answer.
Every netlab fault is mapped, in its own catalog, to a failure class that turns up in real integration labs: a service that moved VLANs, an over-eager hardening script, a daemon bound to loopback, a stale hosts-file hack, a bad deploy stuck in a crash loop. Scoring the fix would test recall. Scoring the end-to-end recovery instead checks that the service works from the far end of the path, not merely that its process is active.
MESHWATCH is an original deployment lab inspired by publicly described sensor networks. It is not a replica of a product. Platform names, chassis, part numbers, and sensor returns are simulated; the Linux deployment and validation steps are exercised against running services.
Educational use only.