MruDocs

Bench

Layer 1 measured on real processors with injected faults: the shrinking quorum against fixed TMR, on the same faults, over 300 runs per host.

Every number on this page comes from scripts/bench.sh in mruspace/flight, run on the hosts listed under Hosts. The raw results, one JSON and one CSV file per host, are in bench/results. This is lab evidence on general-purpose processors. It is not radiation testing and not flight data. See Limits.

Method

Each run starts the quorum program: three replica processes and a voter. Every decision comes from the quorum crate, the same no_std code meant for flight. Each tick, the replicas hash a 4 KiB block of working memory. An upset flips one bit in that block during the computation, so the fault travels through the real calculation.

The bench runs every combination of:

  • Policy: fixed TMR and the shrinking quorum, on the same seed and faults.
  • Seed: 1 to 10. The seed sets where and when upsets land.
  • Upset rate: 0.0001, 0.001 and 0.01, the chance per replica per tick of a bit flip.
  • Fault schedule: five schedules, below.

Each run is 20,000 ticks. That makes 300 runs per host.

ScheduleFaults (tick)What it tests
noneUpsets onlyVoting masks random upsets.
staggered-killsReplica 0 dies at 5,000, replica 1 at 10,000Loss of two replicas, one at a time.
common-modeReplica 0 dies at 2,000, replica 1 at 4,000Self-check on the last replica for 80% of the run. An upset can hit both of its runs alike (share 0.05).
stuckReplica 1 stuck at 4,000, replica 0 dies at 9,000A replica that returns the same wrong value, found by health scoring.
last-stuckReplicas 0 and 1 die at 5,000 and 10,000, replica 2 stuck at 15,000A stuck fault on the last replica, found only by the periodic known-answer test.

Known-answer test on the last replica every 64 ticks. Common-mode share 0.05. Both are the program's defaults.

Metrics

Section 14 of the whitepaper names six metrics. A run of three processes supports three of them, plus two measures of its own.

MetricHow the bench measures it
Mission ScoreCorrect results: ticks that delivered the right value.
Survival DurationThe tick of the last correct result.
Graceful Degradation IndexCorrect results divided by capacity. Capacity is the most a checked design can deliver on the replicas still healthy at each tick: one result per tick with two or more, one every other tick with one (it must compute twice to check itself). 1.0 means every healthy replica did useful work.
Wrong resultsDelivered values that were not correct. Nothing flagged them.
Detection latencyTicks from a stuck, hang or corrupt fault to the voter's first reaction to that replica: an outvote, a caught disagreement, or its removal.

Energy Efficiency, Decision Quality and Thermal Compliance need power, science and thermal models. The bench does not measure them.

Results

At an upset rate of 0.001, mean of 10 seeds. The ratio column gives the mean, and the lowest single seed in brackets.

ScheduleCorrect, TMRCorrect, shrinkShrink / TMRGDI, TMRGDI, shrink
none19,999.919,999.91.00 (1.00)1.0001.000
staggered-kills9,988.514,985.01.50 (1.50)0.6660.999
common-mode3,994.711,987.33.00 (3.00)0.3330.999
stuck8,989.514,483.81.61 (1.61)0.6200.999
last-stuck9,988.512,486.21.25 (1.25)0.7990.999

20,000 ticks per run. Results for all three upset rates are in the JSON files.

In all 150 pairs of runs, the shrinking quorum delivered at least as many correct results as fixed TMR on the same seed and faults. While two or more replicas live, both policies make the same decision, as the Kani proofs show. The gain comes after TMR stops: TMR halts at the second loss, or, in the stuck schedule, keeps comparing a good replica with a stuck one and rejects every result.

Wrong results

Fixed TMR delivered no wrong result in any run. It stops instead. The shrinking quorum pays for its extra work with a small number of wrong results on its last replica. The table gives the mean per run.

ScheduleUpset rate 0.00010.0010.01
none000
staggered-kills00.31.7
common-mode00.33.6
stuck00.21.7
last-stuck20.020.120.7

Shrinking quorum, mean of 10 seeds per cell, 20,000 ticks per run.

There are two sources:

  • Common-mode upsets in self-check. An upset that hits both runs alike gives two equal wrong values. The count follows the model: in the common-mode schedule at 0.01, about 8,000 self-checked results, 80 upsets and a share of 0.05 give 4.0 expected. The bench measured 3.6.
  • A stuck last replica. Self-check cannot see a fault that gives the same wrong value twice. Only the known-answer test finds it. In last-stuck, the fault came at tick 15,000 and the test caught it at 15,040, after 20 wrong results. The run then stopped.

Survival and detection

ScheduleLast correct result, TMRLast correct result, shrinkDetection latency
staggered-kills9,99920,000None to detect: the voter injects kills itself
common-mode3,99920,000As above
stuck8,99920,0000 ticks
last-stuck9,99914,99840 ticks

Upset rate 0.001, mean of 10 seeds. Runs end at tick 20,000. At 0.01, two shrinking-quorum runs caught a disagreement on the last tick, so their last correct result was tick 19,998.

With three replicas, a stuck replica is outvoted on the tick it fails, and health scoring retires it two ticks later. On the last replica, detection waits for the next known-answer test, at most 64 ticks. A shorter --kat-period gives fewer wrong results and less throughput.

Cost of the core

The bench crate times quorum::decide() and Health::strike() with std::time::Instant: the median of 9 rounds of 2,000,000 calls. Each figure includes the loop and the copy of the replies, so it is an upper bound.

Hostdecide(), per callHealth::strike(), per call
Apple M42.45 to 3.24 ns1.33 ns
Arm Neoverse-N22.51 to 4.41 ns3.40 ns
AMD EPYC 77633.79 to 7.63 ns6.15 ns

The range covers all six input cases and both policies. Per-case figures are in the JSON files. On the shared CI machines, two runs of the same code differed by up to 10%.

Built for a bare-metal ARM Cortex-M target (thumbv7em-none-eabihf), the core is 902 bytes of code with rustc 1.99.0, and 1,178 bytes with rustc 1.86.0. In the runs, no process used more than 2.1 MB of memory.

Hosts

The same 300 runs on three hosts gave the same counts, run for run. Only the timing differs. In CI, the static ARM binaries also run a short bench under emulation (64-bit ARM, and 32-bit ARM with hard and soft float), and the build fails if any count differs from a native run.

ProcessorSystemCompilerResults file
Apple M4macOS, Darwin 27.0rustc 1.86.02026-10-11-darwin-arm64
Arm Neoverse-N2Linux 6.17, GitHub Actions ubuntu-24.04-armrustc 1.99.02026-10-11-linux-aarch64
AMD EPYC 7763Linux 6.17, GitHub Actions ubuntu-latestrustc 1.99.02026-10-11-linux-x86_64

Release build: lto = true, codegen-units = 1. Each file records the commit it ran.

Limits

  • Not radiation. Upsets are bit flips that the voter injects at chosen rates. The rates are test settings, not a model of any orbit.
  • Not target hardware. The hosts are general-purpose processors. Flight-class boards come next. This is why the architecture status keeps Layer 1 at Prototype.
  • Processes, not separate processors. The three replicas share one processor and one memory. A real failure can take more than one of them at once.
  • Chosen faults. The five schedules set when replicas fail. Failure times in a real mission are random. For statistics over thousands of missions, see Dusk.
  • Short runs. 20,000 ticks per run. The bench does not model aging, power or temperature.

Run it

# in a clone of mruspace/flight
./scripts/bench.sh           # 300 runs, about 2 minutes on an Apple M4
./scripts/bench.sh --quick   # 2 seeds, 5,000 ticks
cargo run --release -p bench # the cost of the core only

The script writes bench/results/<date>-<host>.json and .csv, with the host, processor, commit and build flags. The same seeds give the same counts on every run.

Model check

dusk --bench replays these schedules through Dusk's policy model with the same deaths, upset rate and self-check coverage, and compares the means. On each host's results, 17 of 18 rows agree within 3 standard errors on both correct and wrong results. The 18th, the shrinking quorum's correct results in the common-mode schedule at 0.01, is 3.4 standard errors low. Its two phases, rerun over 200 seeds, matched the model: 40.38 results lost against 40.42 expected with two replicas, and 19.17 caught and 0.90 wrong against 19.00 and 1.00 with one.

This checks the policy model against the code. It does not calibrate the environment: the bench sets the upset rate and the failure times itself. The stuck and last-stuck schedules are not modeled, because Dusk has only processors that die, not ones that return wrong values.

What comes next

  • Flight-class boards. The same bench on target-class processors, with reboots and power cuts. The scripts and procedure are in docs/board.md. No board run has been done yet.
  • Calibrate Dusk's environment with real upset and failure data, from a radiation beam test or from flight.

Updated 11 Oct 2026 · Source: scripts/bench.sh · Edit on GitHub · Questions? Get in touch