Bench
Layer 1 measured on real processors with injected faults: the shrinking quorum against fixed TMR, on the same faults, over 300 runs per host.
Every number on this page comes from scripts/bench.sh in mruspace/flight, run on the hosts listed under Hosts. The raw results, one JSON and one CSV file per host, are in bench/results. This is lab evidence on general-purpose processors. It is not radiation testing and not flight data. See Limits.
Method
Each run starts the quorum program: three replica processes and a voter. Every decision comes from the quorum crate, the same no_std code meant for flight. Each tick, the replicas hash a 4 KiB block of working memory. An upset flips one bit in that block during the computation, so the fault travels through the real calculation.
The bench runs every combination of:
- Policy: fixed TMR and the shrinking quorum, on the same seed and faults.
- Seed: 1 to 10. The seed sets where and when upsets land.
- Upset rate: 0.0001, 0.001 and 0.01, the chance per replica per tick of a bit flip.
- Fault schedule: five schedules, below.
Each run is 20,000 ticks. That makes 300 runs per host.
| Schedule | Faults (tick) | What it tests |
|---|---|---|
| none | Upsets only | Voting masks random upsets. |
| staggered-kills | Replica 0 dies at 5,000, replica 1 at 10,000 | Loss of two replicas, one at a time. |
| common-mode | Replica 0 dies at 2,000, replica 1 at 4,000 | Self-check on the last replica for 80% of the run. An upset can hit both of its runs alike (share 0.05). |
| stuck | Replica 1 stuck at 4,000, replica 0 dies at 9,000 | A replica that returns the same wrong value, found by health scoring. |
| last-stuck | Replicas 0 and 1 die at 5,000 and 10,000, replica 2 stuck at 15,000 | A stuck fault on the last replica, found only by the periodic known-answer test. |
Known-answer test on the last replica every 64 ticks. Common-mode share 0.05. Both are the program's defaults.
Metrics
Section 14 of the whitepaper names six metrics. A run of three processes supports three of them, plus two measures of its own.
| Metric | How the bench measures it |
|---|---|
| Mission Score | Correct results: ticks that delivered the right value. |
| Survival Duration | The tick of the last correct result. |
| Graceful Degradation Index | Correct results divided by capacity. Capacity is the most a checked design can deliver on the replicas still healthy at each tick: one result per tick with two or more, one every other tick with one (it must compute twice to check itself). 1.0 means every healthy replica did useful work. |
| Wrong results | Delivered values that were not correct. Nothing flagged them. |
| Detection latency | Ticks from a stuck, hang or corrupt fault to the voter's first reaction to that replica: an outvote, a caught disagreement, or its removal. |
Energy Efficiency, Decision Quality and Thermal Compliance need power, science and thermal models. The bench does not measure them.
Results
At an upset rate of 0.001, mean of 10 seeds. The ratio column gives the mean, and the lowest single seed in brackets.
| Schedule | Correct, TMR | Correct, shrink | Shrink / TMR | GDI, TMR | GDI, shrink |
|---|---|---|---|---|---|
| none | 19,999.9 | 19,999.9 | 1.00 (1.00) | 1.000 | 1.000 |
| staggered-kills | 9,988.5 | 14,985.0 | 1.50 (1.50) | 0.666 | 0.999 |
| common-mode | 3,994.7 | 11,987.3 | 3.00 (3.00) | 0.333 | 0.999 |
| stuck | 8,989.5 | 14,483.8 | 1.61 (1.61) | 0.620 | 0.999 |
| last-stuck | 9,988.5 | 12,486.2 | 1.25 (1.25) | 0.799 | 0.999 |
20,000 ticks per run. Results for all three upset rates are in the JSON files.
In all 150 pairs of runs, the shrinking quorum delivered at least as many correct results as fixed TMR on the same seed and faults. While two or more replicas live, both policies make the same decision, as the Kani proofs show. The gain comes after TMR stops: TMR halts at the second loss, or, in the stuck schedule, keeps comparing a good replica with a stuck one and rejects every result.
Wrong results
Fixed TMR delivered no wrong result in any run. It stops instead. The shrinking quorum pays for its extra work with a small number of wrong results on its last replica. The table gives the mean per run.
| Schedule | Upset rate 0.0001 | 0.001 | 0.01 |
|---|---|---|---|
| none | 0 | 0 | 0 |
| staggered-kills | 0 | 0.3 | 1.7 |
| common-mode | 0 | 0.3 | 3.6 |
| stuck | 0 | 0.2 | 1.7 |
| last-stuck | 20.0 | 20.1 | 20.7 |
Shrinking quorum, mean of 10 seeds per cell, 20,000 ticks per run.
There are two sources:
- Common-mode upsets in self-check. An upset that hits both runs alike gives two equal wrong values. The count follows the model: in the common-mode schedule at 0.01, about 8,000 self-checked results, 80 upsets and a share of 0.05 give 4.0 expected. The bench measured 3.6.
- A stuck last replica. Self-check cannot see a fault that gives the same wrong value twice. Only the known-answer test finds it. In last-stuck, the fault came at tick 15,000 and the test caught it at 15,040, after 20 wrong results. The run then stopped.
Survival and detection
| Schedule | Last correct result, TMR | Last correct result, shrink | Detection latency |
|---|---|---|---|
| staggered-kills | 9,999 | 20,000 | None to detect: the voter injects kills itself |
| common-mode | 3,999 | 20,000 | As above |
| stuck | 8,999 | 20,000 | 0 ticks |
| last-stuck | 9,999 | 14,998 | 40 ticks |
Upset rate 0.001, mean of 10 seeds. Runs end at tick 20,000. At 0.01, two shrinking-quorum runs caught a disagreement on the last tick, so their last correct result was tick 19,998.
With three replicas, a stuck replica is outvoted on the tick it fails, and health scoring retires it two ticks later. On the last replica, detection waits for the next known-answer test, at most 64 ticks. A shorter --kat-period gives fewer wrong results and less throughput.
Cost of the core
The bench crate times quorum::decide() and Health::strike() with std::time::Instant: the median of 9 rounds of 2,000,000 calls. Each figure includes the loop and the copy of the replies, so it is an upper bound.
| Host | decide(), per call | Health::strike(), per call |
|---|---|---|
| Apple M4 | 2.45 to 3.24 ns | 1.33 ns |
| Arm Neoverse-N2 | 2.51 to 4.41 ns | 3.40 ns |
| AMD EPYC 7763 | 3.79 to 7.63 ns | 6.15 ns |
The range covers all six input cases and both policies. Per-case figures are in the JSON files. On the shared CI machines, two runs of the same code differed by up to 10%.
Built for a bare-metal ARM Cortex-M target (thumbv7em-none-eabihf), the core is 902 bytes of code with rustc 1.99.0, and 1,178 bytes with rustc 1.86.0. In the runs, no process used more than 2.1 MB of memory.
Hosts
The same 300 runs on three hosts gave the same counts, run for run. Only the timing differs. In CI, the static ARM binaries also run a short bench under emulation (64-bit ARM, and 32-bit ARM with hard and soft float), and the build fails if any count differs from a native run.
| Processor | System | Compiler | Results file |
|---|---|---|---|
| Apple M4 | macOS, Darwin 27.0 | rustc 1.86.0 | 2026-10-11-darwin-arm64 |
| Arm Neoverse-N2 | Linux 6.17, GitHub Actions ubuntu-24.04-arm | rustc 1.99.0 | 2026-10-11-linux-aarch64 |
| AMD EPYC 7763 | Linux 6.17, GitHub Actions ubuntu-latest | rustc 1.99.0 | 2026-10-11-linux-x86_64 |
Release build: lto = true, codegen-units = 1. Each file records the commit it ran.
Limits
- Not radiation. Upsets are bit flips that the voter injects at chosen rates. The rates are test settings, not a model of any orbit.
- Not target hardware. The hosts are general-purpose processors. Flight-class boards come next. This is why the architecture status keeps Layer 1 at Prototype.
- Processes, not separate processors. The three replicas share one processor and one memory. A real failure can take more than one of them at once.
- Chosen faults. The five schedules set when replicas fail. Failure times in a real mission are random. For statistics over thousands of missions, see Dusk.
- Short runs. 20,000 ticks per run. The bench does not model aging, power or temperature.
Run it
# in a clone of mruspace/flight ./scripts/bench.sh # 300 runs, about 2 minutes on an Apple M4 ./scripts/bench.sh --quick # 2 seeds, 5,000 ticks cargo run --release -p bench # the cost of the core only
The script writes bench/results/<date>-<host>.json and .csv, with the host, processor, commit and build flags. The same seeds give the same counts on every run.
Model check
dusk --bench replays these schedules through Dusk's policy model with the same deaths, upset rate and self-check coverage, and compares the means. On each host's results, 17 of 18 rows agree within 3 standard errors on both correct and wrong results. The 18th, the shrinking quorum's correct results in the common-mode schedule at 0.01, is 3.4 standard errors low. Its two phases, rerun over 200 seeds, matched the model: 40.38 results lost against 40.42 expected with two replicas, and 19.17 caught and 0.90 wrong against 19.00 and 1.00 with one.
This checks the policy model against the code. It does not calibrate the environment: the bench sets the upset rate and the failure times itself. The stuck and last-stuck schedules are not modeled, because Dusk has only processors that die, not ones that return wrong values.
What comes next
- Flight-class boards. The same bench on target-class processors, with reboots and power cuts. The scripts and procedure are in docs/board.md. No board run has been done yet.
- Calibrate Dusk's environment with real upset and failure data, from a radiation beam test or from flight.
Updated 11 Oct 2026 · Source: scripts/bench.sh · Edit on GitHub · Questions? Get in touch

