Dusk
A Monte Carlo of redundancy policies on hardware that only ever decays. It asks what a shrinking quorum buys, and what it costs, on the same simulated hardware.
Run it
cargo run --release # the default scenario, about 2 s cargo run --release -- --spares 2 # add salvaged processors cargo run --release -- --sweep # sensitivity, about 30 s cargo run --release -- --ground # against TMR with a ground team, about 10 s cargo run --release --bin charts # redraw the charts and data.json
Times are for an 8-core laptop. Results depend only on the seed, never on the thread count.
Options
| Option | Default | Meaning |
|---|---|---|
| --runs | 10000 | Missions per policy |
| --seed | 1 | Base seed |
| --threads | all cores | Worker threads |
| --horizon-years | 1000 | Mission horizon |
| --nodes | 3 | Processors at launch |
| --spares | 0 | Salvaged processors |
| --shape | 1.5 | Weibull shape |
| --scale-years | 125 | Weibull scale |
| --p-corr | 0.1 | Chance a death kills another |
| --p-upset | 0.001 | Bad-result upsets per node-day |
| --coverage | 0.99 | Lone self-check coverage |
| --selfcheck-throughput | 0.5 | Lone output vs a vote |
The fault model
Each mission is a day-by-day fault-injection run.
- Processors die on a Weibull lifetime. Shape 1.5 and a 125-year scale give a median life of about 98 years.
- Deaths can be correlated. With probability
p_corr, a death also kills one other processor that day. - Spares are salvaged from instruments as they die, already aged by the same radiation.
- Upsets corrupt a day's result, at the rate that gets past EDAC and scrubbing.
- The comparison is paired. Every policy sees the same fault history, so any difference is the policy.
Default results
policy useful p10 p50 p90 service wrong days
mean y mean y mean / run
fixed TMR 99.0 37.6 92.7 168.9 99.0 0.000
standby simplex 83.1 38.4 78.3 134.0 166.4 0.603
shrinking quorum 132.6 64.3 126.6 208.3 166.4 0.2461.34× the useful work of fixed TMR in total (95% CI 1.33× to 1.35×). Never less on any single mission.
Sweeps and ground support
--sweep varies each parameter the model cannot pin down. Across every row, the gain over fixed TMR stays between 1.11× and 1.68×.
--ground adds a ground team that commands the fallback by hand while support lasts. With 49 years of support, Voyager's so far, the gain is still 1.25×. Answer time barely matters. What matters is whether anyone is still there.
What this does not claim
Updated 29 Sep 2026 · Edit on GitHub · Questions? Request information