№ 002 · Results
Same hardware, 1.34× the useful work
Dusk simulates 10,000 missions on processors that fail and are never repaired. A shrinking quorum delivered a third more useful work than fixed triple redundancy, and never less on any single mission.
The question
Classic triple modular redundancy has a fixed threshold. Three processors vote. Two can still detect a disagreement. One cannot outvote anything, so the system stops.
On a mission with no repairs and no ground team in reach, that throws away the last processor's whole remaining life. The whitepaper proposes a shrinking quorum instead: vote while three are alive, compare while two are, self-check on one. Falling back from TMR to simplex is not new. The question is what it is worth when no ground team can command it.
The model
Each mission is a day-by-day fault-injection run. Processors die on a Weibull lifetime. Deaths can be correlated, as a shared thermal or power fault would be. Upsets corrupt a day's result at the rate that gets past EDAC and scrubbing.
Every policy runs against the same fault history and the same upsets. Any difference between policies is the policy.
The result
Mean useful years per mission, over 10,000 missions:
That is 1.34× the useful work of fixed TMR (95% CI 1.33× to 1.35×). It is never less on any single mission, because the two policies are identical until TMR stops.
dusk --bin charts; nothing by hand.The cost
A lone self-checking processor misses a few of its upsets. Those become wrong results that nobody knew were bad. Fixed TMR never produces them, because it never lives long enough.
So the honest comparison is a price. Fixed TMR comes out ahead only if one wrong result costs more than 136 years of useful work (95% CI 131 to 142). Against one computer with spares, the shrinking quorum gives 1.6× the useful work and 59% fewer wrong results.
When the ground can help
Real spacecraft do not run TMR alone. When a string fails, a ground team commands a fallback by hand. Dusk adds that team as a baseline, for as long as support lasts.
| Ground support ends | Gain over TMR |
|---|---|
| At launch | 1.34× |
| After 25 years | 1.31× |
| After 50 years | 1.25× |
| After 100 years | 1.11× |
| Never | 1.00× |
Answer time barely matters. A fallback a day or six months late gives the same result to two decimals. What matters is whether anyone is still there.
How sure are we
The sweep varies every parameter the model cannot pin down. Across every row, the shrinking quorum returned 1.11× to 1.68× the useful work of fixed TMR. The gain is largest with early, random failures and smallest with sharp wear-out, where all three processors die close together.
What this does not claim
The parameters are plausible orders of magnitude, not values fitted to flight data. The claim is the shape of the trade, and the sweep shows where it holds. Power, thermal limits and duty cycling are not modelled. The ground team in the baseline is perfect, which flatters the ground.
Cite this
@software{binns2026dusk,
author = {Binns, Will},
title = {dusk: A Monte Carlo of Redundancy Policies
on Hardware That Only Ever Decays},
year = {2026},
url = {https://github.com/mruspace/dusk}
}Sources
- R. E. Lyons and W. Vanderkulk, “The Use of Triple-Modular Redundancy to Improve Computer Reliability,” IBM J. Res. Dev., 1962.
- W. Binns, “Mru: A Fault-Tolerant Operating System for Thousand-Year Autonomous Operation,” 2026. doi:10.5281/zenodo.20579438
- J. F. Meyer, “On Evaluating the Performability of Degradable Computing Systems,” IEEE Trans. Computers, 1980.
- B. Efron, “Bootstrap Methods: Another Look at the Jackknife,” Annals of Statistics, 1979.