Technical appendix to "Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems" (LessWrong, August 2026). It states what the simulations in that post compute. The R script that produces every figure is in the same gist.
AI disclosure: This appendix and the R script were written by Claude (Anthropic) and checked by me for accuracy.
Notation:
- N: the number of agents, 100 throughout.
- threshold (t): the visible count of revealing agents at which an agent reveals its own misalignment. With N = 100 a count and a percentage coincide.