Hase & Potts turn this into a training objective, making a model's CoT legible to a monitor. Karvonen et al. use the tested output of counterfactuals to build an interpretability eval.
Counterfactuals don't explain the mechanism, but predictability is a good base to build on.
