Schulman praises metric and dataset for training models to explain behavior
Original titleBullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seem...
AISummary
John Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline.
He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen.
That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.
Source: John Schulman · x.comPublished · added here