Skip to content
Read the original: John Schulman· Published 34/100AI score34/100

Schulman praises metric and dataset for training models to explain behavior

Original titleBullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seem...

AISummary

John Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline.

He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen.

That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

Read the original x.com

Source: John Schulman · x.comPublished · added here