Do VPD's Explanations Aggregate? An Audit of the Released Decomposition
TL;DR: adVersarial Parameter Decomposition (VPD) decomposes model weights into simple components and then labels each component as "needed here" or "safe to remove here" for each token of the input. These labels make up the "explanation" of that input. A core aspiration of VPD is that inputs' explanations can aggregate...