When we ablate the mean-diff direction from the B vector of a single rank-1 adapter model at layer 24, misalignment
What if you use the B vector to ablate from the residual activations ?
(0.04 cosine similarity)
Couldn't it possible that 0.04 is actually high in high dimensional space ? One could compare it to samples of random directions.
At which point are those downstream changes measured ?