The explanation for why DPO suppresses VEA may be incomplete.
In a very simple setting where we assume that the win and lose datapoints in DPO are independent Bernoulli random variables and we learn a single Bernoulli parameter , DPO does not push toward the chosen marginal probability .
The global minimum of the DPO loss is in fact:
where is the DPO hyperparameter and is a reference.
So DPO reinforces if and only if (this is the condition for the second term above being positive).
Of course, this is an extreme caricature of what is going on here (they are... (read more)
The explanation for why DPO suppresses VEA may be incomplete.
In a very simple setting where we assume that the win and lose datapoints in DPO are independent Bernoulli random variables and we learn a single Bernoulli parameter , DPO does not push toward the chosen marginal probability .
The global minimum of the DPO loss is in fact:
where is the DPO hyperparameter and is a reference.
So DPO reinforces if and only if (this is the condition for the second term above being positive).
Of course, this is an extreme caricature of what is going on here (they are... (read more)