DPO implicit-reward logit mathematical notation
DPO implicit-reward logit is a recurring research-paper notation family. Scores a response by its learned-policy log probability relative to the reference policy, scaled by beta before pairwise comparison.
DPO implicit-reward logit: Scores a response by its learned-policy log probability relative to the reference policy, scaled by beta before pairwise comparison. Example: Reward responses whose relative probability rose above the reference.