DPO implicit-reward logit mathematical notation

DPO implicit-reward logit is a recurring research-paper notation family. Scores a response by its learned-policy log probability relative to the reference policy, scaled by beta before pairwise comparison.

DPO implicit-reward logit: Scores a response by its learned-policy log probability relative to the reference policy, scaled by beta before pairwise comparison. Example: Reward responses whose relative probability rose above the reference.