PLC-DPO Teaches Preference Optimization to Correct Noisy Labels
DPO treats pairwise preferences as reliable supervision, but real datasets often contain reversed, weak, or ambiguous judgments. PLC-DPO uses a calibrated policy–reference margin to route each pair toward reinforcement, reversal, or tie regularization.
Read more