RLAIF: The Model as Preference Labeller
RLAIF replaces human preference labeling with a model that chooses between two responses, keeping the downstream RLHF pipeline unchanged. The method, exemplified by Constitutional AI, offers scalability and consistency b…