Video Reward Model
As I’m currently working on a project related to RL with video generation models, my dear boss asked me to study reward model training practices, and translating a pairwise model to a ranking model during inference. So here’s my study notes.
