Nearby in the stack

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling · arXivDesk