Research
PreferenceEKF Tracks Reward-Model Uncertainty With a Kalman Filter in a Low-Dimensional Subspace
RLHF is sample-inefficient, so active learning matters for choosing informative preference queries, but quantifying uncertainty over a large neural reward model is the blocker. PreferenceEKF frames active preference learning as sequential Bayesian filtering, running an extended Kalman filter within a low-dimensional parameter subspace instead of attempting posterior inference over the full network, and updating the reward-model posterior as each new preference query arrives. That makes sampling network parameters cheap enough to compute acquisition functions, with results demonstrated on the D4RL and V-D4RL benchmarks.
↳ Follow the thread