NeurIPS 2026 - Evaluations and Datasets TrackPreference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for speech.