Preference Learning: Aligning AI to Human Preferences
Overview
Preference-based reinforcement learning aligns AI systems to human judgment by showing annotators two outputs and asking which is better. The annotation protocols, agreement statistics, and training objectives supporting this approach were designed and validated on text.
Much of what this lab builds produces speech rather than text, where timing, prosody, and tone carry meaning alongside the words. While preference learning is well established for language models, its transfer to audio is largely untested: a PRISMA-guided review of roughly 500 papers found that only 6% apply it to audio tasks.
Modality also changes the judgment itself. In a controlled cross-modal study of human and synthetic annotation, identical content produced different preference ratings depending on whether raters read it or heard it. Agreement statistics imported from the text literature can therefore measure something other than what they report.
This work supports the lab’s speech and conversation projects, where feedback on how a clinician sounded, an AI persona that adapts its affect, and a simulated patient whose tone shifts all rest on defensible judgments that one generated utterance is better than another.
Collaborations
This work is carried out with collaborators across UC San Diego, including Prithviraj Ammanabrolu in Computer Science and Engineering, and with Eshin Jolly.