preference-learning

Same Words, Different Judgments: How Preferences Vary Across Modalities

arXiv preprint 2026Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for speech.

Preference-Based Learning in Audio Applications: A Systematic Analysis

arXiv preprint 2025Despite the parallel challenges that audio and text domains face in evaluating generative model outputs, preference learning remains remarkably underexplored in audio applications. Through a PRISMA-guided systematic review of approximately 500 papers, we find that only 30 (6%) apply preference learning to audio tasks.