Skip to main content

ASER: Attribute-Serialized Emotion Recognition With Generative Multi-Task Learning

By
Yingming Gao; Yiming Gao; Ruixi Jiang; Guoping Du; Jingwei Zhao; Yuhua Wen; Junyang Wu; Qifei Li; Ya Li

Speech emotion recognition (SER) is commonly formulated as direct utterance-to-label classification, leaving the acoustic and semantic evidence behind each prediction implicit. This letter proposes Attribute-Serialized Emotion Recognition (ASER), an attribute-serialized generative multi-task framework for SER, which organizes transcription, speaker-relative prosodic attributes (PROS), speaker gender, dimensional affect, textual sentiment, and emotion into a unified autoregressive target. Compact fixed-vocabulary attribute tokens are generated before the emotion token, enabling intermediate affective evidence to be modeled within the same decoding trajectory while adding only a short non-transcription segment. To improve the reliability of serialized attribute learning, the framework further incorporates speaker-relative prosodic coding, input-adaptive decoder-layer fusion, and task-normalized auxiliary supervision. Experiments on IEMOCAP and MELD show that ASER outperforms most recent audio-only baselines and approaches strong multimodal references.

Read on IEEE Xplore