Skip to main content

Emotion-Complexity Guided Expert Activation for Speech Emotion Recognition

By
Bao Thang Ta; Huynh Thi Thanh Binh; Van Hai Do

Mixture of Experts (MoE) models are a promising approach for Speech Emotion Recognition (SER), but they often rely on a fixed top-$k$ routing rule, which limits their ability to adapt expert activation to input complexity and training dynamics. We introduce Emotion-Complexity Guided Expert Activation (EC-MoE), a simple yet effective framework that replaces fixed routing with a per-utterance, entropy-derived activation threshold. A normalized Shannon entropy score, computed directly from the gating distribution, serves as an unsupervised measure of emotional complexity, enabling the dynamic activation of more experts for ambiguous utterances and fewer for prototypical ones. A temperature-sharpening schedule further adapts routing breadth during training, promoting early exploration and later specialization, while a warm-up phase stabilizes optimization before entropy-based filtering begins. Experiments on the IEMOCAP and Vietnamese ViSEC datasets demonstrate that EC-MoE consistently outperforms fixed top-$k$ baselines and recent competitive SER methods, achieving statistically significant improvements.

Read on IEEE Xplore