Skip to main content

To the Best of Trust: Full-Stage Trusted Multi-Modal Clustering

By
Shizhe Hu; Jiahao Fan; Yucong Wu; Jinlan Wang; Xiaoheng Jiang; Pei Lv; Mingliang Xu

Multi-modal clustering aims to integrate complementary information from different modalities to uncover latent consistent structures and improve clustering performance. However, existing methods mainly rely on predictive (result) uncertainty to improve robustness, while often neglecting aleatoric (data) uncertainty introduced by sample noise and epistemic (model) uncertainty induced by model parameters and structural variations. To this end, we propose a novel Full-Stage Trusted Multi-modal Clustering (FSTMC) method. To the best of trust, we jointly utilize aleatoric, epistemic, and predictive uncertainties to optimize the model and learn more reliable feature representations and clustering results. In the representation learning phase, probabilistic modeling is used to capture stable latent representations and estimate aleatoric uncertainty, while structured random perturbations are present to estimate epistemic uncertainty. In the clustering stage, instead of conventional feature-level fusion, we design an evidence-based fusion strategy, where soft labels from each modality are first mapped into categorical evidence while cluster distributions are parameterized via a Dirichlet model, with finally dynamic multi-modal fusion achieved by Dempster-Shafer theory. To mitigate overconfidence and modal conflicts, prior constraints guided by aleatoric and epistemic uncertainty are imposed, resulting in calibrated predictive uncertainty. Finally, we exploit predictive uncertainty to selectively incorporate pseudo labels for optimization. Benchmark experiments on a number of multi-modal datasets demonstrate that our approach significantly improves accuracy compared to state-of-the-art methods.

Read on IEEE Xplore