Skip to main content

Sparsity-Controllable Normality Learning With Vision–Language Models for Scenario-Related Video Anomaly Detection

By
Jiangyun Chen; Yuanjie Dang; Peng Chen; Ronghua Liang; Haidong Gao; Nan Gao; Ruohong Huan; Dongdong Zhao; Xiang Tian; Shifeng Zhang

Video anomaly detection (VAD) is critical for automation systems and security surveillance. Recently, multimodal vision–language models (MLLMs) have attracted increasing attention due to their rich pre-trained knowledge and strong explainability. However, existing MLLM-based approaches struggle to adapt to real-world settings where anomaly definitions are complex: they either rely on the model’s built-in knowledge and mainly capture only generic anomalies, or require anomalous samples for supervised fine-tuning—which are often rare and may raise legal or privacy concerns. To address this challenge, we propose a Sparsity-Controllable Vision-Language Model (SCVLM) for scenario-related anomaly detection. SCVLM learns normality from unlabeled normal data by jointly reconstructing multimodal representations and summarizing textual descriptions, thus enabling anomalies to be detected as deviations from the learned normal patterns. We introduce a Sparsity-Controllable Memory Block (SCMB) to improve memory addressing mechanism for pretrained multimodal representations. During inference, anomalies are detected by fusing two modality-specific reconstruction-errors with an LLM-based textual anomaly scoring mechanism. Meanwhile, we fuse the interpretable cues from each detection branch to derive anomaly reasoning consistent with human commonsense. Extensive experiments on challenging benchmarks demonstrate that SCVLM achieves state-of-the-art detection performance. Our code is available at https://github.com/SCVLM/SCVLM

Read on IEEE Xplore