RGB-Thermal semantic segmentation is essential for robust visual understanding in autonomous driving, surveillance, and search-and-rescue applications. While RGB and thermal modalities provide complementary information, cross-modal difficult regions with ambiguous boundaries and conflicting signals represent the primary performance bottleneck. Existing methods treat all regions uniformly and lack principled approaches for targeting these critical challenging areas. To address this limitation, we propose FocusSeg, a prompt-guided refinement framework embodying “Focus on the Hard” to systematically identify and refine difficult cross-modal regions. Our approach employs a two-stage framework: first establishing multi-modal understanding, then introducing intelligent point-based correction. Specifically, the framework integrates a Difficult Point Sampling Strategy (DPSS) for dynamic challenging region identification, a Multimodal Correction Module (MMCM) for training specialized correction heads, and a Point Correction Strategy (PCS) combining correction-head predictions with historical priors to generate reliable corrective prompts. These refined prompts guide a class-aware SAM framework toward superior targeted refinement. Extensive experiments on multiple RGB-T datasets demonstrate substantial improvements over state-of-the-art methods, with particularly notable gains in challenging scenarios involving cross-modal conflicts and boundary ambiguities. Our work establishes a new paradigm for intelligent resource allocation in multimodal segmentation tasks. The source code will be released at https://github.com/Wanda36/FocusSeg
