高级检索

基于矩形自校准与跨域交互的红外与可见光图像融合

Infrared and visible image fusion based on rectangular self-calibration and cross-domain interaction

  • 摘要: 为了解决红外与可见光图像融合在低光照、烟雾等极端环境下纹理细节与全局上下文信息易丢失的问题,采用基于矩形自校准与跨域交互的图像融合方法。该方法中,矩形自校准模块通过水平和垂直池化捕获轴向全局上下文,结合广播加法与大核条带卷积生成自适应注意力图,强化可见光图像对前景目标的鲁棒锁定;热感图卷积模块引入条件位置编码与深度卷积,保留红外图像热辐射结构完整性并减少背景噪声;跨域交互模块整合局部与全局注意力机制,实现多模态特征深度互补。实验结果表明,在 MSRS 数据集上,该方法的信息熵(EN)、标准差(SD)、峰值信噪比(PSNR)和均方误差(MSE)分别比对比方法提升11.2%、11.8%、0.7%、13.2%;在 M3FD 数据集上,信息熵(EN)、相关系数(CC)、标准差(SD)和峰值信噪比(PSNR)分别提升 14.7%、13.3%、5.4%、4.8%,多项指标均优于主流方法。融合图像在细节保留与视觉质量上表现突出。该研究对目标检测、自动驾驶等高级视觉任务的性能提升具有实际意义。

     

    Abstract:
    Infrared and visible image fusion is vital for advanced visual tasks like object detection and autonomous driving, integrating complementary information from the two modalities. Visible images provide rich textures but are sensitive to illumination changes and occlusions, while infrared images highlight thermal targets regardless of lighting but suffer from low resolution and weak textures. Existing methods often fail to balance thermal structure preservation and texture retention, especially under extreme conditions (e.g., low light, smoke), leading to key information loss and poor visual quality. A novel framework is thus needed to enhance target focusing, preserve domain-specific features, and enable deep cross-modal interaction.
    A fusion framework called rectangular self-calibration and cross-domain interaction fusion (RSCFusion) was proposed, comprising three core modules: rectangular self-calibration module (RCM), thermal sense map convolution module (TSMC), and cross-domain interaction module (CIM) (Fig.1). For visible images, RCM captured axial global context via horizontal/vertical pooling, fused features through broadcast addition, and generated adaptive attention maps to focus on foreground targets. For infrared images, TSMC introduced conditional position encoding (CPE) and deep convolution to preserve thermal structures and reduce edge blurring. CIM integrated local spatial attention (LSA) and global interactive semantic attention (GISA) for deep cross-modal complementarity. A comprehensive loss function (content, gradient, semantic) guided network training. Experiments were conducted on MSRS and M3FD datasets, using six metrics for comparison with seven state-of-the-art methods.
    Qualitative results showed that RSCFusion generated naturally bright fused images with clear thermal targets and intact details on MSRS (Fig.5) and handled extreme conditions effectively on M3FD (Fig.6). Quantitative evaluations on 20 image pairs revealed outstanding performance: on MSRS, entropy (6.649), standard deviation (8.053), peak signal-to-noise ratio (PSNR=64.405), and mean squared error (0.026) improved by 11.2%, 11.8%, 0.7%, and 13.2% (Table 1). On M3FD, entropy (7.498), correlation coefficient (0.599), standard deviation (10.305), and PSNR (64.516) increased by 14.7%, 13.3%, 5.4%, and 4.8% (Table 2). Ablation experiments confirmed module synergy, with the full configuration achieving the highest entropy (7.310) and PSNR (63.558) (Table 3).
    RSCFusion effectively addresses existing limitations via self-calibration and cross-domain interaction, generating fused images with richer information and stronger robustness under extreme conditions. Future work will focus on lightweight design and optimizing cross-modal alignment to improve real-time performance and handle extreme occlusions.

     

/

返回文章
返回