融合热红外与RGB相机的露天矿道路障碍物检测方法研究

Research on Open−pit Mine Road Obstacle Detection Method Based on Thermal Infrared and RGB Camera Fusion

  • 摘要: 针对露天矿场夜间光线暗弱及高浓度粉尘等恶劣低能见度工况下,单模态目标检测算法在道路障碍物检测中容易导致图像质量下降、检测准确率欠佳以及模型鲁棒性薄弱的问题,提出了一种基于迭代式GPT−Fusion的露天矿区可见光与热红外融合道路障碍物检测算法。借助深度网络驱动的局部特征匹配算法,实现两种图像间的特征配准;构建跨模态双分支特征融合网络,在特征融合阶段中引入GPT−Fusion模块,将二维视觉特征重组为一维词元序列以完成全局上下文关系的构建及深层特征互交;同时,网络内嵌了小目标感受野增强结构,通过多分支空洞卷积机制显著强化了算法对远距离细小障碍物的探测水平,最后在检测预测头中集成了空间上下文金字塔模块,自适应合并空间与通道维度信息来有效消解模态差异。研究结果表明,改进后模型在自建露天矿区双模态数据集中具有较好的检测性能,其平均精度mAP50达到89.4%,平均召回率达到88.7%,整体效能优于当前主流单模态和融合网络,可满足无人矿卡在复杂多变的露天矿区环境下进行安全的障碍检测要求。

     

    Abstract: Aiming at the drawbacks of single−modal object detection algorithms for road obstacle detection under harsh low−visibility operating conditions in open−pit mines—including dim nighttime illumination and high−concentration dust, which degrade image quality, reduce detection accuracy, and weaken model robustness—this paper proposes a visible−light and thermal infrared fused road obstacle detection algorithm for open−pit mines based on iterative GPT−Fusion. A deep network−driven local feature matching algorithm is adopted to realize feature registration between the two types of images. A cross−modal two−branch feature fusion network is constructed, where the GPT−Fusion module is embedded in the feature fusion stage. This module reorganizes two−dimensional visual features into one−dimensional token sequences to establish global contextual dependencies and conduct deep cross−modal feature interaction. Meanwhile, a small−object receptive field enhancement structure is embedded into the network, which significantly boosts the detection capability for distant tiny obstacles via a multi−branch dilated convolution mechanism. Finally, a spatial context pyramid module is integrated into the detection prediction head to adaptively aggregate spatial and channel−wise information, thereby effectively mitigating cross−modal discrepancies. Experimental results demonstrate that the improved model achieves superior detection performance on a self−built bimodal open−pit mine dataset, with a mean average precision mAP50 of 89.4% and an average recall of 88.7%. The overall performance outperforms state−of−the−art single−modal and fusion networks, satisfying the safe obstacle detection requirements of unmanned mining trucks in complex and variable open−pit mine environments.

     

/

返回文章
返回