| 摘 要: 随着城市既有建筑老化问题日益增多,传统人工巡检效率低下,而单一可见光视觉方法难以识别缺乏表面纹理特征的内部空鼓,单一红外热成像又受限于低分辨率导致的边界模糊。针对以上问题,本文提出了一种基于可见光与红外热成像双模态协同的CT-FuseNet(CNN-Transformer Fusion Network)融合模型缺陷智能检测方法,该方法在模型架构上,构建了异构双流并行架构:利用改进的ResNet-CNN分支提取可见光图像中的高频纹理与边缘细节,同时利用Swin Transformer分支捕捉热成像数据中的全局拓扑结构与长距离语义依赖。为解决异构特征融合中的语义鸿沟,本文设计了“语义-尺度双重对齐机制”与“边缘感知增强模块”,通过双向门控交互实现特征的深度互补,并引入Sobel梯度先验显式优化缺陷边界。此外,提出了包含分类、分割、严重度回归及边界距离预测的协同多任务学习框架,配合梯度归一化策略解决任务间的优化冲突。基于自主构建的12,869组时空严格对齐的双模态数据集验证,该方法在缺陷分类准确率上达到 91.1%,分割 mIoU 达到 88.1%,显著优于现有主流模型,能够为房屋室内缺陷的自动化检测提供有效的技术支持。 |
| 关键词: 房屋室内缺陷检测 CNN-Transformer融合 双模态融合 多任务学习 边缘感知增强模块 |
|
中图分类号:
文献标识码:
|
|
| Research on Indoor Building Defect Detection Method Based on Bimodal Data Synergy and CT-FuseNet Fusion Model |
|
sunjiahui
|
Anhui Construction Engineering Testing and Research Institute Co., Ltd.
|
| Abstract: As the aging problem of existing urban buildings grows increasingly severe, traditional manual inspection suffers from low efficiency. Moreover, single visible-light vision methods struggle to identify internal hollowing that lacks surface texture features, while single infrared thermography is limited by low resolution, leading to blurred boundaries. To address these issues, this paper proposes an intelligent defect detection method based on visible light and infrared thermal imaging bimodal synergy using the CT-FuseNet (CNN-Transformer Fusion Network) fusion model. In terms of model architecture, a heterogeneous dual-stream parallel architecture is constructed: an improved ResNet-CNN branch is used to extract high-frequency texture and edge details from visible light images, while a Swin Transformer branch captures global topological structures and long-range semantic dependencies from thermal imaging data. To bridge the semantic gap in heterogeneous feature fusion, this paper designs a "semantic-scale dual alignment mechanism" and an "edge-aware enhancement module" to achieve deep feature complementarity through bidirectional gated interaction, and introduces a Sobel gradient prior to explicitly optimize defect boundaries. Furthermore, a collaborative multi-task learning framework is proposed, including classification, segmentation, severity regression, and boundary distance prediction, together with a gradient normalization strategy to resolve optimization conflicts among tasks. Validated on a self-constructed bimodal dataset of 12,869 spatially and temporally strictly aligned image pairs, the proposed method achieves a defect classification accuracy of 91.1% and a segmentation mIoU of 88.1%, significantly outperforming existing mainstream models, and can provide effective technical support for automated indoor building defect detection. |
| Keywords: Indoor building defect detection CNN-Transformer fusion Bimodal fusion Multi-task learning Edge Perception Enhancement Module |