| 摘 要: 针对现有视觉位置识别(VPR)方法多尺度表征不足、空间特征不稳定及全局聚合效果有限等问题,提出一种基于空间-频域多尺度语义增强的 VPR方法。利用金字塔视觉 Transformer(PVT)提取分层多尺度特征,以兼顾全局语义与局部细节;通过小波变换设计空间-频域联合增强机制,实现频域分解与门控调制,并与空间特征融合以增强鲁棒性;引入语义驱动的多尺度特征聚合模块,获得紧凑且判别性强的全局表示。在 Pitts30k、SPED、Pitts250k、MSLS_val、Nordland、Eynsham 和St-Lucia等7个公开街景影像数据集上的实验结果表明,该方法在Recall@1指标上分别取得了94.2%、92.5%、95.6%、92.0%、90.0%、92.1%和99.9%的性能,整体表现优于对比的主流方法,在复杂环境下有效提升了 VPR的鲁棒性与精度。 |
| 关键词: 视觉位置识别 金字塔视觉Transformer 小波变换 特征聚合 |
|
中图分类号:
文献标识码: A
|
| 基金项目: 国家自然科学基金项目资助(62031021) |
|
| Visual Place Recognition via Multi-Scale Semantic Enhancement with Spatial-Frequency Fusion |
|
WANG Weiwei, MA Lingkun
|
(School of Electronic Information and Artificial Intelligence, Shaanxi University of Science and Technology, Xi’an 710021, China)
231611014@sust.edu.cn; malingkun@sust.edu.cn
|
| Abstract: To address the limitations of existing Visual Place Recognition ( VPR) methods in multi-scale representation, spatial feature stability, and global aggregation, this paper proposes a VPR approach based on spatial-frequency domain multi-scale semantic enhancement. Specifically, a Pyramid Vision Transformer (PVT) is employed to extract hierarchical multi-scale features, balancing global semantic information and local details. A spatia-l frequency joint enhancement mechanism is designed via wavelet transform, which performs frequency decomposition and gated modulation, and then integrates frequency and spatial features to improve robustness. Furthermore, a semantic-driven multi-scale feature aggregation module is introduced to obtain compact and highly discriminative global representations. Experimental results on seven public street-view datasets—Pitts30k, SPED, Pitts250k, MSLS _ val, Nordland, Eynsham, and St-Lucia—demonstrate that the proposed method achieves Recall@1 scores of 94.2%,92.5%,95.6%,92.0%,90.0%,92.1%,and 99.9% , respectively. Overall, the proposed method outperforms mainstream competing approaches and significantly enhances robustness and accuracy of VPR in complex environments. |
| Keywords: visual place recognition Pyramid Vision Transformer wavelet transform feature aggregation |