PoseViTNet: Multi-Scene Absolute Pose Regression Using Vision Transformers
Creators
- 1. Sabanci Univ, Fac Engn & Nat Sci, TR-34956 Istanbul, Turkiye
Description
Accurate camera pose estimation is crucial for autonomous driving and vehicle networking. Traditional pipelines based on geometric models and feature matching struggle in dynamic, featureless environments which are common in many environments. Inspired by the success of vision transformers (ViT), our approach uses a ViT backbone with an attention-based mask to extract a global image descriptor, which is then passed through fully connected layers for pose regression. The multi-headed self-attention in ViT helps the model learn scene layouts and focus on relevant features. We introduce an attention mask to improve performance in challenging scenes, especially dynamic or featureless ones. We compare three backbones: ViT (multi-headed self-attention throughout), ConViT (self-attention in the last two layers, gated positional self-attention elsewhere), and ResNet (pure convolution). We evaluate our model on two commonly used benchmarks for outdoor and indoor localization and we show that our model which uses ViT backbone achieves the state of the art results for both indoor and outdoor multi-scene absolute localization benchmarks.
Files
bib-994531e3-af49-4008-9e23-4aa72ea976ef.txt
Files
(150 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:7a1ba02cef0630978f5cd95b516f9b1a
|
150 Bytes | Preview Download |