PoseViTNet: Multi-Scene Absolute Pose Regression Using Vision Transformers
Oluşturanlar
- 1. Sabanci Univ, Fac Engn & Nat Sci, TR-34956 Istanbul, Turkiye
Açıklama
Accurate camera pose estimation is crucial for autonomous driving and vehicle networking. Traditional pipelines based on geometric models and feature matching struggle in dynamic, featureless environments which are common in many environments. Inspired by the success of vision transformers (ViT), our approach uses a ViT backbone with an attention-based mask to extract a global image descriptor, which is then passed through fully connected layers for pose regression. The multi-headed self-attention in ViT helps the model learn scene layouts and focus on relevant features. We introduce an attention mask to improve performance in challenging scenes, especially dynamic or featureless ones. We compare three backbones: ViT (multi-headed self-attention throughout), ConViT (self-attention in the last two layers, gated positional self-attention elsewhere), and ResNet (pure convolution). We evaluate our model on two commonly used benchmarks for outdoor and indoor localization and we show that our model which uses ViT backbone achieves the state of the art results for both indoor and outdoor multi-scene absolute localization benchmarks.
Dosyalar
bib-994531e3-af49-4008-9e23-4aa72ea976ef.txt
Dosyalar
(150 Bytes)
| Ad | Boyut | Hepisini indir |
|---|---|---|
|
md5:7a1ba02cef0630978f5cd95b516f9b1a
|
150 Bytes | Ön İzleme İndir |