Published January 1, 2025 | Version v1
Conference paper Open

PoseViTNet: Multi-Scene Absolute Pose Regression Using Vision Transformers

  • 1. Sabanci Univ, Fac Engn & Nat Sci, TR-34956 Istanbul, Turkiye

Description

Accurate camera pose estimation is crucial for autonomous driving and vehicle networking. Traditional pipelines based on geometric models and feature matching struggle in dynamic, featureless environments which are common in many environments. Inspired by the success of vision transformers (ViT), our approach uses a ViT backbone with an attention-based mask to extract a global image descriptor, which is then passed through fully connected layers for pose regression. The multi-headed self-attention in ViT helps the model learn scene layouts and focus on relevant features. We introduce an attention mask to improve performance in challenging scenes, especially dynamic or featureless ones. We compare three backbones: ViT (multi-headed self-attention throughout), ConViT (self-attention in the last two layers, gated positional self-attention elsewhere), and ResNet (pure convolution). We evaluate our model on two commonly used benchmarks for outdoor and indoor localization and we show that our model which uses ViT backbone achieves the state of the art results for both indoor and outdoor multi-scene absolute localization benchmarks.

Files

bib-994531e3-af49-4008-9e23-4aa72ea976ef.txt

Files (150 Bytes)

Name Size Download all
md5:7a1ba02cef0630978f5cd95b516f9b1a
150 Bytes Preview Download