Published January 1, 2025 | Version v1
Conference paper Open

VISUAL STATE-SPACE BASED MULTI-TASK LEARNING FOR BUILDING SEGMENTATION AND HEIGHT ESTIMATION

  • 1. Ozyegin Univ, Fac Engn, TR-34794 Istanbul, Turkiye
  • 2. Istanbul Medipol Univ, TR-34810 Istanbul, Turkiye
  • 3. Huawei Turkiye R&D Ctr, TR-34768 Istanbul, Turkiye

Description

Building segmentation and height estimation from remote sensing data are critical tasks in urban planning, telecommunications, and geospatial sciences. However, current methods often rely on multi-modal data, such as SAR imagery, or face challenges with computational inefficiency and boundary inaccuracies. In this work, we address these limitations by proposing BuildMamba, a novel dual-path multi-task learning framework that performs joint building segmentation and height estimation using only RGB satellite imagery. To our knowledge, this is the first model to utilize a pretrained VMamba encoder, trained on segmentation tasks, combined with a Mamba Attention Module (MAM) for enhanced feature representation and a boundary-aware Huber loss to improve accuracy at building edges. Our model integrates global and local context through a dual-path backbone, refined using MAM, and outputs precise segmentation masks and height maps. BuildMamba achieves state-of-the-art performance on the DFC23 dataset, with a segmentation mIoU of 0.922 and a height estimation delta(1) score of 0.872, outperforming prior approaches such as HGDNet and MFTSC. These findings demonstrate the efficiency and scalability of BuildMamba in addressing key challenges in remote sensing, offering a robust solution for height estimation and segmentation tasks.

Files

bib-dc94256d-d126-4690-b198-1d983e6fe167.txt

Files (231 Bytes)

Name Size Download all
md5:eded109c20531a1aed1a93a6acd9b7c0
231 Bytes Preview Download