Published January 1, 2025 | Version v1
Journal article Open

TinyRS-R1: Compact Vision Language Model for Remote Sensing

  • 1. Middle East Tech Univ, Ctr Image Anal OGAM, Dept Elect & Elect Engn, TR-06800 Ankara, Turkiye

Description

Remote sensing (RS) applications often rely on edge hardware that cannot host the models in the 7B parametric vision language of today. This letter presents TinyRS, the first 2B-parameter vision language models (VLMs) optimized for RS, and TinyRS-R1, its reasoning-augmented variant. Based on Qwen2-VL-2B, TinyRS is trained via a four-stage pipeline: pretraining on million-scale satellite images, instruction tuning, fine-tuning with chain-of-thought (CoT) annotations from a new reasoning dataset, and group relative policy optimization (GRPO)-based alignment. TinyRS-R1 matches or surpasses recent 7B RS models in classification, visual question answering (VQA), grounding, and open-ended QA-while using one third of the memory and latency. CoT reasoning improves grounding and scene understanding, while TinyRS excels at concise, low-latency VQA. TinyRS-R1 is the first domain-specialized small VLM with GRPO-aligned CoT reasoning for general-purpose RS. The code, models, and caption datasets are available at https://github.com/aybora/TinyRS

Files

bib-7317cf84-cc8f-4ac3-ab0d-c4e0f9f733b5.txt

Files (153 Bytes)

Name Size Download all
md5:3270e40d4de26d34437ec8321b442ecf
153 Bytes Preview Download