TinyRS-R1: Compact Vision Language Model for Remote Sensing
Creators
- 1. Middle East Tech Univ, Ctr Image Anal OGAM, Dept Elect & Elect Engn, TR-06800 Ankara, Turkiye
Description
Remote sensing (RS) applications often rely on edge hardware that cannot host the models in the 7B parametric vision language of today. This letter presents TinyRS, the first 2B-parameter vision language models (VLMs) optimized for RS, and TinyRS-R1, its reasoning-augmented variant. Based on Qwen2-VL-2B, TinyRS is trained via a four-stage pipeline: pretraining on million-scale satellite images, instruction tuning, fine-tuning with chain-of-thought (CoT) annotations from a new reasoning dataset, and group relative policy optimization (GRPO)-based alignment. TinyRS-R1 matches or surpasses recent 7B RS models in classification, visual question answering (VQA), grounding, and open-ended QA-while using one third of the memory and latency. CoT reasoning improves grounding and scene understanding, while TinyRS excels at concise, low-latency VQA. TinyRS-R1 is the first domain-specialized small VLM with GRPO-aligned CoT reasoning for general-purpose RS. The code, models, and caption datasets are available at https://github.com/aybora/TinyRS
Files
bib-7317cf84-cc8f-4ac3-ab0d-c4e0f9f733b5.txt
Files
(153 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:3270e40d4de26d34437ec8321b442ecf
|
153 Bytes | Preview Download |