Adapter-Based Parameter-Efficient Fine-Tuning for Vision-Language Models in Radiology
Creators
- 1. Bilkent Univ, Elect & Elect Engn, Ankara, Turkiye
Description
Since the emergence of Vision-Language Models (VLMs) in natural image analysis, researchers have been working to extend their success to medical imaging, particularly for the joint processing of radiological images and reports. However, training a VLM from scratch is both compute-intensive and time-consuming, and it typically yields suboptimal performance compared to models in the natural image domain. Inspired by the recent successes of Parameter-Efficient Fine-Tuning (PEFT) methods, we propose a novel approach for fine-tuning a pre-trained VLM instead of training all parameters or starting from scratch. Our proposed approach is a fine-tuning strategy that efficiently adapts the relationship between images and reports using linear adapter layers added to the end of the model. Experiments based on the BiomedCLIP model demonstrate that our approach outperforms many state-of-the-art VLMs in radiology while significantly reducing training time and the number of trainable parameters.
Files
bib-30ca845c-42b2-43d6-81b5-5203d094616a.txt
Files
(236 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:e9b8e84b28c1b19ece2dd556371125fd
|
236 Bytes | Preview Download |