Published January 1, 2025 | Version v1
Conference paper Open

Adapter-Based Parameter-Efficient Fine-Tuning for Vision-Language Models in Radiology

  • 1. Bilkent Univ, Elect & Elect Engn, Ankara, Turkiye

Description

Since the emergence of Vision-Language Models (VLMs) in natural image analysis, researchers have been working to extend their success to medical imaging, particularly for the joint processing of radiological images and reports. However, training a VLM from scratch is both compute-intensive and time-consuming, and it typically yields suboptimal performance compared to models in the natural image domain. Inspired by the recent successes of Parameter-Efficient Fine-Tuning (PEFT) methods, we propose a novel approach for fine-tuning a pre-trained VLM instead of training all parameters or starting from scratch. Our proposed approach is a fine-tuning strategy that efficiently adapts the relationship between images and reports using linear adapter layers added to the end of the model. Experiments based on the BiomedCLIP model demonstrate that our approach outperforms many state-of-the-art VLMs in radiology while significantly reducing training time and the number of trainable parameters.

Files

bib-30ca845c-42b2-43d6-81b5-5203d094616a.txt

Files (236 Bytes)

Name Size Download all
md5:e9b8e84b28c1b19ece2dd556371125fd
236 Bytes Preview Download