QUBVIS: query based multi-modal summarization system using CLIP based transformer and vision language models
Oluşturanlar
- 1. Manisa Celal Bayar Univ, Manisa, Turkiye
- 2. Firat Univ, Elazig, Turkiye
Açıklama
In this study, a new approach is proposed for user-interactive summarization of online videos. In the proposed approach, video-to-video summarization is performed with a very high success rate using a multimodal transformer architecture (QUBVIS) that also takes activity queries from the user as input, and the resulting summary video is subjected to captioning using a Vision Language Model with a GPT-2 decoder. The developed models are integrated with a Flask API and presented in a way that online video platforms can easily integrate into their systems. In addition, a simple web interface using this API is developed to provide API communication with the user. The performance evaluations of both models of the proposed method show our superiority over similar studies in the literature.
Dosyalar
bib-97bce212-465d-44c3-89b9-3f815e914288.txt
Dosyalar
(175 Bytes)
| Ad | Boyut | Hepisini indir |
|---|---|---|
|
md5:14248dde2a3adabcb5e40a3091dca6cb
|
175 Bytes | Ön İzleme İndir |