Spatial-Temporal Coherence in Extreme Video Retargeting for Consumer Screening Devices
Creators
- 1. Bahcesehir Univ, Dept Comp Engn, TR-34349 Istanbul, Turkiye
- 2. Florida Gulf Coast Univ, Dept Comp & Software Engn, Ft Myers, FL 33965 USA
Description
The accessibility of diverse display devices and their aspect ratios has drawn much research attention to video retargeting. Non-consistent video retargeting can significantly affect a video's spatial and temporal quality, particularly in extreme retargeting cases. Since there are no perfectly annotated datasets for video retargeting, deep learning-based techniques are rarely utilized. This paper proposes a method that learns to retarget videos by detecting the salient areas and shifting them to the appropriate location. First, we segment the salient objects using a unified Transformer model. Using convolutional layers and a shifting strategy, we shift and warp objects to the appropriate size and location in the frame. We use 1D convolution to move the salient items in the scene. Additionally, we employ a frame interpolation technique to preserve temporal information. To train the network, we feed the retargeted frames to a variational auto-encoder network to map the retargeted frames back to the input frames. Furthermore, we design perceptual and wavelet-based loss functions to train our model. Thus, we train the network unsupervised. Extensive qualitative and quantitative experiments on the DAVIS dataset show the superiority of the proposed method over existing image and video-based methods.
Files
bib-b995773c-4c07-4c6e-9fd2-bb08152a0748.txt
Files
(177 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:d4b9282bb8cd2d33ea9c8f4ac33ed1a3
|
177 Bytes | Preview Download |