Published January 1, 2018 | Version v1
Conference paper Open

Audio-Visual Prediction of Head-Nod and Turn-Taking Events in Dyadic Interactions

  • 1. Koc Univ, Istanbul, Turkey

Description

Head-nods and turn-taking both significantly contribute conversational dynamics in dyadic interactions. Timely prediction and use of these events is quite valuable for dialog management systems in human-robot interaction. In this study, we present an audio-visual prediction framework for the head-nod and turn taking events that can also be utilized in real-time systems. Prediction systems based on Support Vector Machines (SVM) and Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN) are trained on human-human conversational data. Unimodal and multi-modal classification performances of head-nod and turn-taking events are reported over the IEMOCAP dataset.

Files

bib-16344ddc-cfe6-4ed5-b238-8724106e530e.txt

Files (242 Bytes)

Name Size Download all
md5:ab6f4953912d5c9e09a2dc2c2a6138c9
242 Bytes Preview Download