DECEPTIcON: Bridging Gaps in In-the-Wild Deception Research
- 1. Bilkent Univ, Dept Comp Engn, TR-06800 Ankara, Turkiye
Description
We present DECEPTIcON, a new large-scale dataset for automatic deception detection. It contains video clips from 100 public figures, mostly politicians, along with manually aligned text transcripts and extracted audio-visual features. Each video is labeled with one of six truth levels from the PolitiFact fact-checking platform, allowing both fine-grained and binary classification tasks. Unlike earlier datasets, DECEPTIcON is designed to study deception in the wild, meaning it includes real-life, unscripted speech from a wide range of people and topics. We test and compare several baseline models for text, audio, and visual input separately, using state of the art pretrained architectures such as MPNet (text), Wav2Vec2 (audio), and VideoMAE (vision). Each model is trained for deception classification using 5-fold subject-independent cross-validation. We report CCR, F1-score, and MAE to evaluate performance. Our results show that text performs best overall, while fusion of multiple inputs leads to small but meaningful improvements. We also analyze the effect of different truth-level grouping strategies and show how attention-based interpretability tools help explain which parts of the input influence model predictions. DECEPTIcON aims to support fair, generalizable, and reproducible research in multimodal deception detection, and the dataset will be made available for research purposes.
Files
bib-aaefc1ce-1e9f-4171-bdc6-5a06aea2be02.txt
Files
(177 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:bc3414b5343a818b4f38e495e2c8d599
|
177 Bytes | Preview Download |