Published January 1, 2025 | Version v1
Conference paper Open

Financial Statement Fraud Detection with a Categorical-to-Numerical Data Representation

  • 1. Ozyegin Univ, Artificial Intelligence & Data Engn, Istanbul, Select State, Turkiye

Description

Identifying fraudulent financial reports and elucidating the mechanisms of fraud are critical for safeguarding investors from substantial losses. Financial statements present detailed accounting entries in tabular form; they inherently combine categorical and numerical variables governed by accounting dependencies, yet most existing methods fail to model interpretable interactions between these feature types. In this case, handling categorical variables together with numerical variables is important in enhancing the financial statement fraud detection performance. Here, we compare the methods for transforming categorical to numerical attributes, which are then used for financial statement fraud detection. We perform comprehensive experiments on two real-world datasets: FiGraph and USFSD. We compare 4 state-of-the-art specialized categorical-to-numerical transformation techniques with several other simpler statistical encoding mechanisms, such as target, label, Helmert, and GLMM encodings, as well as methods that can directly work on categorical data, such as CatBoost. These specialized transformation techniques are Hierarchical Coupling Learning-based CURE, Graph-based Categorical Embedding GCE, and Transitive Distance Learning-based embedding. The results reveal that the performance of CURE and XGBoost together surpasses all state-of-the-art techniques, achieving significant relative gains in macro-level recall over the second-best performing approaches, CatBoost and FT-Transformer, while also providing clear and interpretable insights into the discovered fraud pathways.

Files

bib-145f30d9-bc9f-4ecb-859b-52e6b3824dc4.txt

Files (182 Bytes)

Name Size Download all
md5:36c2ee4c712c8902a4248fa92e6a26be
182 Bytes Preview Download