Published January 1, 2018 | Version v1
Conference paper Open

Key Extraction in Table Form Documents: Insurance Policy as an Example

  • 1. SIBERTEK Inc, Insurance & Mobile App Unit, Ankara, Turkey
  • 2. Ankara Yildirim Beyazit Univ, Dept Comp Engn, Ankara, Turkey

Description

Automatic keyword/key-phrase extraction is an popular research area in text mining, and information retrieval which provides us a brief summary of documents and enables us to analyze large amount of textual material efficiently. Key-phrase extraction methodologies are generally based on the assumption that key-phrases are statistically important phrases and have a good coverage of documents. In this study, we propose a method to extract key-phrases in electronic documents such as insurance policy, bank receipts or e-invoices that appear as keys and have corresponding values. Unlike key-phrases, keys do not cover the whole document and their term frequencies are also relatively low, most of which appear only once in a document. They tend to appear in tables and forms at the beginning of the documents. Based on these observations, we propose a classification method for key extraction for table form electronic documents. Four machine learning based classification algorithms are applied and compared including decision trees, random forests, logistic regression and extreme gradient boosting. Experimental results show that random forests have perlOrmed well with an overall accuracy over 98%.

Files

bib-96e76e02-fbf6-419c-8bd8-bac767818f4a.txt

Files (197 Bytes)

Name Size Download all
md5:62cfd59d9a4202eb9d86b551abf09ef2
197 Bytes Preview Download