Published January 1, 2010 | Version v1
Conference paper Open

SYNTACTIC AND SUB-LEXICAL FEATURES FOR TURKISH DISCRIMINATIVE LANGUAGE MODELS

  • 1. Bogazici Univ, Dept Elect & Elect Engn, Istanbul, Turkey
  • 2. OHSU, OGI, Ctr Spoken Language Understanding, Beaverton, OR USA

Description

This paper investigates syntactic and sub-lexical features in Turkish discriminative language models (DLMs). DLM is a feature-based language modeling approach. It reranks the ASR output with discriminatively trained feature parameters. Syntactic information is incorporated into DLM as part-of-speech (PoS) tag n-gram features and head-to-head dependency relations. Sub-lexical units are first utilized as language modeling units in the baseline recognizer. Then, sub-lexical features are used to rerank the sub-lexical hypotheses. We explore features, similar to syntactic features, on sub-lexical units to reveal the implicit morpho-syntactic information conveyed by these units. We find out that DLM yields more improvement for sub-lexical units than for words. Basic sub-lexical n-gram features result in 0.6% reduction over the baseline and morpho-syntactic features yield an additional 0.4% reduction on the test set.

Files

bib-cc5b22e0-d8d0-419d-bf6c-7e19c046b934.txt

Files (215 Bytes)

Name Size Download all
md5:069e09d7c6a787bac0365431248458ea
215 Bytes Preview Download