Published January 1, 2010
| Version v1
Conference paper
Open
SYNTACTIC AND SUB-LEXICAL FEATURES FOR TURKISH DISCRIMINATIVE LANGUAGE MODELS
- 1. Bogazici Univ, Dept Elect & Elect Engn, Istanbul, Turkey
- 2. OHSU, OGI, Ctr Spoken Language Understanding, Beaverton, OR USA
Description
This paper investigates syntactic and sub-lexical features in Turkish discriminative language models (DLMs). DLM is a feature-based language modeling approach. It reranks the ASR output with discriminatively trained feature parameters. Syntactic information is incorporated into DLM as part-of-speech (PoS) tag n-gram features and head-to-head dependency relations. Sub-lexical units are first utilized as language modeling units in the baseline recognizer. Then, sub-lexical features are used to rerank the sub-lexical hypotheses. We explore features, similar to syntactic features, on sub-lexical units to reveal the implicit morpho-syntactic information conveyed by these units. We find out that DLM yields more improvement for sub-lexical units than for words. Basic sub-lexical n-gram features result in 0.6% reduction over the baseline and morpho-syntactic features yield an additional 0.4% reduction on the test set.
Files
bib-cc5b22e0-d8d0-419d-bf6c-7e19c046b934.txt
Files
(215 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:069e09d7c6a787bac0365431248458ea
|
215 Bytes | Preview Download |