Published January 1, 2011 | Version v1
Conference paper Open

Data Sampling and Dimensionality Reduction Approaches for Reranking ASR Outputs Using Discriminative Language Models

  • 1. Bogazici Univ, Dept Elect & Elect Engn, Istanbul, Turkey
  • 2. Bogazici Univ, Dept Comp Engn, Istanbul, Turkey

Description

This paper investigates various approaches to data sampling and dimensionality reduction for discriminative language models (DLM). Being a feature based language modeling approach, the aim of DLM is to rerank the ASR output with discriminatively trained feature parameters. Using a Turkish morphology based feature set, we examine the use of online Principal Component Analysis (PCA) as a dimensionality reduction method. We exploit ranking perceptron and ranking SVM as two alternative discriminative modeling techniques, and apply data sampling to improve their efficiency. We obtain a reduction in word error rate (WER) of 0.4%, significant at p < 0.001 over the baseline perceptron result.

Files

bib-7b76c4ab-5924-4d3c-92c7-6a817a2a461b.txt

Files (289 Bytes)

Name Size Download all
md5:5b0dde49d10452e03ce87aac96726289
289 Bytes Preview Download