The inclusion of more potentiallycorrect words in the candidate sets is important to improve the accuracy of Large Vocabulary Continuous Speech Recognition(LVCSR).A candidate expansion algorithm based on theWeighted Syllable Confusion Matrix(WSCM)is proposed.First,WSCM is derived from aconfusion network.Then,the recognised candidates in the confusion network is used to conjecture the most likely correct words based onWSCM,after which,the conjectured wordsare combined with the recognised candidatesto produce an expanded candidate set.Finally,a combined model having mutual informationand a trigram language model is used to rerank the candidates.The experiments on Mandarin film data show that an improvement of9.57% in the character correction rate is obtained over the initial recognition performanceon those light erroneous utterances.
展开▼