Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

Huu-anh Tran; Yuhang Guo; Ping Jian; Shumin Shi; Heyan Huang

首页> 中文期刊> 《北京理工大学学报：英文版》 >Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

开具论文收录证明 >>

期刊封面封底目录下载 >>

文献代查 >>

页面导航

摘要
著录项
相似文献
相关主题

摘要

The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource. However,getting a parallel corpus,which has a large scale and is of high quality,is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately,multilingual user generated contents( UGC),such as bilingual movie subtitles,provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable,the original corpus is not suitable for statistical machine translation( SMT) systems. The corpus may contain translation errors,sentence mismatching,free translations,etc. To improve the quality of the bilingual corpus for SMT systems,three filtering methods are proposed: sentence length difference,the semantic of sentence pairs,and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus.Experimental results demonstrate that all the three methods effectively improve the corpus quality,and the machine translation performance( BLEU score) can be improved by 1. 32.

著录项

来源
《北京理工大学学报：英文版》 |2018年第1期|127-136|共10页
作者
Huu-anh Tran; Yuhang Guo; Ping Jian; Shumin Shi; Heyan Huang;
展开▼
作者单位

Department of Computer Science and Technology;

Beijing Institute of Technology;

Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Application;

Beijing Institute of Technology;

展开▼
原文格式 PDF
正文语种 chi
中图分类文字信息处理;
关键词
parallel corpus filtering; low resource languages; bilingual movie subtitles; machine translation; Chinese-Vietnamese translation;

相似文献

中文文献
外文文献
专利

1. Corpus Augmentation for Improving Neural Machine Translation [J] . Zijian Li ,Chengying Chi ,Yunyun Zhan . 计算机、材料和连续体(英文) . 2020,第7期
2. The Virtual Corpus Approach to DerivingNgram Statistics from Large Scale Corpora [C] . . 1998中文信息处理国际会议 . 1998
3. A Comparative Study of Rule-based and Statistical-based Machine Translation Models [A] . 王茜 . 2011

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

摘要

著录项

相似文献

相关主题

期刊订阅