首页> 外文会议>International conference on web information systems engineering >Near Duplicate Text Detection Using Frequency-Biased Signatures
【24h】

Near Duplicate Text Detection Using Frequency-Biased Signatures

机译:使用频率偏置签名近重复文本检测

获取原文

摘要

As the use of electronic documents are becoming more popular, people want to find documents completely or partially duplicate. In this paper, we propose a near duplicate text detection framework using signatures to save space and query time. We also propose a novel signature selection algorithm which uses collection frequency of q-grams. We compare our algorithm with Winnowing, which is one of the state-of-the-art signature selection algorithms. We show that our algorithm acquires much better accuracy with less time and space cost. We perform extensive experiments to verify our conclusion.
机译:随着电子文件的使用变得越来越受欢迎,人们希望完全或部分复制文件。在本文中,我们提出了使用签名的近副本文本检测框架来节省空间和查询时间。我们还提出了一种新颖的签名选择算法,它使用Q-Grams的集合频率。我们将算法与WinNowing进行比较,这是最先进的签名选择算法之一。我们表明我们的算法通过更少的时间和空间成本获取更好的准确性。我们进行广泛的实验以验证我们的结论。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号