首页> 外文会议>4th International Moratuwa Engineering Research Conference >Distinguishing Real Web Crawlers from Fakes: Googlebot Example
【24h】

Distinguishing Real Web Crawlers from Fakes: Googlebot Example

机译:区分真实的Web爬虫与假冒:Googlebot示例

获取原文
获取原文并翻译 | 示例

摘要

Web crawlers are programs or automated scripts that scan web pages methodically to create indexes. such as Google, Bing use crawlers in order to provide web surfers with relevant information. Today there are also many crawlers that impersonate well-known web crawlers. For example, it has been observed that Google's Googlebot crawler is impersonated to a high degree. This raises ethical and security concerns as they can potentially be used for malicious purposes. In this paper, we present an effective methodology to detect fake Googlebot crawlers by analyzing web access logs. We propose using Markov chain mohels to learn profiles of real and fake Googlebots based on their patterns of web resource access sequences. We have calculated log-odds ratios for a given set of crawler sessions and our results show that the higher the log-odds score, the higher the probability that a given sequence comes from the real Googlebot. Experimental results show, at a threshold log-odds score we can distinguish the real Googlebot from the fake.
机译:Web搜寻器是程序或自动脚本,可以有条不紊地扫描网页以创建索引。例如Google,Bing等Bing使用搜寻器,以便向网络冲浪者提供相关信息。如今,也有许多搜寻器模仿了著名的Web搜寻器。例如,已经观察到Google的Googlebot搜寻器被高度模仿。这引起了道德和安全问题,因为它们有可能被用于恶意目的。在本文中,我们提出了一种通过分析网络访问日志来检测伪造的Googlebot抓取工具的有效方法。我们建议使用Markov链式莫赫尔基于网络资源访问序列的模式来学习真实和虚假Googlebot的配置文件。我们已经计算出一组给定的搜寻器会话的对数比,结果表明,对数比值越高,给定序列来自真实Googlebot的概率就越高。实验结果表明,在对数奇数阈值下,我们可以区分真实的Googlebot和假冒的Googlebot。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号