应用聚类技术分类提取Web页面.docVIP

  • 0
  • 0
  • 约5.15千字
  • 约 8页
  • 2018-04-06 发布于北京
  • 举报
应用聚类技术分类提取Web页面   摘要:针对Web中数据密集型的动态页面,文本数据少,网页结构化程度高的特点,介绍了一种基于HTML结构的web信息提取方法。该方法先将去噪处理后的Web页面进行解析,然后根据树编辑距离计算页面之间的相似度,对页面进行聚类,再对每一类簇生成相应的提取规则,对Web页面进行数据提取。   关键词:Web信息提取;树编辑距离;聚类;提取规则   中图分类号:TP391文献标识码:A文章编号:1009-3044(2010)01-212-02   Application of Clustering Technology Category Extraction Web Pages   CUI Hui-chao, LIU Li   (Southwest Jiaotong University, College of Information Science and Technology, Chengdu 610031, China)   Abstract: According to the characteristic of data-intensive dynamic web pages, insufficient text data and page structure with a high degree in web, this paper

文档评论(0)

1亿VIP精品文档

相关文档