- 3
- 0
- 约3.11千字
- 约 49页
- 2020-09-02 发布于福建
- 举报
浙江大学本科生《数据挖掘导论》课件
第2课数据预处理技术
徐从富,副教授
浙江大学人工智能研究所
内容提纲
a Why preprocess the data?
a Data cleaning
a Data integration and transformation
■ Data reduction
a Discretization and concept hierarch
generation
a Summary
I. Why Data Preprocessing
Data in the real world is dirty
D incomplete: lacking attribute values, lacking certain
attributes of interest or containing only aggregate data
g, occupation“”
L noisy: containing errors or outliers
e.g
ary-
D inconsistent: containing discrepancies in codes or names
eg,Age=“4
“03/07/1997”
g. Was rating“1,2,3”, now rating“A,B,C
e. g, discrepancy between dupl
原创力文档

文档评论(0)