- 1、本文档共54页,可阅读全部内容。
- 2、原创力文档(book118)网站文档一经付费(服务费),不意味着购买了该文档的版权,仅供个人/单位学习、研究之用,不得用于商业用途,未经授权,严禁复制、发行、汇编、翻译或者网络传播等,侵权必究。
- 3、本站所有内容均由合作方或网友上传,本站不对文档的完整性、权威性及其观点立场正确性做任何保证或承诺!文档内容仅供研究参考,付费前请自行鉴别。如您付费,意味着您自己接受本站规则且自行承担风险,本站不退款、不进行额外附加服务;查看《如何避免下载的几个坑》。如果您已付费下载过本站文档,您可以点击 这里二次下载。
- 4、如文档侵犯商业秘密、侵犯著作权、侵犯人身权等,请点击“版权申诉”(推荐),也可以打举报电话:400-050-0827(电话支持时间:9:00-18:30)。
查看更多
中文分词应用研究现状.ppt
* Bakeoff 2007 – 法国电信北京研发中心 Problems of NER with only local information “Many empirical approaches…make decision only on local context for extract inference, which is based on the data independent assumption. But often this assumption does not hold because non-local dependencies are prevalent in natural language.” Observation from Experiments: There are many seen named entities are missed; At least 10% of unseen and missed named entities have been labeled out correctly for at least once. “If the context surrounding one occurrence of a token sequence is very indicative of it being an entity, then this should also influence the labeling of another occurrence of the same token sequence in a different context that is not indicative of entity”. * Bakeoff 2007 – 法国电信北京研发中心 * Bakeoff 2007 – 法国电信北京研发中心 Local Features Unigram:Cn(n=-2,-1,0,1,2) Bigram:CnCn+1(n=-2,-1,0,1) and C-1C1 0/1 Features Assign 1 to all the characters which are labeled as entity and 0 to all the characters which are labeled as NONE in training data. In such way, the class distribution can be alleviated greatly , taking Bakeoff 2006 MSRA NER training data for example, if we label the corpus with 10 classes, the class distribution is: 0.81(B-PER), 1.70(B-LOC), 0.95(BORG), 0.81(I-PER), 0.88(I-LOC), 2.87(I-ORG), 0.76(EPER), 1.42(E-LOC), 0.94(E-ORG), 88.86(NONE) if we change the label scheme to 2 labels(0/1), the class distribution is: 11.14 (entity), 88.86(NONE) * Bakeoff 2007 – 法国电信北京研发中心 Non-local Features Token-position features(NF1) These refer to the position information(start, middle and last) assigned to the token sequence which is matched with the entity list exactly. These features enable us to capture the dependencies between the identical candidate entities and their boundaries. Entity-majority features(NF2) These refer to the majority label assigned to the token sequence which is matched with the entity list exactly. These features enable us to capture the dependencies between the identical e
您可能关注的文档
- 中国古典建筑。.ppt.ppt
- 中国对外开放格局.ppt
- 中国建筑艺术赏析.ppt
- 中国式云计算服务模式.ppt
- 中国文化经典研读:儒道互补精讲课课件.ppt
- 中国文字话健康(莫志兵)中华讲师网.ppt
- 中国文学史第二讲北宋前期词坛.ppt
- 中国文学史第八讲姜夔、吴文英的宋末词坛.ppt
- 中国文学史第六讲南渡词人.ppt
- 中国新诗现代诗ppt_ppt.ppt
- 房地产企业产品创新策略规划与2025年目标客群画像定位分析.docx
- 2025年幼儿园保育员五级业务能力考试试题附解析.docx
- 航空物流市场需求动态变化对航空货运枢纽选址影响报告.docx
- 基于2025年抖音社交平台的短视频变现策略研究报告.docx
- 2025年幼儿园保育员业务考试试题(I卷)含答案.docx
- 2025-2026学年高中英语选择性必修第二册冀教版(2019)教学设计合集.docx
- 2025年无人机适航认证案例在安防监控领域的应用报告.docx
- 中小学教师心理健康教育能力提升规划与培训方案.docx
- 绿色环保产业扶持资金2025年申请政策红利与项目实施路径报告.docx
- 2025-2026学年高中英语选择性必修第二册上教版(2020)教学设计合集.docx
文档评论(0)