教程大数据平台与编程spark.pptxVIP

  • 2
  • 0
  • 约1.46万字
  • 约 64页
  • 2021-08-30 发布于北京
  • 举报
Apache SparkKishore Pusukuri, Spring 2018HTTP://WWW.CS.CORNELL.EDU/COURSES/CS5412/2018SPRecapMapReduceFor easily writing applications to process vast amounts of data in-parallel on large clusters in a reliable, fault-tolerant mannerTakes care of scheduling tasks, monitoring them and re-executes the failed tasksHDFS MapReduce: Running on the same set of nodes ? compute nodes and storage nodes same (keeping data close to the computation) ? very high throughputYARN MapReduce: A single master resource manager, one slave node manager per node, and AppMaster per applicationToday’s TopicsMotivation

文档评论(0)

1亿VIP精品文档

相关文档