7.7日 实验03-Spark批处理开发(2)

使用Spark处理数据文件

检查数据

检查$DATA_EXERCISE/activations里的数据,每个XML文件包含了客户在指定月份活跃的设备数据。

拷贝数据到HDFS的/dw目录

样本数据示例:

复制代码
<activations>
  <activation timestamp="1225499258" type="phone">
    <account-number>316</account-number>
    <device-id>d61b6971-33e1-42f0-bb15-aa2ae3cd8680</device-id>
    <phone-number>5108307062</phone-number>
    <model>iFruit 1</model>
  </activation>
  ...
</activations>

处理文件

读取XML文件并抽取账户号和设备型号,把结果保存到/dw/account-models,格式为account_number:model

输出示例:

复制代码
1234:iFruit 1
987:Sorrento F00L
4566:iFruit 1
...

提供了解析XML的函数如下:

复制代码
// Stub code to copy into Spark Shell

import scala.xml._

// Given a string containing XML, parse the string, and 
// return an iterator of activation XML records (Nodes) contained in the string

def getActivations(xmlstring: String): Iterator[Node] = {
    val nodes = XML.loadString(xmlstring) \\ "activation"
    nodes.toIterator
}

// Given an activation record (XML Node), return the model name
def getModel(activation: Node): String = {
   (activation \ "model").text
}

// Given an activation record (XML Node), return the account number
def getAccount(activation: Node): String = {
   (activation \ "account-number").text
}

上传数据

复制代码
# 1. 检查并创建HDFS目录
hdfs dfs -mkdir -p /dw

# 2. 将本地数据上传到HDFS(替换$DATA_EXERCISE为实际路径)
hdfs dfs -put $DATA_EXERCISE/activations /dw/

# 3. 检查文件是否上传成功
hdfs dfs -ls /dw/activations
复制代码
定义题目提供的解析函数
复制代码
def getActivations(xmlstring: String): Iterator[Node] = {
    (XML.loadString(xmlstring) \\ "activation").toIterator
}

def getModel(activation: Node): String = (activation \ "model").text
def getAccount(activation: Node): String = (activation \ "account-number").text
复制代码
读取数据(像处理日志一样)
复制代码
val xmlRDD = sc.wholeTextFiles("/dw/activations/*.xml")
复制代码
测试解析(查看第一条记录)
复制代码
val firstRecord = getActivations(xmlRDD.first()._2).next()
println(s"测试解析结果: ${getAccount(firstRecord)}:${getModel(firstRecord)}")
复制代码
处理全部数据
复制代码
val resultRDD = xmlRDD.flatMap { case (_, xml) => 
  getActivations(xml).map(act => s"${getAccount(act)}:${getModel(act)}")
}
复制代码
查看结果样例(10条)
复制代码
resultRDD.take(10).foreach(println)
复制代码
保存结果(先清理旧数据)
复制代码
import org.apache.hadoop.fs._
val outputPath = "/dw/account-models"
val fs = FileSystem.get(sc.hadoopConfiguration)
if (fs.exists(new Path(outputPath))) fs.delete(new Path(outputPath), true)

resultRDD.saveAsTextFile(outputPath)
println(s"结果已保存到 hdfs://$outputPath")

验证结果(在Linux终端执行)

复制代码
# 查看输出结果
hdfs dfs -cat /dw/account-models/part-* | head -n 10

# 如果需要合并结果到单个文件
hdfs dfs -getmerge /dw/account-models ./account_models.txt
head account_models.txt
相关推荐
杜大哥34 分钟前
python程序:如何查看电脑【电池电量的剩余百分比】 和 【是否插入连接着充电器】?
开发语言·python
wuyk55535 分钟前
Python网络爬虫入门到实战 第01章:爬虫到底是什么?原理、流程、合法性、风险全解析(零基础必看)
开发语言·爬虫·python
yurenpai(27届找实习中)42 分钟前
Java 10 的 var 为什么不能随便用?从订单汇总看懂类型推断与泛型陷阱
java·开发语言·java18
shmily麻瓜小菜鸡1 小时前
JS/TS 易踩坑知识点 — 模块化与工程化类
开发语言·javascript·ecmascript
geovindu1 小时前
CSharp: Observer Pattern
开发语言·后端·观察者模式·设计模式·c#·.netcore·行为模式
郝学胜-神的一滴2 小时前
C++11 工程级应用 09:告别无谓拷贝,解锁高性能移动语义
开发语言·数据结构·c++·vscode·软件工程·visual studio
MC皮蛋侠客3 小时前
OPC UA 系列(一):标准全景与 Python 最小闭环——让第一条设备数据流动起来
开发语言·python·opcua
默_笙3 小时前
🍕 AI 的嘴巴装了水管(上):从"等它说完"到"边说边听"的流式输出指南
前端·javascript
SamChan903 小时前
PDF翻译后的格式完整性校验:用Python自动比对译文与原文档的表格与段落结构
开发语言·python·ai·pdf·机器翻译
暖焰核心3 小时前
继承全解——继承、默认成员函数、切片、隐藏与虚继承
java·前端·javascript