详解 Spark 核心编程之 RDD 分区器

一、RDD 分区器简介

  • Spark 分区器的父类是 Partitioner 抽象类
  • 分区器直接决定了 RDD 中分区的个数、RDD 中每条数据经过 Shuffle 后进入哪个分区,进而决定了 Reduce 的个数
  • 只有 Key-Value 类型的 RDD 才有分区器,非 Key-Value 类型的 RDD 分区的值是 None
  • 每个 RDD 的分区索引的范围:0~(numPartitions - 1)

二、HashPartitioner

默认的分区器,对于给定的 key,计算其 hashCode 并除以分区个数取余获得数据所在的分区索引

scala 复制代码
class HashPartitioner(partitions: Int) extends Partitioner {
    require(partitions >= 0, s"Number of partitions ($partitions) cannot be negative.")
    
    def numPartitions: Int = partitions
    
    def getPartition(key: Any): Int = key match {
    	case null => 0
    	case _ => Utils.nonNegativeMod(key.hashCode, numPartitions)
    }
    
    override def equals(other: Any): Boolean = other match {
    	case h: HashPartitioner => h.numPartitions == numPartitions
    	case _ => false
    }
    
    override def hashCode: Int = numPartitions
}

三、RangePartitioner

将一定范围内的数据映射到一个分区中,尽量保证每个分区数据均匀,而且分区间有序

scala 复制代码
class RangePartitioner[K: Ordering: ClassTag, V](partitions: Int, rdd: RDD[_ <: Product2[K, V]], private var ascending: Boolean = true) extends Partitioner {
    // We allow partitions = 0, which happens when sorting an empty RDD under the default settings.
    require(partitions >= 0, s"Number of partitions cannot be negative but found 
    $partitions.")
    
    private var ordering = implicitly[Ordering[K]]
    // An array of upper bounds for the first (partitions - 1) partitions
    private var rangeBounds: Array[K] = {
    	...
    }
    
    def numPartitions: Int = rangeBounds.length + 1
    
    private var binarySearch: ((Array[K], K) => Int) =  CollectionsUtils.makeBinarySearch[K]
    
    def getPartition(key: Any): Int = {
        val k = key.asInstanceOf[K]
        var partition = 0
        if (rangeBounds.length <= 128) {
            // If we have less than 128 partitions naive search
            while(partition < rangeBounds.length && ordering.gt(k, rangeBounds(partition))) {
                partition += 1
            }
        } else {
            // Determine which binary search method to use only once.
            partition = binarySearch(rangeBounds, k)
            // binarySearch either returns the match location or -[insertion point]-1
            if (partition < 0) {
            	partition = -partition-1
            }
            
            if (partition > rangeBounds.length) {
                partition = rangeBounds.length
            }
    	}
        
        if (ascending) {
            partition
        } else {
            rangeBounds.length - partition
        }
    }
    
    override def equals(other: Any): Boolean = other match {
    	...
    }
    
    override def hashCode(): Int = {
    	...
    }
    
    @throws(classOf[IOException])
    private def writeObject(out: ObjectOutputStream): Unit =  Utils.tryOrIOException 
    {
    	...
    }
    
    @throws(classOf[IOException])
    private def readObject(in: ObjectInputStream): Unit = Utils.tryOrIOException {
    	...
    }
}

四、自定义 Partitioner

scala 复制代码
/**
	1.继承 Partitioner 抽象类
	2.重写 numPartitions: Int 和 getPartition(key: Any): Int 方法
*/
object TestRDDPartitioner {
    def main(args: Array[String]): Unit = {
        val conf = new SparkConf().setMaster("local[*]").setAppName("partition")
    	val sc = new SparkContext(conf)
        
        val rdd = sc.makeRDD(List(
        	("nba", "xxxxxxxxxxx"),
            ("cba", "xxxxxxxxxxx"),
            ("nba", "xxxxxxxxxxx"),
            ("ncaa", "xxxxxxxxxxx"),
            ("cuba", "xxxxxxxxxxx")
        ))
        
        val partRdd = rdd.partitionBy(new MyPartitioner)
        
        partRdd.saveAsTextFile("output")
        
    }
}

class MyPartitioner extends Partitioner {
    // 重写返回分区数量的方法
    override def numPartitions: Int = 3
    
    // 重写根据数据的key返回数据所在的分区索引的方法
    override def getPartition(key: Any): Int = {
        key match {
            case "nba" => 0
            case "cba" => 1
            case _ => 2
        }
    }
    
}
相关推荐
和裕16 分钟前
全纸结构重型纸箱能否满足 1 吨以上设备出口熏蒸豁免要求?通关合规性全解析
大数据·运维·网络·人工智能·算法
代码方舟33 分钟前
零信任架构实战:基于天远运营商三要素简版V即时版查询构建自动化电子投保实名核验网关
大数据·人工智能·架构·自动化
xiaohaiAIgeo2 小时前
【2026年】实验室应急预案中通风系统的关键作用
大数据·人工智能·科普知识
Sayai2 小时前
Elasticsearch 日志检索 DSL 实战:时间范围查询、字段去重、分钟级统计与最新日志获取
大数据·运维·elasticsearch·搜索引擎·日志分析
W***25923 小时前
2026 企业 AI 办公工具选型指南:可完成端到端任务的平台怎么评估
大数据·人工智能
今年下半年3 小时前
【Spring Boot】多种存储文件(FastDFS / MinIO / 阿里云 OSS)接入设计说明(附源码)
spring boot·分布式·中间件·简单工厂模式·策略模式
龙亘川4 小时前
数字化赋能基层协同治理:亘川智城一网统管平台落地实践思考
大数据·人工智能·智慧城市·开源软件·数据可视化
ShineWinsu5 小时前
对于Redis:string类型的解析
java·c++·redis·分布式·缓存·面试·string
可乐ea5 小时前
AI Agent 工具调用准确性评测:选择错误与参数错误分开测
大数据·人工智能·算法·大模型·工具调用·ai智能体·agent评测
数智顾问5 小时前
(90页PPT)IBM集团管理驾驶舱项目蓝图规划(附下载方式)
大数据·人工智能·物联网