Flink基础

Flink
architecture

job manager is master

task managers are workers

task slot is a unit of resource in cluster, number of slot is equal to number of cores(超线程则slot=2*cores), slot=一组内存+一些线程+共享CPU

when starting a cluster,job manager will allocate a certaion number of slots to each taskManager in cluster,

each slots can run one parallel instance of a task or operator
tasks as a basic unit of work execution physically

each task corresponds to a logical reperesentation of data processiong (entire job chain excution )

a subtask represents some operators physically. which is concrete and excutable with other subtasks run in paralle in the same task slot,Flink will process the excution by chaining compatible oeprators if can be chained in same slot to reduce data shuffling
Subtask 是 Flink 作业中 Operator 的并行实例。每个 Operator 都可以拥有一个或多个 subtask,这些 subtask 是并行执行的,运算符子任务(subtask)的数量是该特定运算符的并行度

subtask scheduling

if parallelism is 6, six parallel instances will go across the available task slots.

Flink will process the excution by chaining compatible oeprators if can be chained in same slot to reduce data shuffling

if key by,then all data with same key will be processed in the same slot for accurate state management

**key by group by or window operation need data shuffling(**data movement between nodes)

operator会被chain在同一subtask的情况

(1)手动设置setChainingStrategy(ChainingStrategy.ALWAYS)

.map(x => x * 2)

.filter(x => x > 2)

.setChainingStrategy(ChainingStrategy.ALWAYS)

(2)keyby分区后,相同数据的后续所有操作都在同一个subtask中

keyBy(keySelector).map(...).filter(...) .print();

(3)并行度相同的operators通常可能被chain在一起减少data shuffling

flink Window窗口

在一个无界流中设置起始位置和终止位置,让无界流变成有界流,并且在有界流中进行数据处理,流批转化

  • window窗口在无界流中设置起始位置和终止位置的方式可以有两种 ,基于时间或者基于窗口数据量,
  • 分组和未分组窗口。自定义窗口
  • 时间窗口:
  • 滚动窗口: 数据不重复
  • 滑动窗口:数据有重复
  • 窗口聚合函数:
  • 增量聚合:ReduceFunction、AggregateFunction
  • 全量聚合 ProcessWindowFunction、WindowFunction属于全量窗口函数
相关推荐
航飞光电市场经理9 分钟前
UWB定位技术选型指南:从芯片架构到定位引擎的底层能力分析
大数据·人工智能·物联网·安全·人员定位
starzy19909 分钟前
Flink Sliding Window 详解及代码实现:从窗口重叠到状态爆炸防控
运维·数据库·flink
Raas10037 分钟前
MAI Gateway(魔芋企业级AI网关)功能全解:AI网关支持本地模型吗?一文看懂AI网关能力矩阵
大数据·人工智能·网关·ai网关·mai gateway·企业级产品
爱签AI电子合同1 小时前
电子合同服务稳定性怎么测?可用性保障维度专项测评
大数据·人工智能·电子合同·电子签名
跨境数据猎手1 小时前
从零搭建多平台二手ERP中台:闲鱼、淘宝、京东、拼多多、Mercari统一调度架构
大数据·系统架构·团队开发
Francek Chen1 小时前
【大数据处理与分析】数据仓库Hive:04 数据仓库Hive概述
大数据·数据仓库·hive·hadoop·分布式
Leo.yuan2 小时前
2026国产数据仓库软件有哪些?从数据库、云数仓到数据集成平台一次讲清
大数据
红姐跨境书2 小时前
高并发场景下的本地缓存进化论:从 Go sync.Map 到 BigCache 的性能调优实践
大数据
shujudang2 小时前
业务数据分析项目中的分析方法与团队协作
大数据·数据挖掘·数据分析
Gl�ria2 小时前
Yarn NM 常驻Flink任务下线:stop/savepoint 释放容器
flink·yarn