Flink基础

Flink
architecture

job manager is master

task managers are workers

task slot is a unit of resource in cluster, number of slot is equal to number of cores(超线程则slot=2*cores), slot=一组内存+一些线程+共享CPU

when starting a cluster,job manager will allocate a certaion number of slots to each taskManager in cluster,

each slots can run one parallel instance of a task or operator
tasks as a basic unit of work execution physically

each task corresponds to a logical reperesentation of data processiong (entire job chain excution )

a subtask represents some operators physically. which is concrete and excutable with other subtasks run in paralle in the same task slot,Flink will process the excution by chaining compatible oeprators if can be chained in same slot to reduce data shuffling
Subtask 是 Flink 作业中 Operator 的并行实例。每个 Operator 都可以拥有一个或多个 subtask,这些 subtask 是并行执行的,运算符子任务(subtask)的数量是该特定运算符的并行度

subtask scheduling

if parallelism is 6, six parallel instances will go across the available task slots.

Flink will process the excution by chaining compatible oeprators if can be chained in same slot to reduce data shuffling

if key by,then all data with same key will be processed in the same slot for accurate state management

**key by group by or window operation need data shuffling(**data movement between nodes)

operator会被chain在同一subtask的情况

(1)手动设置setChainingStrategy(ChainingStrategy.ALWAYS)

.map(x => x * 2)

.filter(x => x > 2)

.setChainingStrategy(ChainingStrategy.ALWAYS)

(2)keyby分区后,相同数据的后续所有操作都在同一个subtask中

keyBy(keySelector).map(...).filter(...) .print();

(3)并行度相同的operators通常可能被chain在一起减少data shuffling

flink Window窗口

在一个无界流中设置起始位置和终止位置,让无界流变成有界流,并且在有界流中进行数据处理,流批转化

  • window窗口在无界流中设置起始位置和终止位置的方式可以有两种 ,基于时间或者基于窗口数据量,
  • 分组和未分组窗口。自定义窗口
  • 时间窗口:
  • 滚动窗口: 数据不重复
  • 滑动窗口:数据有重复
  • 窗口聚合函数:
  • 增量聚合:ReduceFunction、AggregateFunction
  • 全量聚合 ProcessWindowFunction、WindowFunction属于全量窗口函数
相关推荐
夜郎king39 分钟前
解决AI图文解析偏差:CodeBuddy多模型交叉校验+腾讯地图Skill经纬度定位实战
大数据·人工智能
智慧物业老杨40 分钟前
物业费公共收益维修资金三类资金分账的合规管控
大数据·人工智能·物联网
不动明王198444 分钟前
Presto 查询引擎内核详解:Worker 本地执行模型——从物理计划到可调度执行单元
大数据·presto·湖仓·查询引擎内核·pipeline执行
tkevinjd1 小时前
MiniCode 项目详解7:Memory 记忆系统
大数据·python·搜索引擎·llm·agent
流量猎手2 小时前
GitHub 使用说明
大数据·elasticsearch·github
中国搜索直付通2 小时前
防沉迷新规下的棋牌游戏生存术:从合规底线到用户体验升级
大数据·人工智能·游戏
志栋智能3 小时前
超自动化安全:提升安全服务满意度的隐形引擎
大数据·安全·自动化
BerrySen1783 小时前
一个Java项目改成AI流程后,最难的部分完全变了
java·大数据·人工智能·可观测性·大模型应用开发·工程思维
核数聚3 小时前
【赛迪专访核数聚】深耕数据治理,打通数据孤岛夯实 AI 发展根基
大数据·人工智能·算法
北京晶数信息科技3 小时前
加油站成品油智慧监管平台+交易即开票一体化解决方案 (一)
大数据·人工智能·物联网·产品经理·需求分析