flink 写入数据到 kafka 后,数据过一段时间自动删除

版本

  • flink 1.16.0
  • kafka 2.3

流程描述:

flink利用KafkaSource ,读取kafka的数据,然后经过一系列的处理,通过KafkaSink ,采用 EXACTLY_ONCE 的模式,将处理后的数据再写入到新的topic中。

问题描述:

数据写入到新的topic后,过上几分钟的时间,利用工具offset explorer观察对应topic的数据量,显示为0。

刚写入没多久的数据消失了 ???大写的懵 ???

定位问题:

  • 首先查看kafka的日志:
  • 阅读flink 官方文档 kafkaSink的介绍:

DeliveryGuarantee.EXACTLY_ONCE: In this mode, the KafkaSink will write

all messages in a Kafka transaction that will be committed to Kafka on

a checkpoint. Thus, if the consumer reads only committed data (see

Kafka consumer config isolation.level), no duplicates will be seen in

case of a Flink restart. However, this delays record visibility

effectively until a checkpoint is written, so adjust the checkpoint

duration accordingly. Please ensure that you use unique

transactionalIdPrefix across your applications running on the same

Kafka cluster such that multiple running jobs do not interfere in

their transactions! Additionally, it is highly recommended to tweak

Kafka transaction timeout (see Kafka producer transaction.timeout.ms)>>

maximum checkpoint duration + maximum restart duration or data loss

may happen when Kafka expires an uncommitted transaction.

  • 翻译过来的意思大概就是:

在EXACTLY_ONCE这种模式下,KafkaSink在事务中写入所有的消息,这些消息在checkpoint上提交给kafka。因此,在flink重启的情况下,如果消费者值读取提交的数据,不会看到重复的数据。缺点就是延迟记录可见性,知道写入检查点为止。强烈建议调整kafka的事务超时时间(见Kafka producer transaction.timeout.ms),超时时间要大于【最大检查点持续时间+最大重启持续时间】,否则当Kafka过期未提交的事务时可能会发生数据丢失。

  • 阅读kafka的官网介绍:

Producer Configs:
transaction.timeout.ms:60000(默认值)

参数描述:

The maximum amount of time in ms that the transaction coordinator will

wait for a transaction status update from the producer before

proactively aborting the ongoing transaction.If this value is larger

than the transaction.max.timeout.ms setting in the broker, the request

will fail with a InvalidTransactionTimeout error.

Broker Configs
transaction.max.timeout.ms:900000(默认值)

参数描述:

The maximum allowed timeout for transactions. If a client's requested

transaction time exceed this, then the broker will return an error in

InitProducerIdRequest. This prevents a client from too large of a

timeout, which can stall consumers reading from topics included in the

transaction.

  • 最后排查
    在flink中设置的超时时间违反了kafka producer对应的参数规定。

解决问题

在kafkaSink的配置中,加入

java 复制代码
Properties properties = new Properties();
// 根据上面的介绍自己计算这边的超时时间,满足条件即可
properties.setProperty("transaction.timeout.ms","900000");

KafkaSink<String> sink = KafkaSink.<String>builder()
                .setBootstrapServers(bootstrapServers)
                .setRecordSerializer(KafkaRecordSerializationSchema.<String>builder()
                        .setTopic(sinkTopic)
                        .setValueSerializationSchema(new SimpleStringSchema())
                        .build()
                )
                .setKafkaProducerConfig(properties)
                .setDeliveryGuarantee(DeliveryGuarantee.EXACTLY_ONCE)
                .setTransactionalIdPrefix("flink-xhaodream-")
                .build();

总结

在使用现有框架和工具的时候,往往只是懂得怎么用,具体底层的逻辑、原理,了解的很少。往往只有真正理解了原理,遇到了问题,才会更快、更准确的定位问题、解决问题。

相关推荐
java1234_小锋10 小时前
【免费】基于Spark实时电商用户行为分析与预测(Java版本+可视化大屏+Kafka+SpringBoot+Vue3) 锋哥原创出品,必属精品
java·spark·kafka·实时电商用户行为分析与预测系统
abcy07121320 小时前
flink datastream调用8种分区策略实例
flink
AI人工智能+电脑小能手1 天前
【大白话说Java面试题 第194题】【08_Kafka篇】第10题:简述 Kafka 的 Rebalance 机制?
java·kafka·消费者组·rebalance·分布式消息队列
yqj2342 天前
Flink集群配置与部署全攻略
大数据·flink
AI人工智能+电脑小能手2 天前
【大白话说Java面试题 第192题】【08_Kafka篇】第8题:死信队列是什么?延时队列是什么?
java·kafka·消息队列·死信队列·延时队列
abcy0712132 天前
spark核心组件
flink·spark
Apache Flink3 天前
全模态入湖,把大模型接入实时湖仓:Flink OSS CDC + DLF Paimon 实现零代码以图搜图
大数据·flink
AI人工智能+电脑小能手3 天前
【大白话说Java面试题 第190题】【08_Kafka篇】第6题:消息队列有什么作用?
java·kafka·消息队列·系统设计·分布式架构
风中凌乱3 天前
kafka新版本集群的安装与部署
分布式·kafka
abcy0712133 天前
kafka消息丢失与消息重复消费
kafka