《Towards Black-Box Membership Inference Attack for Diffusion Models》论文笔记

《Towards Black-Box Membership Inference Attack for Diffusion Models》

Abstract

  1. 识别艺术品是否用于训练扩散模型的挑战,重点是人工智能生成的艺术品中的成员推断攻击------copyright protection
  2. 不需要访问内部模型组件的新型黑盒攻击方法
  3. 展示了在评估 DALL-E 生成的数据集方面的卓越性能。

作者主张

previous methods are not yet ready for copyright protection in diffusion models.

Contributions(文章里有三点,我觉得只有两点)

  1. ReDiffuse:using the model's variation API to alter an image and compare it with the original one.
  2. A new MIA evaluation dataset:use the image titles from LAION-5B as prompts for DALL-E's API 31 to generate images of the same contents but different styles.

Algorithm Design

target model:DDIM

为什么要强行引入一个版权保护的概念???

定义black-box variation API

x ^ = V θ ( x , t ) \hat{x}=V_{\theta}(x,t) x^=Vθ(x,t)

细节如下:

总结为: x x x加噪变为 x t x_t xt,再通过DDIM连续降噪变为 x ^ \hat{x} x^

intuition

Our key intuition comes from the reverse SDE dynamics in continuous diffusion models.

one simplified form of the reverse SDE (i.e., the denoise step)
X t = ( X t / 2 − ∇ x log ⁡ p ( X t ) ) + d W t , t ∈ 0 , T (3) X_t=(X_t/2-\nabla_x\log p(X_t))+dW_t,t\in0,T\tag{3} Xt=(Xt/2−∇xlogp(Xt))+dWt,t∈0,T(3)

The key guarantee is that when the score function is learned for a data point x, then the reconstructed image x ^ i \hat{x}_i x^i is an unbiased estimator of x x x.(算是过拟合的另一种说法吧)

Hence,averaging over multiple independent samples x ^ i \hat{x}_i x^i would greatly reduce the estimation error (see Theorem 1).

On the other hand, for a non-member image x ′ x' x′, the unbiasedness of the denoised image is not guaranteed.

details of algorithm:

  1. independently apply the black-box variation API n times with our target image x as input
  2. average the output images
  3. compare the average result x ^ \hat{x} x^ with the original image.

evaluate the difference between the images using an indicator function:
f ( x ) = 1 D ( x , x \^ ) \< τ f(x)=1D(x,\\hat{x})\<\\tau f(x)=1D(x,x\^)\<τ

A sample is classified to be in the training set if D ( x , x ^ ) D(x,\hat{x}) D(x,x^) is smaller than a threshold τ \tau τ ( D ( x , x ^ ) D(x,\hat{x}) D(x,x^) represents the difference between the two images)

ReDiffuse
Theoretical Analysis

什么是sampling interval???

MIA on Latent Diffusion Models

泛化到latent diffusion model,即Stable Diffusion

ReDiffuse+

variation API for stable diffusion is different from DDIM, as it includes the encoder-decoder process.
z = E n c o d e r ( x ) , z t = α ‾ t z + 1 − α ‾ t ϵ , z ^ = Φ θ ( z t , 0 ) , x ^ = D e c o d e r ( z ^ ) (4) z={\rm Encoder}(x),\quad z_t=\sqrt{\overline{\alpha}_t}z+\sqrt{1-\overline{\alpha}t}\epsilon,\quad \hat{z}=\Phi{\theta}(z_t,0),\quad \hat{x}={\rm Decoder}(\hat{z})\tag{4} z=Encoder(x),zt=αt z+1−αt ϵ,z^=Φθ(zt,0),x^=Decoder(z^)(4)
modification of the algorithm

independently adding random noise to the original image twice and then comparing the differences between the two restored images x ^ 1 \hat{x}_1 x^1 and x ^ 2 \hat{x}_2 x^2:
f ( x ) = 1 D ( x \^ 1 , x \^ 2 ) \< τ f(x)=1D(\\hat{x}_1,\\hat{x}_2)\<\\tau f(x)=1D(x\^1,x\^2)\<τ

Experiments

Evaluation Metrics
  1. AUC
  2. ASR
  3. TPR@1%FPR
same experiment's setup in previous papers 5, 18.
target model DDIM Stable Diffusion
version 《Are diffusion models vulnerable to membership inference attacks?》 original:stable diffusion-v1-5 provided by Huggingface
dataset CIFAR10/100,STL10-Unlabeled,Tiny-Imagenet member set:LAION-5B,corresponding 500 images from LAION-5;non-member set:COCO2017-val,500 images from DALL-E3
T 1000 1000
k 100 10
baseline methods 5Are diffusion models vulnerable to membership inference attacks?: SecMIA 18An efficient membership inference attack for the diffusion model by proximal initialization. 28Membership inference attacks against diffusion models
publication International Conference on Machine Learning arXiv preprint 2023 IEEE Security and Privacy Workshops (SPW)
Ablation Studies
  1. The impact of average numbers
  2. The impact of diffusion steps
  3. The impact of sampling intervals
相关推荐
菩提小狗7 分钟前
每日安全情报报告 · 2026-10-01
网络安全·漏洞·cve·安全情报·每日安全
m4Rk_18 小时前
【论文阅读】Agent 记忆机制(86):Skill-Pro——用 Non-Parametric PPO 将交互经验演化为可复用技能
论文阅读·人工智能·学习·开源·github
黄金龙PLUS1 天前
5个1280比特大状态置换算法的设计与分析
算法·网络安全·密码学·哈希算法·同态加密
三8441 天前
TLS + Web API 安全 · 03 · TLS 实战:证书、漏洞与抓包中间人
网络安全·tls·中间人·ca证书
Rocky Ding*1 天前
DeepSeek DSec技术深度解析:Agent规模化训练的真正瓶颈,是沙箱基础设施
论文阅读·人工智能·深度学习·机器学习·aigc·agent·ai-native
大棉花哥哥1 天前
太空安全,正式上线 —— 博客改版与双专栏发布记
安全·网络安全·卫星安全
光依旧1 天前
MCP实战手记(九):生产化MCP Server的6层安全防护
java·人工智能·spring boot·安全·网络安全·ai agent·mcp
admin and root2 天前
「AI安全篇」实战AntiDebug自动化JS逆向加解密MCP
javascript·人工智能·网络安全·自动化·漏洞挖掘·cnvd·src赏金
泛联新安2 天前
软件定义汽车时代,如何让AI研发“可信”?——泛联新安构建汽车企业级可信AI体系的落地路径
大数据·人工智能·安全·网络安全·汽车·漏洞挖掘·代码安全
菩提小狗2 天前
每日安全情报报告 · 2026-09-29
网络安全·漏洞·cve·安全情报·每日安全