1999.VLDB.Finding intensional knowledge of distance-based outliers

1999.VLDB.Finding intensional knowledge of distance-based outliers

paper

pdf

main idea

intensional knowledge:

a description or an explanation of why an identified outlier is exceptional.

two main issue:

what kinds of intensional knowledge to provide;

how to optimize the computation of such knowledge.

contribution

1、define two notions of outliers and the corresponding intensional knowledge: strongest outlier and weak outlier.

2、develop a naive and semi-naive algorithm for computing strongest outlier and weak outlier and the corresponding intensional knowledge.

3、effective sharing of IO and experiment.

method

other's citation

citation 1

authors identify outliers in subspaces of the features space using a distance-based anomaly detection method. This serves as explanation since the identified anomalies are outliers in the specific subspaces found, meaning that the features constituting the subspace are those that discriminate the most the instance. The authors introduce the notions of strongest, weak and trivial outliers. An outlier is non-trivial in a subspace A if it is not

an outlier in any subspace included in A. A strongest outlier is an outlier in a strongest outlying feature space (if no outlier exists in any subspace included in A, then A is a strongest feature space). A weak outlier is a non-trivial not strongest outlier. Algorithms are provided to identify (and thus explain) strong and weak outliers. This anomaly explanation method is model-specific because it is designed for distance-based methods. It is also local because it helps explaining one outlier at a time.1

citation 2

For example, Knorr and Ng 50 define the outlier categories C = {"trivial outlier," "weak outlier," "strongest outlier"} to help gain better insights about the nature of outliers. They define an anomalous data point o as the "strongest outlier" in a subspace A if it meets two criteria: (i) o is not an outlier in any subspace B ⊂ A, and (ii) no outlier exists in any subspace B ⊂ A.If o does not satisfy the criteria in (i), then it is a "weak outlier." If it does not fit the two criteria, then it is a "trivial outlier." The terms "trivial outlier," "weak outlier," and "strongest outlier" are used to separate noise from meaningful abnormal data 2.

Figure 3 shows an illustration of strongest, weak, and trivial outliers in the 3D space {A, B, C}. P1 and P5 are non-trivial outliers in the subspace AB because they are not outliers in subspace A or subspace B. They are also the strongest outliers in AB because there is no other anomalous point in the subspace A or B. P20 is a weak outlier in the subspace AC because there is another outlier point P11 in the subspace C. P11 is a trivial outlier in the subspace AC because it is also an outlier in the subspace C.3

definitions



experiment

reference

12022.DKE.Anomaly explanation A review :3.1.1. Non-weighted feature importance

22022.VLDB.A survey on outlier explanations:2.1.2 Categorical ranking of outliers

32022.VLDB.A survey on outlier explanations:5.2 Techniques to find categorical rankings of

outliers

相关推荐
这张生成的图像能检测吗2 天前
(论文速读)FE-CLIP:把频域信息注入 CLIP,做零样本异常检测与分割
人工智能·opencv·目标检测·计算机视觉·缺陷检测·异常检测
这张生成的图像能检测吗2 天前
(论文速读)LogSAD:无需训练的结构异常与逻辑异常统一检测
人工智能·机器学习·计算机视觉·大模型·多模态·异常检测
zxsz_com_cn4 天前
预测性维护中的时间序列异常检测:从3sigma到Transformer
transformer·工业4.0·异常检测·时间序列·预测性维护
这张生成的图像能检测吗5 天前
(论文速读)DefectDiffu:基于一致性建模的少样本工业缺陷图像生成
人工智能·深度学习·计算机视觉·异常检测·少样本学习·扩散生成
EDPJ1 个月前
(2026|IPI|我的论文投稿中,PSP 超轻量采样算法,轻量化架构探索/早融合+头压缩)LUMIN:面向工业异常检测的轻量级通用制造检测网络
算法·计算机视觉·架构·异常检测·采样算法
EDPJ1 个月前
(2026|CVPR|安大,可学习正常/异常 token,空间感知交叉注意力)VisualAD:基于 ViT 的语言无关零样本异常检测
深度学习·计算机视觉·缺陷检测·异常检测
这张生成的图像能检测吗1 个月前
(论文速读)SubspaceAD:用 PCA 子空间做 Training-Free Few-Shot 异常检测
人工智能·机器学习·pca·异常检测·少样本学习
程序员学习Chat5 个月前
异常检测Anomalib库使用说明
异常检测·anomalib
程序员学习Chat5 个月前
计算机视觉-异常检测
人工智能·计算机视觉·异常检测
__土块__5 个月前
AI 管理后台首页信息过载治理:从指标泛滥到决策摘要的视图重构实践
异常检测·可观测性·故障排查·信息架构·ai工程·管理后台设计·状态机建模