Abstract
Computing image anomaly score from the maximum of the anomaly segmentation prediction result has been widely adopted for end-to-end anomaly detection approaches. However, slight discrepancy in predicted pixel-level anomaly scores for normal and anomalous features often leads to high segmentation accuracy but unmatched poor detection performance. To overcome this problem, we propose a novel siamese-based U-Net model based on a contrastive learning framework combined with deviation-based detection finetuning strategy. The model is trained to drag normal features together while alienating the anomaly samples. Moreover, we introduce a novel channel-positional attention module (CPAM) in our U-Net decoder for refined feature upsampling. Our model reaches SOTA performance on the well-known 2D MVTecAD dataset and outperforms all other methods on the challenging dataset MVTec3D-AD by a large margin.