ACM Multimedia 2026
The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. However, existing detectors suffer from two critical limitations: poor generalization to unseen generative architectures for manipulation detection and lack of interpretability. In this context, we present HexMIL (Hierarchical EXplainable Multiple Instance Learning), a mask-free medical deepfake detector that simultaneously addresses both limitations using only binary volume-level supervision. HexMIL decomposes each CT volume into a two-level hierarchy of patches and slices, aggregated via independent Gated Attention modules whose weights are directly combined into a full-resolution 3D attention volume that localizes the manipulated sub-region without any pixel-level annotation. Unlike post-hoc methods such as Grad-CAM, HexMIL's attention weights constitute the exact forward computation driving the classification decision, providing ante-hoc and structurally faithful spatial attribution. We evaluate HexMIL on M3DSynth and CT-GAN datasets under a rigorous cross-generator generalization protocol, training on a single generative architecture and testing on unseen ones. HexMIL outperforms all baselines by +9.1 AUC and +9.4 F1 in out-of-domain classification, and achieves the best average IoU and Pointing Game score in localization.
Evaluated under a rigorous out-of-distribution protocol: the detector is trained on a single generative architecture and tested on the others. Reported per cross-generator transfer and averaged. Bold = best, underline = second.
Cross-generator generalization across Pix2Pix, CycleGAN, Diffusion Model (DM) and CT-GAN.
| Method | Train: Pix2Pix | Train: CycleGAN | Train: DM | Train: CT-GAN | Avg | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CycleGAN | DM | CT-GAN | Pix2Pix | DM | CT-GAN | Pix2Pix | CycleGAN | CT-GAN | Pix2Pix | CycleGAN | DM | ||
| R3D-18 | 66.8/67.0 | 66.5/64.4 | 62.1/65.4 | 70.8/64.0 | 70.1/67.2 | 64.3/61.8 | 69.5/64.6 | 67.6/68.4 | 71.2/74.9 | 60.4/66.2 | 61.3/62.1 | 65.5/68.0 | 66.3/66.2 |
| ResNet50-ABMIL | 64.6/62.8 | 77.6/70.2 | 65.8/66.0 | 69.4/62.2 | 61.1/58.3 | 68.0/72.3 | 61.9/58.4 | 54.3/59.6 | 57.8/61.0 | 60.7/59.8 | 70.1/65.6 | 68.0/67.9 | 65.3/63.7 |
| ViT-ABMIL | 65.3/70.6 | 68.5/68.0 | 61.2/66.4 | 68.3/64.4 | 68.2/66.7 | 64.9/62.1 | 66.8/64.5 | 64.7/68.5 | 69.1/73.3 | 58.7/65.0 | 62.4/60.5 | 60.1/68.8 | 64.9/66.6 |
| DFX-SN | 73.1/76.0 | 70.9/70.6 | 68.4/74.2 | 63.4/63.1 | 75.0/65.8 | 59.7/61.3 | 67.4/68.5 | 69.7/69.9 | 72.1/70.4 | 66.2/71.5 | 64.8/63.4 | 67.9/75.8 | 68.2/70.0 |
| HP-FCN | 60.4/65.4 | 74.0/67.6 | 55.2/63.8 | 64.5/60.3 | 61.1/63.1 | 67.8/65.4 | 72.1/60.4 | 55.0/65.2 | 64.3/71.2 | 53.4/59.1 | 66.2/74.4 | 57.9/56.6 | 62.1/64.4 |
| MVSS-Net | 63.4/66.9 | 72.9/63.2 | 58.7/64.2 | 65.4/59.1 | 64.9/63.1 | 69.2/71.5 | 74.8/58.9 | 59.0/66.8 | 67.5/65.9 | 55.1/62.3 | 68.4/70.1 | 60.8/63.4 | 65.0/64.6 |
| FreqNet | 62.0/71.4 | 71.9/67.4 | 57.8/63.2 | 53.4/63.1 | 62.0/64.7 | 55.6/62.1 | 72.4/75.5 | 68.7/65.0 | 66.5/72.3 | 58.9/66.4 | 64.2/67.8 | 61.3/60.9 | 62.9/66.6 |
| NPR | 70.4/71.0 | 64.9/70.9 | 66.8/72.4 | 77.4/73.1 | 79.0/65.3 | 71.2/70.5 | 90.4/85.1 | 75.7/76.9 | 84.1/88.3 | 65.5/70.1 | 62.4/68.0 | 73.0/71.6 | 73.4/73.6 |
| TruFor | 84.0/75.9 | 84.9/70.4 | 71.2/76.8 | 83.4/63.1 | 82.0/65.6 | 78.5/82.4 | 90.4/65.5 | 78.7/74.9 | 85.1/89.3 | 71.4/74.2 | 81.3/79.5 | 72.8/75.0 | 80.3/73.3 |
| D³ | 75.0/76.0 | 84.9/75.4 | 68.4/74.2 | 88.0/87.1 | 92.2/82.6 | 83.2/81.5 | 89.4/90.5 | 79.7/76.9 | 85.7/93.3 | 72.1/78.9 | 79.4/77.2 | 70.3/76.5 | 80.7/80.8 |
| ManTraNet | 86.1/75.3 | 88.9/63.3 | 73.5/70.2 | 89.0/61.1 | 85.7/64.7 | 81.2/88.4 | 93.4/62.6 | 76.8/76.5 | 87.9/85.3 | 72.4/77.1 | 83.5/81.2 | 75.8/73.5 | 82.9/73.7 |
| HexMIL (Ours) | 91.6/86.2 | 99.3/97.5 | 88.4/92.1 | 97.1/93.5 | 98.3/96.0 | 94.2/91.8 | 98.7/96.5 | 90.1/80.8 | 90.5/84.2 | 85.1/90.3 | 86.4/83.0 | 84.2/90.0 | 92.0/90.2 |
Spatial grounding on unseen manipulation types, evaluated for the pixel-grounding detectors.
| Method | Train: Pix2Pix | Train: CycleGAN | Train: DM | Train: CT-GAN | Avg | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CycleGAN | DM | CT-GAN | Pix2Pix | DM | CT-GAN | Pix2Pix | CycleGAN | CT-GAN | Pix2Pix | CycleGAN | DM | ||
| HP-FCN | 14.9/20.8 | 23.1/32.3 | 16.9/21.5 | 26.4/34.0 | 8.8/17.3 | 25.3/29.8 | 35.9/31.2 | 7.2/13.4 | 34.3/28.8 | 11.5/16.3 | 22.7/33.8 | 16.5/16.8 | 20.3/24.7 |
| MVSS-Net | 41.1/60.3 | 33.9/50.6 | 38.5/63.1 | 40.8/56.8 | 30.3/43.8 | 38.1/58.9 | 46.2/66.5 | 29.3/45.4 | 46.1/67.4 | 42.7/56.2 | 33.1/49.9 | 37.3/59.1 | 38.1/56.5 |
| TruFor | 34.4/62.8 | 35.1/63.2 | 33.7/57.8 | 37.8/71.0 | 35.5/65.6 | 38.2/69.2 | 53.9/80.7 | 19.5/47.8 | 49.7/83.7 | 31.4/60.3 | 37.7/62.9 | 32.4/58.7 | 36.6/65.3 |
| ManTraNet | 38.2/67.2 | 43.7/69.0 | 33.7/63.1 | 41.4/71.9 | 33.9/63.7 | 42.4/73.7 | 50.4/83.5 | 10.6/33.0 | 47.9/81.9 | 34.3/65.5 | 44.5/71.2 | 34.9/66.0 | 38.0/67.5 |
| HexMIL (Ours) | 40.0/66.1 | 46.7/75.6 | 37.8/63.4 | 47.1/77.0 | 44.1/70.7 | 43.5/73.2 | 50.7/84.8 | 31.7/55.5 | 49.2/81.1 | 36.4/62.5 | 42.1/78.4 | 38.9/64.0 | 42.4/70.6 |
The manipulated sub-region emerges purely as a by-product of the binary classification objective — no pixel-level supervision is ever used.
If you find our work useful, please consider citing:
@inproceedings{pontorno2026hexmil,
title = {{HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes}},
author = {Pontorno, Orazio and Guarnera, Luca and Akhtar, Zahid and Battiato, Sebastiano},
booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
year = {2026},
}