Beat 01
The Dilemma
Whole-body bone scans are the go-to screen for cancer that has spread to bone. Reading them is slow, tiring work, readers don't always agree, and harmless wear-and-tear hotspots are easy to mistake for disease.
Status: Critical · Clinical screening bottleneck
- Screening mode
- Manual whole-body reads
- Reader fatigue
- High at volume
- Confounders
- Harmless wear-and-tear hotspots
- Failure mode
- Reader disagreement
Baseline: qualitative problem signals
Beat 02
The Hypothesis
If a team of very different CNNs can agree on a diagnosis, one small network should be able to learn their judgment, and small enough to be practical in a busy clinic.
System blueprint
- STAGE 01Whole-body scintigraphyAnterior / posterior bone scans
- STAGE 0212 CNN teachersEight families, fine-tuned from ImageNet
- STAGE 03Best-4 ensemble0.979 F1 · 0.997 AUC, AUC-weighted voting
- STAGE 04Knowledge distillationBlended focal loss on teacher + true labels
- STAGE 05ScintNet3.27M params, anisotropic kernels, hybrid pooling
- STAGE 06Dynamic quantization13.11 MB → 9.96 MB (head only)
- STAGE 07Grad-CAM++ · LIME · uptake reportIs it looking at real hotspots?
Beat 03
The Execution
Fine-tuned 12 ImageNet CNNs from eight families, then searched for the four whose errors complement each other: ResNet18, ResNet50, InceptionV3 and EfficientNetB3 (0.979 F1, 0.997 AUC together). Distilled them into ScintNet, a 3.27M-parameter student shaped for tall 1024×256 scans, and quantized its head from 13.11 MB to 9.96 MB.
Knowledge distillation · interactive
Numbers from the paper
Teacher pool
- ResNet18
- DenseNet121
- MobileNetV3-L
- EfficientNetB3
- ConvNeXt-T
- ResNet50
- EfficientNetB0
- Xception
- VGG16
- EffNetV2-S
- InceptionV3
- EfficientNetB1
Student
ScintNet
3.27M params · hybrid pooling head
○ Full-precision weights · 13.11 MB
Twelve ImageNet CNNs from eight families are fine-tuned on whole-body scans, each one a diagnostician in its own right.
Validation
Every positive call comes with Grad-CAM++ heatmaps, LIME regions and an uptake report by body region, so a clinician can check it's looking at real hotspots and not scanner noise. Honest caveat: it still needs external, prospective validation.
Illustrative synthetic render · not patient data
Explainability audit
Right for the right reasons
A model can score well by latching onto artefacts. Overlaying Grad-CAM and LIME attributions checks that ScintNet's evidence sits on true skeletal lesions, and that degenerative uptake such as arthritic knees is not what drives a positive call.
- Thoracic spine lesionAttended
- Left rib lesionAttended
- Iliac lesionAttended
- Degenerative knee uptakeIgnored
Beat 04
The Impact
A student 3.3× smaller than its lightest teacher that keeps 97.7% of the ensemble's AUC, and shows its evidence.
- Accuracy
- 94.9%
- Held-out test set
- F1 Score
- 0.944
- Distilled student
- AUC
- 0.975
- Distilled student
- Specificity
- 99.0%
- 1 false alarm in 99 healthy scans
- Sensitivity
- 87.9%
- Caught 51 of 58 metastases
- Ensemble AUC
- 0.997
- Best-4 teacher ensemble
ScintNet on the 157-scan held-out test set at the recall-aware threshold (0.570). Ensemble AUC for reference. Needs external validation.