BetaEntity Annotation Prototype
← Back to diseases

Annotated abstract

IDDF2026-ABS-0273 Deep learning as a safety net for capsule endoscopy interpretation

gutjnl · 2026-06-26 · canonical JSON source

2 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Background Video capsule endoscopy (VCE) generates thousands of frames per examination, requiring 1-2 hours of manual review, limiting throughput and increasing reader fatigue. We aimed to develop a deep learning system to automatically classify VCE frames across 14 small bowel findings and generate anomaly scores to prioritise clinically significant segments for physician review.Methods We utilised the Kvasir-Capsule dataset (47,238 images, 14 classes; 37,790/9,448 train/validation split). EfficientNet-B0 pretrained on ImageNet was fine-tuned with focal loss (γ=2.0, inverse-frequency α-weighting) to address severe class imbalance (3,434:1 ratio). Weighted random sampling ensured adequate representation of rare findings including polyps (n=55), ampulla of Vater (n=10), and blood-hematin (n=12). Data augmentation included random flips, rotations, and colour jitter. Frame-level anomaly scores were computed as 1-P(normal), simulating clinical pre-screening. Model interpretability was assessed using Grad-CAM activation maps.Results Training loss converged rapidly with no overgap between training and validation curves, confirming stable generalisation ( IDDF2026-ABS-0273 Figure 1. Training and validation focal loss curves across 10 epochs). The model achieved macro F1=0.595, validation accuracy=0.674, and mean per-class recall of 0.906 (IDDF2026-ABS-0273 Figure 3. Normalized per-class recall confusion matrix). Per-class performance (IDDF2026-ABS-0273 Figure 2. Validation macro F1 and accuracy across 10 epochs) demonstrated high sensitivity for clinically critical findings: polyp (1.000), blood-hematin (1.000), foreign body (0.994), blood-fresh (0.989), ulcer (0.982), and angiectasia (0.977). The most challenging class was normal clean mucosa (recall=0.56), reflecting expected difficulty distinguishing normal from subtle pathology, not a safety concern as the system is designed to flag abnormalities.In a simulated clinical reading of 2,000 sequential frames, the anomaly detection system achieved 99.8% sensitivity for abnormal frames, reducing frames requiring physician review by 35%, compressing 2-hour manual reads to a prioritised subset (IDDF2026-ABS-0273 Figure 3. Normalized per class recall confusion matrix).Conclusions A lightweight EfficientNet-B0 model with focal loss and weighted sampling effectively classifies 14 VCE findings despite extreme class imbalance. Near-perfect sensitivity for abnormal frames supports its use as a clinical pre-screening tool, substantially reducing reading time without compromising safety ( IDDF2026-ABS-0273 Figure 4. Model performance summary). Grad-CAM interpretability enhances clinical trust. Future work will incorporate temporal modelling across consecutive frames and prospective validation.Abstract IDDF2026-ABS-0273 Figure 1Abstract IDDF2026-ABS-0273 Figure 2Abstract IDDF2026-ABS-0273 Figure 3Abstract IDDF2026-ABS-0273 Figure 4