Evaluating the Generalisability of Convolutional Neural Networks for Diabetic Retinopathy Detection in Latin America and Sub-Saharan Africa

Abstract

Diabetic retinopathy is a leading cause of vision loss worldwide, particularly impacting individuals in low- and middle-income countries with limited healthcare access. Early detection through automated screening systems is essential for improving outcomes, as timely intervention can prevent severe vision impairment. However, most of the available AI models have not been evaluated in low-resource settings. Hence, this study presents an evaluation of the efficacy of advanced deep learning architectures for detecting rDR across diverse population datasets. A dual-phase validation approach was employed to assess model performance. Internal validation utilised the BrSET dataset to establish baseline performance metrics, while external validation was conducted on the MoDRIA dataset, which encompasses various conditions and demographics, to evaluate model robustness. Key performance metrics, including accuracy, specificity, sensitivity, F1-score, and calibration scores, were systematically recorded and analysed. Internal validation revealed high accuracy across all models, EfficientNetB0 achieved the highest classification accuracy (0.9561; 95% CI 0.9490–0.9630), EfficientNetB3 demonstrated superior overall discriminative performance, achieving the highest AUROC (0.9892; 95% CI 0.9841–0.9934) highest sensitivity (0.9573), and lowest Brier score (0.0168). Meanwhile, DenseNet ex hibited the most balanced clinical screening performance, achieving the highest F1-score (0.7259; 95% CI 0.6797–0.7669) and Youden Index (0.2381), indicating improved balance between sensitivity and specificity. In contrast, external validation revealed substantial deterioration in model performance across all architectures, highlighting major limitations in cross-population generalisability. Although EfficientNetB0 achieved the highest external accuracy (0.8821; 95% CI 0.8746–0.8898), AUROC values declined markedly across models (0.5140–0.6104), accompanied by poor sensitivity, reduced F1-scores, and substantial calibra tion instability. EfficientNetB3 achieved the highest external sensitivity (0.5939), whereas calibration analyses demonstrated unreliable probability estimation under domain-shift conditions. These findings suggest that AI models trained on geographically homogeneous retinal imaging datasets may not generalise reliably across underrepresented populations. Population differences and imaging variability substantially affected external model per formance, highlighting the need for diverse datasets, rigorous external validation, and adaptive recalibration before clinical deployment of AI-driven DR screening systems.

Description

Citation

Mwavu, R., Kaggwa, F., Arunga, S., & Wasswa, W. (2026). Evaluating the Generalisability of Convolutional Neural Networks for Diabetic Retinopathy Detection in Latin America and Sub-Saharan Africa. Information, 17(6), 552.

Endorsement

Review

Supplemented By

Referenced By