
AI's reach in retinal disease: Lab bench to bedside to backpack
New research covers a 3D OCT foundation model, an OCTA biomarker review and offline smartphone screening for three retinal diseases.
Retinal artificial intelligence (AI) is often discussed in isolation, with studies focusing on a single setting such as the reading centre, biomarker research or point-of-care screening. Three recent studies illustrate how these applications are evolving across that spectrum. One study introduces a foundation model designed to automate the 3D optical coherence tomography (OCT) reading workflow in tertiary referral centres. Another study systematically reviews the reliability of current OCT angiography (OCTA) biomarkers for detecting early diabetic retinopathy (DR). A third study evaluates an offline smartphone platform designed to screen for multiple eye diseases in real-world clinical settings in India.
Together, the studies trace the expanding role of retinal AI from specialist interpretation and biomarker development to accessible, point-of-care screening.
The lab bench: Automating the full OCT workflow
OCT has become central to retinal disease diagnosis because of its high-resolution, 3D imaging, but its use for large-scale screening has been limited by a labour-intensive workflow: image acquisition, selection of pathology-relevant slices and expert interpretation. A team led by researchers at Zhongshan Ophthalmic Center, Sun Yat-sen University, along with collaborators at several Chinese institutions, set out to automate that entire pipeline with a system they call the Full-process OCT-based Clinical Utility System (FOCUS), described in npj Digital Medicine.¹
FOCUS operates in two stages. First, an image quality assessment module built on EfficientNetV2-S screens out unusable OCT slices. Usable slices are then moved to a diagnostic stage built on a fine-tuned Vision Foundation Model, which extracts features from individual 2D B-scans. A component the authors call the Unified Adaptive Aggregation Classifier then fuses those slice-level predictions into a single patient-level diagnosis, weighting each slice by its diagnostic relevance and prediction confidence rather than treating all slices equally. The system was designed to detect nine categories: normal, AMD, choroidal neovascularization, central serous chorioretinopathy, DR, epiretinal membrane, macular hole, macular oedema and retinitis pigmentosa.¹
FOCUS was trained and tested on 3,300 patients (40,672 slices) and externally validated on 1,345 patients (18,498 slices) drawn from four tiered clinical centres in China using different OCT devices (Spectralis SD-OCT, Cirrus HD-OCT and DRI OCT Triton). It achieved F1-scores of 99.01% for quality assessment, 97.46% for abnormality detection and 94.39% for patient-level diagnosis on internal testing, with real-world external validation holding stable at F1-scores between 90.22% and 95.24% across the four centres.¹
In a human-machine comparison using 842 OCT images from 105 patients, FOCUS matched or exceeded ophthalmic technicians in flagging abnormal scans (F1-score, 95.47% versus an average of 90.91% for four technicians) and matched retinal specialists in multi-disease diagnosis (F1-score, 93.49% versus 91.35% for the specialist average), while substantially outperforming junior doctors (62.35%) and senior doctors (76.28%) on the same task. On efficiency, FOCUS processed the full 842-image test set in 22.7 seconds, compared with an average of 29.3 minutes for junior doctors, 40.8 minutes for senior doctors and 32.1 minutes for retinal experts versus 83.4 minutes for technicians completing the separate screening task.¹
The authors note that their training and validation data came primarily from Chinese medical centres, which may limit generalizability to other populations, and that performance on several public international OCT benchmarks was lower than on their in-domain test sets. According to the authors, this is evidence that broader validation across different populations, disease spectra and device vendors is still needed before wider deployment.¹
In an exclusive quote for Ophthalmology Times Europe, Peng Xiao, PhD, professor at the Ophthalmic Engineering Research Center, State Key Laboratory of Ophthalmology, Zhongshan Ophthalmic Center, Sun Yat-sen University, in Guangzhou, China, and corresponding author on the FOCUS study, shared the following:
“FOCUS represents a shift from developing isolated AI algorithms to engineering a clinically aligned, deployable system. In our validation, it matched retinal specialist-level accuracy while processing an entire patient volume in seconds rather than the tens of minutes required by human graders. That speed advantage is not merely incremental—it changes the fundamental economics of OCT screening. It means we can move OCT from a concentrated, hospital-based expert procedure to a decentralized screening modality deployed in primary care or even community settings, where the bottleneck has always been the availability of trained interpreters.
“Our current training and validation data were drawn primarily from Chinese medical centers, and we are transparent about the need for broader validation. For FOCUS to become a truly global tool, we must test it prospectively across diverse ethnic populations, disease spectra, and healthcare systems—particularly in settings where disease prevalence and imaging protocols differ from our development cohorts. International multi-centre prospective trials, and extending the framework to handle multi-label, co-morbid cases, are critical next steps.
“Regarding its role in the clinic, we do not see FOCUS as a replacement for specialists, but as an intelligent triage and quality-assurance layer. It automates the labor-intensive stages of quality control, anomaly flagging, and initial classification, while providing an uncertainty score that tells the clinician when to trust the output and when to intervene. In this way, it standardizes the consistency of the screening pipeline and allows specialists to concentrate their expertise on the complex, ambiguous cases that most need human judgement.”
The bedside: The biomarker evidence behind the imaging
Automated systems like FOCUS depend on the underlying imaging data being diagnostically reliable—and for OCTA specifically, that evidence base is still being sorted out. A systematic review by Loukia A. Politi and colleagues, published in the European Journal of Ophthalmology, examined 21 studies evaluating OCTA-derived biomarkers for early DR detection, following PRISMA methodology and searching PubMed, PubMed Central and the Cochrane Clinical Trials Database through December 2025.²
Because of substantial heterogeneity across study designs, OCTA protocols and outcome measures, the authors determined a formal meta-analysis was not feasible and instead conducted a qualitative synthesis. Among the biomarkers examined, vessel density (VD) emerged as the most consistently reported: 13 of 15 studies that assessed it found reduced VD in diabetic eyes with early or no clinically apparent retinopathy compared with healthy controls, and seven of eight studies examining disease progression found VD declining further with increasing severity. Enlargement of the foveal avascular zone (FAZ) was reported in nine of 17 studies but not found in the remaining eight—a split the authors attribute partly to individual variation in FAZ size (influenced by age, axial length and refractive error) and partly to inconsistent scan protocols and segmentation methods across studies. Microaneurysms were more frequent in diabetic eyes in three of five studies evaluating them, and increased vessel tortuosity was reported in two of five.²
Only three of the 21 studies reported formal diagnostic accuracy metrics. One found VD and FAZ parameters achieving sensitivities of 93.62% and 91.64%, respectively, for distinguishing DR stages from controls; another reported a VD-based parameter (FD-300) with 77.2% sensitivity and 70.9% specificity; a third reported more modest performance for a different VD measure (area under the receiver operating characteristic (AUROC), 0.713).²
The review’s conclusion is measured: VD is the most promising OCTA-derived biomarker for early DR detection, but the current literature—limited in diagnostic accuracy studies and inconsistent in imaging protocols—is not yet sufficient to support OCTA’s routine clinical implementation for this purpose. The authors call for larger, standardized studies using consistent scan sizes (3 × 3 mm and 6 × 6 mm) and consistent evaluation of both the superficial and deep capillary plexuses.²
The backpack: Multi-disease screening on a smartphone
At the other end of the deployment spectrum, a team based at the National Institute of Ophthalmology in Pune, India, with a co-author also affiliated with Future Vision Eye Care in Mumbai, evaluated a multi-disease offline AI system integrated into a smartphone-based fundus camera. Their study, led by Aditya Kelkar and colleagues and published in the European Journal of Ophthalmology, tested Medios AI (MAI)—a system that had previously been validated separately for DR, glaucoma and AMD—in its newer, consolidated form, which screens for all three conditions using a single AI module rather than separate disease-specific algorithms.³
In this prospective cross-sectional study, 193 adults (371 eyes) with confirmed DR, glaucoma, AMD or normal fundus findings were enrolled between May and December 2024 at a tertiary eye care centre. Dilated fundus images were captured with both the Remidio Fundus on Phone (FoP)—a smartphone-mounted camera running the offline MAI algorithm on an iPhone 13—and a Zeiss Clarus 500 desktop camera, which served as the imaging benchmark. MAI-generated reports were compared against masked grading by two fellowship-trained ophthalmologists.³
For detecting any retinal disease, MAI achieved a sensitivity of 99.3% and a specificity of 95.7% (AUROC, 0.99). Performance by disease was as follows: glaucoma (n = 109), sensitivity 98.2% and specificity 99.0% (AUROC, 0.99); AMD (n = 56), sensitivity 88.9% and specificity 97.5% (AUROC, 0.93); and DR (n = 78), sensitivity 84.6% and specificity 99.0% (AUROC, 0.92), with sensitivity rising to 95% for referrable DR specifically. Agreement between MAI and human graders on vertical cup:disc ratio was strong, with Bland-Altman agreement ranging from −0.1 to 0.1; the two human graders showed excellent agreement with each other on the same measure, with an intergrader intraclass correlation coefficient of 0.97. The system generates an annotated report in under 15 seconds and functions entirely offline, without relying on internet connectivity or cloud infrastructure.³
The authors note that DR sensitivity in this multi-disease configuration (84.6%) was somewhat lower than sensitivities reported for earlier, single-disease Medios models built specifically for DR (previously, 93%-98.8%), which they suggest may reflect trade-offs introduced by having the system simultaneously process signatures for three different diseases. They also flag limitations, including a relatively small AMD subgroup, a single-centre design and a high-prevalence quaternary referral population that likely inflates positive predictive value relative to what would be seen in general community screening. They call for multicentre studies in lower-prevalence, community-based settings to confirm the system’s real-world scalability.³
In an exclusive quote for OTE, Sabyasachi Sengupta, DO, DNB, FMRF, FRCS, FICO, consultant vitreoretinal surgeon at Future Vision Eye Care and Research Centre in Mumbai, India, and corresponding author on the Medios-AI study, was asked about the trade-off in disease-specific sensitivity, the challenges of validating Medios-AI in lower-prevalence community settings and the broader role of offline AI in expanding access. Sengupta replied:
“Some reduction in disease-specific sensitivity is an expected trade-off when an AI system moves from detecting a single condition to distinguishing among several diseases, and further iterations will undoubtedly improve this. However, an important nuance in our study was that some eyes with DR—particularly those with prominent, confluent hard exudates—were classified as ‘DR or AMD’ rather than being considered normal; thus, the pathology was detected, even if the precise diagnostic label was uncertain, pulling down the accuracy of DR. In a screening setting, the patient would still be referred for specialist evaluation, which is ultimately the primary objective of screening.
“Our next step must be community-based validation, where the principal challenge may be image acquisition rather than AI interpretation—small pupils, well-lit screening environments and cataract-related media haze could substantially increase ungradable-image rates. Nearly one-third (prior experience, not validated numbers) may require brief pharmacological dilation, so any real-world study must assess not only diagnostic accuracy, but also gradability and its impact on screening logistics. This is especially true in regions of the world with dark colored iris, but thats where the real disease burden lies.
“I think retinal screening can be substantially offloaded to AI. With very high sensitivity, some over-referral is an acceptable trade-off if it means physicians no longer need to screen millions of normal eyes. Offline systems are particularly powerful: a locally trained operator can acquire images, while AI provides real-time quality checks, prompts retakes or mydriasis when needed, and identifies patients requiring referral. This could bring retinal screening deep into underserved regions of Africa, South America, Asia and Eastern Europe. The technology is largely here; the next challenge is last-mile delivery. Indeed, the Gates Foundation is also backing this vision and is funding Remidio's activities with a recent grant, I believe.”
Where this leaves retinal AI
Taken together, the findings from these three studies show how retinal AI is advancing across different stages of clinical care. FOCUS demonstrates the potential of a foundation model to automate a complex, multistep 3D OCT workflow in tertiary centres, achieving specialist-level accuracy while operating substantially faster than human graders. The OCTA review highlights an important limitation: the imaging biomarkers that underpin these systems still require stronger validation. VD appears to be a consistent early marker of DR, whereas other commonly used OCTA measures, particularly FAZ parameters, have produced less consistent results across studies. Meanwhile, the MAI study shows that an offline, smartphone-based platform can provide accurate multi-disease screening in a real-world clinical setting, supporting the potential for scalable, point-of-care screening in resource-limited settings.
Despite their different applications, the three studies point to a common next step: broader, multicentre validation. More diverse patient populations and clinical settings will be needed to determine how well these AI tools and imaging biomarkers perform beyond the populations and environments in which they were developed.
References
Zhang J, Zhong J, Lin L, et al. Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases. npj Digit Med. Published online August 19, 2026. doi:10.1038/s41746-026-03151-x
Politi LA, Pitoulias AG, Pitoulias MG, Topouzis F. Diagnostic potential of OCTA for early detection of diabetic retinopathy: a literature review. Eur J Ophthalmol. Published online August 8, 2026. doi:10.1177/11206721261476180
Kelkar A, Kelkar J, Garg Y, Jain HH, Sengupta S. Smartphone-based offline AI for multi-disease retinal screening: real-world accuracy. Eur J Ophthalmol. Published online June 24, 2026. doi:10.1177/11206721261462321
















