Researchers have created a pioneering platform that enables commercial artificial intelligence systems to be tested head-to-head under real clinical conditions, marking the first time such a large-scale evaluation has been possible for NHS use. The platform is designed to assess whether AI tools detect disease in a fair, equitable, transparent and trustworthy manner, with diabetic eye disease chosen as the initial exemplar. By operating independently from industry, the system removes the risk of commercial bias and ensures every participating company is judged on identical, rigorously controlled criteria.
Although NHS AI procurement currently focuses on cost-effectiveness and human-level accuracy, deeper challenges persist. There is a pressing need for stronger digital infrastructure and far more robust testing of commercial algorithms before they reach patients. Crucially, AI tools used as medical devices have rarely been assessed for fairness across large and diverse populations. The consequences of neglecting these issues are already known from devices such as pulse oximeters, which have been shown to perform less accurately on people with darker skin, prompting government reviews of equity in medical technologies.
The new platform, described in The Lancet Digital Health, was developed by researchers led by Professor Alicja Rudnicka at City St George’s, University of London, and Adnan Tufail at Moorfields Eye Hospital NHS Foundation Trust, in partnership with Kingston University and Homerton Healthcare NHS Trust. It was used to test commercial diabetic eye-detection algorithms on 1.2 million retinal images drawn from one of England’s largest and most diverse diabetic screening programmes. With over three million people screened every one to two years and roughly 18 million images generated annually, the current dependence on multiple human graders is increasingly unsustainable.
Twenty-five companies with CE-marked algorithms were invited to participate, and eight agreed to submit their systems. These were integrated into a secure ‘trusted research environment’ where they could analyse images without ever accessing patient data or the human grading results. Their performance was directly compared with that of human graders using established NHS protocols. The algorithms processed each patient’s images within milliseconds to seconds, far faster than the twenty minutes often required for manual assessment.
Accuracy for detecting cases potentially requiring clinical intervention ranged from 83.7% to 98.7%. For moderate-to-severe and proliferative diabetic eye disease, accuracy ranged from 95.8% to 99.8%, matching or surpassing human performance reported in earlier studies. The platform also measured how often healthy eyes were incorrectly flagged as diseased —an additional vital safety indicator —and confirmed that accuracy was consistent across ethnic groups—the first large-scale demonstration of fairness in diabetic eye-screening AI.
The researchers envision scaling the platform nationally, creating a centralised AI infrastructure where approved algorithms could analyse retinal images sent from screening centres, with results returned directly to electronic health records. Such an approach would reduce duplication, lower costs and ensure consistent, equitable service. It would allow clinicians to focus on higher-risk cases while giving companies independent feedback to refine their technologies. Ultimately, the model offers a blueprint for evaluating AI across other chronic conditions, supporting safer, fairer and more transparent adoption of AI throughout healthcare.
More information: Alicja Rudnicka et al, Automated retinal image analysis systems to triage for grading of diabetic retinopathy: a large-scale, open-label, national screening programme in England, The Lancet Digital Health. DOI: 10.1016/j.landig.2025.100914
Journal information: The Lancet Digital Health Provided by City St George’s, University of London
