Why DICOM Anonymisation Is Critical for AI Training
Medical AI models rely on large volumes of images and data to reach accurate results, but using patient images in training and research requires protecting the personal information embedded within them. DICOM Anonymisation has become an essential step before sharing images or using them in AI model development. This article explains why DICOM Anonymisation is critical for AI training, starting with understanding what data exists inside imaging files and how to handle it correctly.
Why Every NHS AI Project Starts With a DICOM Anonymisation Problem
AI projects within the NHS require large numbers of medical images, but sharing those images requires removing any information that could reveal a patient’s identity. DICOM anonymisation for NHS AI training has become an important step before using data in research or model training.
NHS England has stated that 90% of AI tools remain stuck in pilot phases due to over-reliance on temporary IT setups in each trust, with each trust required to set up new databases and repeat testing from scratch even when a tool has already been validated elsewhere. Currently, the cost of building the infrastructure required for a multi-site study can reach £3.5 million per study. To address this, NHS England and the NIHR are funding a £6 million AI Research Screening Platform (AIR-SP), expected to launch for research purposes in 2027, which will reduce these costs by an estimated £2 to £3 million per multi-site study and create a shared national infrastructure for AI tool trials, underscoring how critical it is to prepare imaging data securely and at scale from the outset.
The Three Layers of Patient Data in DICOM — Why All Three Must Be Addressed
Patient data within a DICOM file exists across more than one layer, which is why deleting information from the header alone is not sufficient. The most important of these layers are:
- Standard DICOM Header: Contains information such as the patient’s name, date of birth, medical record number, examination date, Trust name, and radiologist, and this data can be handled using standard DICOM tools.
- Private Tags: Added by imaging device manufacturers, these may contain identifying or technical data, and they differ from one device to another.
- Burned-in Pixel PHI: This is information embedded within the image itself, such as the patient’s name, date of birth, or medical record number, and it requires different handling because it forms part of the pixel data.
Why Removing DICOM Headers Alone Is Not Enough for NHS AI Compliance
Some believe that removing the DICOM Header is sufficient to prepare an image for use in research, but patient information may still be visible within the image itself, particularly in ultrasound images, fluoroscopy recordings, and digital subtraction angiography (DSA) studies. The problem is that Burned-in PHI forms part of the image rather than a text field that can simply be deleted. Research on ultrasound image anonymisation has shown that removing metadata alone is not enough, and that data embedded within the pixel layer requires separate handling.
These cases require Pixel-Level DICOM Anonymisation using techniques such as optical character recognition (OCR) to detect and redact text within the image while preserving the diagnostically relevant content.
The Real Cost of Manual DICOM Anonymisation for AI Training Datasets
Some Trusts resort to reviewing images and removing patient data manually, but this approach becomes increasingly difficult as the number of studies grows. The most notable problems it produces are:
- Staff time consumed by reviewing and preparing images.
- A higher likelihood of human error leaving personal information behind within the files.
- The difficulty of preparing the thousands of studies required to train AI models.
- Rising operational costs as data volumes increase.
Processing studies manually can take a considerable amount of time when handling large image sets, which is why Bulk DICOM Anonymisation has become a solution that helps radiology departments and research centres process large numbers of studies more efficiently.
UK GDPR, NHS Secure Data Environments, and What They Require from Your Anonymisation Process
Patient data in the United Kingdom is subject to strict rules when used in research and analysis. The anonymisation process must therefore comply with UK GDPR requirements and NHS Secure Data Environment standards. Medical imaging anonymisation for research also helps prepare medical images before they are used in studies and research projects.
Secure data environments are built on a set of principles designed to protect patient data:
- Safe Data: Minimising data and removing information that could reveal a patient’s identity.
- Safe Settings: Using a secure environment that complies with data protection requirements.
- Safe People: Verifying the identities of those permitted to access the data.
- Safe Projects: Confirming that the research project meets the necessary approvals and requirements.
- Safe Outputs: Reviewing results before they leave the secure environment to ensure no personal information is disclosed.
How BriX Automates Pixel-Level DICOM Anonymisation at Scale for NHS AI Programmes
Radiology departments and research centres need tools capable of processing large numbers of images quickly. BriX from Rosenfield Health is a solution for anonymising medical images at scale, handling both DICOM metadata and pixel-level data. The tool offers a range of features that support the processing of large image sets:
- Processing 100 X-ray images in 4 minutes and 100 CT studies in 16 minutes.
- Integration with existing PACS systems and clinical trial servers.
- Audit files documenting each step of the anonymisation process.
- Support for multi-anonymisation profiles to apply different settings based on data type.
- Support for meso-anonymisation and batch processing for large-scale and scheduled operations.
- Processing of DICOM Headers and Burned-in PHI within a single workflow.
If you are working on a medical AI project within the NHS and need to prepare large volumes of DICOM images, a specialist solution such as BriX can help apply anonymisation at both the data and pixel levels, accelerating image processing while maintaining security and compliance requirements. Contact us today to learn how BriX can streamline your imaging workflow.
Conclusion
Protecting medical images starts with addressing every location where patient data exists within a DICOM file,not just deleting the header, as Private Tags and burned-in patient-identifiable data may also contain identifying information. The answer to why DICOM Anonymisation is critical for AI training lies in its direct connection to protecting patient privacy, preparing data for use in research, and developing AI models, all while maintaining compliance with NHS Medical Imaging GDPR Compliance in the UK.
FAQs
What is the difference between DICOM de-identification and anonymisation?
Anonymisation refers to removing personal information in a way that prevents a patient from being identified,while DICOM de-identification for AI datasets refers to removing certain direct identifying data with the possibility of re-linkage in some cases. Pseudonymisation replaces patient data with a code that can be linked back to the original data using a separate key, which is why the legal requirements and handling differ for each type.
What are the NHS Secure Data Environment requirements for medical imaging anonymisation?
NHS Secure Data Environments require the application of the Five Safes principles, which cover safe data, safe settings, safe people, safe projects, and safe outputs, alongside the protection of patient data before it is made available to researchers or used in analysis.
How many imaging studies does an NHS Trust need to validate an AI tool?
There is no fixed number of studies that applies to all AI tools, because the volume of data required varies according to the type of tool, the nature of the condition, and the level of accuracy needed. External validation also requires diverse data from more than one centre to assess model performance more effectively, particularly when preparing AI training datasets for Radiology.
Why is burned-in pixel PHI particularly dangerous in ultrasound images?
Burned-in PHI is embedded within the image itself, so it cannot be removed simply by deleting the DICOM Header data. Ultrasound images may contain the patient's name, date of birth, or medical record number visibly within the image, which is why they require Pixel-Level DICOM Anonymisation to detect and redact that data.