NHS England Data Science PhD Internships

Fact-Level Privacy Leakage Detection and Mitigation

Keywords: Privacy, Memorisation, Text

Need: Foundation models like LLMs show growing potential in healthcare, but pose novel privacy risks. Prior work by NHS England has shown that:

This project proposes to develop and evaluate more granular methods for detecting privacy leakage from NHS datasets, with a focus on fact-level and context-sensitive exposures.

This project will look to:

Current Knowledge/Examples & Possible Techniques/Approaches:

Related Previous Internship Projects: privfp-experiments; P71 - Investigating Privacy Concerns and Mitigations for Healthcare Language and Foundation Models; P51 - Investigating Privacy Concerns and Mitigations for Language Models in Healthcare

Enables Future Work: A fact-based leakage benchmark could become the basis for national LLM governance testing; Tools from this work could be extended into privacy evaluation modules for NHS LLM deployments; Contributes to Trustworthy AI development pipelines within the NHS

Outcome/Learning Objectives:

Datasets: MIMIC-III/IV and Synthetic Clinical Notes (e.g., from privfp-experiments), Internally generated synthetic clinical notes

Desired skill set: When applying please highlight any experience around privacy in large language models applied to healthcare, privacy preserving ML, coding experience (including any coding in the open), any other data science experience you feel relevant.


Return to list of all available projects.