AI Safety
AI systems fail at their boundaries as often as inside their models — where they meet compromised data, unreliable tools and human decision-makers who must decide how much to trust them. My research treats AI safety as inseparable from cybersecurity, privacy and system assurance.
Why AI Safety Needs Cybersecurity
Many emerging AI safety failures occur at the system boundary rather than inside the model itself: compromised context, poisoned knowledge sources, unsafe tool use, excessive privilege, adversarial manipulation, unreliable outputs, poor provenance, human over-reliance and insufficient assurance evidence. Treating AI safety purely as a model-behaviour problem misses most of this attack surface.
Compromised context
Inputs and retrieved knowledge an AI system relies on can be manipulated upstream of the model.
Unsafe tool use and excessive privilege
Agentic systems that call tools or external services inherit the security posture of everything they touch.
Adversarial manipulation
Systems deployed in contested environments must withstand deliberate attempts to induce unsafe behaviour.
Poor provenance and weak evidence
Without traceable evidence for a system's outputs, safety claims are difficult to verify or audit.
Human over-reliance
Decision-makers who trust AI recommendations uncritically remove the oversight safety depends on.
Insufficient assurance
Deployment decisions are often made without systematic evidence that a system behaves safely under realistic conditions.
Current Research Questions
The questions driving my emerging AI safety and assurance programme.
When should humans trust AI recommendations — and when should the system defer instead?
How can AI confidence be independently verified, rather than self-reported?
How can autonomous AI systems be tested under adversarial conditions before deployment?
How can evidence provenance be preserved across an AI pipeline, from data to decision?
How should organisations evaluate AI systems before deploying them in high-stakes environments?
How do we secure agentic systems that interact with tools and external knowledge sources?