Approach
Assurance as a system, not a checklist.
Trustworthy AI comes from disciplines that reinforce each other: assurance produces evidence, security attacks it, reliability keeps it true in production, and governance and compliance record it.
What I believe
Four convictions behind the work.
Evals are the new tests.
A convincing demo is not evidence. If you cannot measure an AI system’s quality, you are not ready to ship it.
Security cannot be bolted on.
Red-teaming and guardrails belong in the design review, not the incident review.
Governance is how you go faster.
Clear risk tiers and real evidence let an organisation say yes to AI sooner, and with confidence.
AI needs a tester’s mindset.
The habits of good quality engineering, assuming failure and hunting the edges, are what trustworthy AI is missing.
The six disciplines
Six disciplines, one assurance system.
Security
Attack AI systems before adversaries do.
- Prompt-injection, jailbreak and exfiltration testing
- Risk categorisation mapped to OWASP LLM Top 10 and MITRE ATLAS
- SAST, SCA and DAST in a secure SDLC
Reliability
Agents that behave the same way on the thousandth run.
- Failure-mode analysis and fallback design
- Agent SLOs, retries and determinism testing
- Drift and regression detection
Assurance
Measured results decide what ships.
- Evaluation pipelines as release gates
- Golden datasets and LLM-as-a-judge
- Groundedness, faithfulness, hallucination scoring
Governance
Know which models are in use, how risky each is, and what evidence backs it.
- Use-case risk tiering
- Model inventory and model cards
- Audit trails aligned to NIST AI RMF
Compliance
Controls and evidence mapped to OWASP, NIST, MITRE and UAE frameworks.
- OWASP LLM Top 10 and NIST AI RMF mapping
- UAE IA / NESA and ADHICS alignment
- Evaluation evidence and audit trails
Trust
Behaviour you can trace, explain and rely on.
- Guardrails for PII, topic scope and grounding
- Agent tracing with Langfuse, Arize and OpenTelemetry
- Token, cost and latency telemetry; human-in-the-loop review
How an engagement runs
From first assessment to ongoing evidence.
Assess
Map the AI use cases, agents and data flows. Tier each use case by risk so effort goes where impact is highest.
Evaluate
Golden datasets, metrics and thresholds agreed up front, then automated as release gates in the pipeline.
Red-team
Adversarial testing for prompt injection, jailbreaks, data exfiltration, tool and RAG poisoning, and excessive agency.
Guard
Runtime controls at three layers: AI Gateways for access and cost, AI Shield for injection and tool-call abuse, and AI Guardrails for PII, topic scope and grounding, tuned against live traffic.
Observe
Agent and span-level tracing, cost and latency telemetry, drift detection and SLOs for production agents.
Govern
Model inventory, risk registers and audit trails that reuse the evidence the earlier steps produce.
Frameworks
Mapped to recognised frameworks.
Risks, controls and evidence are mapped to these frameworks. Alignment is not a certification.
- LLM Top 10
- AI RMF
- ATLAS
- Top 10
- IA / NESA
- ADHICS
Working together
Ways we can work together.
Assurance advisory
Strategy and roadmap for evaluating, securing and governing AI across a portfolio.
EnquireAgentic AI architecture
Design or review of agentic systems: orchestration, AI Gateways, AI Shield, AI Guardrails and observability.
EnquireAI red-team assessment
Scoped adversarial testing of an LLM or agentic application, mapped to OWASP LLM Top 10 and MITRE ATLAS.
EnquireTalks and workshops
Practical sessions for leadership and engineering teams on evaluation, red-teaming and governance.
Enquire
Work with me
Not sure which fits? Start with a short conversation about what you are building.