Case studies
How I approach engineering problems.
Each case study starts from the question a team actually faces, then works through architecture, implementation, validation and trade-offs.
- Anonymised work from my roles. Product and client details are confidential.
- A reference architecture and implementation approach, not a specific client engagement.
- A detailed technical method, written from hands-on practice. Any numbers shown are illustrative.
From practice
Anonymised work from my roles
From practiceReliability
How do you know an AI agent is reliable once it is in production?
Architecting a reliability and observability platform that evaluates agent behaviour in real time and in batch, so failures surface before users find them.
2evaluation planes, real-time and batch, kept separate
From practiceEvaluation
Can you prove a retrieval-augmented system is telling the truth?
An evaluation pipeline across retrieval, chunking, embedding and generation, enforced as release gates, that cut production hallucination incidents by 40%.
40%fewer hallucination incidents in production
From practiceData Quality
What if most manual data checks could validate themselves?
An LLM-assisted validation framework for big data pipelines that automated more than 70% of manual ETL checks.
70%+of manual ETL checks automated
From practiceAI Security
How do you find out how an AI agent can be attacked, before someone else does?
An adversarial testing practice for LLM and agentic systems, with risks categorised against the OWASP LLM Top 10, NIST AI RMF and MITRE ATLAS, on top of a full application security programme.
Read the case studyFrom practiceGovernance
How can an organisation say yes to AI faster, and still stay in control?
An AI governance framework covering use-case risk tiering, model inventory, evaluation evidence and audit trails, aligned to NIST AI RMF and regional regulation.
Read the case studyFrom practiceEvaluation
What happens when quality engineering is done by a team of AI agents?
Building and deploying agentic AI systems for QA, with testing, data-validation, bug-logging and notification agents, observable end to end.
4specialised agents: testing, data validation, bug logging, notification
Reference architectures & deep dives
How I would design and validate it
How do you evaluate an agent that takes a different path every time it runs?
A reference architecture for scoring agent answers, tool calls and workflows, tracing every step with OpenTelemetry, and gating releases on evidence rather than demos.
How do you turn AI red-teaming from a one-off exercise into a repeatable engineering system?
A reference harness that generates adversarial tests, judges outcomes, scores severity against OWASP and MITRE ATLAS, and re-runs every confirmed finding as a CI regression test.
How do you make a multi-agent workflow reliable in a way you can measure, rather than hope for?
A reference architecture for engineering reliability into multi-agent systems, from failure-mode taxonomy and guarded tool calls to SLOs, error budgets and release gates.
How do you turn seven security scanners into one decision you can trust at release time?
A reference architecture that orchestrates SAST, SCA, container, IaC, DAST, API and pen testing, normalises findings into one schema and gates releases with policy-as-code.
How do you prove that every row that left the source arrived in the lakehouse, correctly transformed?
A reference design for reconciling and validating data across CDC, Kafka, PySpark and a Trino-queried lakehouse, from row counts to keyed diffs and anomaly detection.
How do you benchmark a query engine so the numbers survive a second run?
A reference harness for running a workload derived from TPC-H on Trino and Iceberg, with controlled environments, statistical reporting and CI regression gates.
Work with me
Facing a similar assurance, security or engineering problem?