Insights
Notes from the engineering floor.
How I evaluate, secure and engineer AI and data systems. Original pieces live here; earlier articles on Medium and Towards Dev link to their original home.
Series
Series · 4 parts
TPC-H benchmarking with Trino
Series · 5 parts
Big data testing and data quality
Latest articles
14 original pieces
Beyond accuracy: evaluating agentic AI layer by layer
A practical framework for evaluating agents per layer, choosing metrics and thresholds, calibrating LLM judges against humans, and catching regressions across prompt and model changes.
Turning a query engine benchmark into a performance regression test
Baselines, noise, run counts, statistical comparison and per-query thresholds for catching Trino performance regressions in CI.
Reading Trino query plans to find bottlenecks in TPC-H derived queries
How to read EXPLAIN and EXPLAIN ANALYZE in Trino, and what fragments, exchanges, join distribution, dynamic filters and statistics tell you.
Testing indirect prompt injection in RAG and agent pipelines
Where untrusted content enters an LLM pipeline, how to design planted-instruction tests with canaries, how to measure attack success, and how to verify defences actually hold.
Validating CDC pipelines
How to test incremental and change-data-capture pipelines for ordering, late and duplicate events, deletes, idempotent MERGE, watermarks and reconciliation windows.
Agent failure handling patterns
Retries, timeouts, fallbacks and recovery for AI agents, and the cases where a retry makes things worse.
Testing tool-calling reliability in LLM applications
How to test whether an LLM picks the right tool, passes the right arguments, calls tools in a safe order and recovers from failure, with a pytest-style harness you can adapt.
Making security findings comparable across tools
How to turn SAST, SCA, DAST and container scan output into one schema, with stable fingerprints, CWE and CVE mapping, normalised severity and expiring suppressions.
Evals are the new tests
Why a demo is not evidence, and what it takes to treat AI evaluation as a release gate rather than a research exercise.
Security testing AI agents: permissions and trust boundaries
How to test what an AI agent is allowed to do, on whose behalf, and with whose data. Least privilege, confused deputies, MCP trust and approval gates, with a test matrix.
From expectations to a production data-quality system
Rule catalogues, severity and ownership, where checks run, simple anomaly baselines for volume, freshness and distribution, and alerting that people still read.
Measuring AI reliability
Why a single pass proves little for non-deterministic systems, and how to measure consistency, set SLOs and error budgets, and choose thresholds you can defend.
From detection to verified remediation
Closing the loop on vulnerabilities with ownership routing, severity-based SLAs, targeted re-scans, regression tests and metrics that leadership and auditors can trust.
Governance is how you go faster
Risk tiers, model inventories and evidence trails are not bureaucracy. Done well, they are what lets an organisation say yes to AI.