Topic
AI reliability engineering
Failure modes, fallbacks, SLOs and release readiness for non-deterministic, autonomous systems.
Case studies
How do you know an AI agent is reliable once it is in production?
Architecting a reliability and observability platform that evaluates agent behaviour in real time and in batch, so failures surface before users find them.
How do you make a multi-agent workflow reliable in a way you can measure, rather than hope for?
A reference architecture for engineering reliability into multi-agent systems, from failure-mode taxonomy and guarded tool calls to SLOs, error budgets and release gates.
Articles
Agent failure handling patterns
Retries, timeouts, fallbacks and recovery for AI agents, and the cases where a retry makes things worse.
Measuring AI reliability
Why a single pass proves little for non-deterministic systems, and how to measure consistency, set SLOs and error budgets, and choose thresholds you can defend.
Work with me
Working on a problem like the ones above? I am glad to compare notes.