Topic
AI evaluation & observability
Measuring what agents and LLM systems actually do: evaluation frameworks, tracing, cost and quality feedback loops.
Case studies
Can you prove a retrieval-augmented system is telling the truth?
An evaluation pipeline across retrieval, chunking, embedding and generation, enforced as release gates, that cut production hallucination incidents by 40%.
What happens when quality engineering is done by a team of AI agents?
Building and deploying agentic AI systems for QA, with testing, data-validation, bug-logging and notification agents, observable end to end.
How do you evaluate an agent that takes a different path every time it runs?
A reference architecture for scoring agent answers, tool calls and workflows, tracing every step with OpenTelemetry, and gating releases on evidence rather than demos.
Articles
Beyond accuracy: evaluating agentic AI layer by layer
A practical framework for evaluating agents per layer, choosing metrics and thresholds, calibrating LLM judges against humans, and catching regressions across prompt and model changes.
Testing tool-calling reliability in LLM applications
How to test whether an LLM picks the right tool, passes the right arguments, calls tools in a safe order and recovers from failure, with a pytest-style harness you can adapt.
Evals are the new tests
Why a demo is not evidence, and what it takes to treat AI evaluation as a release gate rather than a research exercise.
Work with me
Working on a problem like the ones above? I am glad to compare notes.