<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Satya Prakash Solanki · Insights</title><description>Articles and essays on AI evaluation, AI security, reliability, governance, application security, data quality and performance engineering.</description><link>https://iamsatya.com/</link><item><title>Beyond accuracy: evaluating agentic AI layer by layer</title><link>https://iamsatya.com/insights/agentic-ai-evaluation-framework/</link><guid isPermaLink="true">https://iamsatya.com/insights/agentic-ai-evaluation-framework/</guid><description>A practical framework for evaluating agents per layer, choosing metrics and thresholds, calibrating LLM judges against humans, and catching regressions across prompt and model changes.</description><pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate></item><item><title>Turning a query engine benchmark into a performance regression test</title><link>https://iamsatya.com/insights/query-engine-performance-regression/</link><guid isPermaLink="true">https://iamsatya.com/insights/query-engine-performance-regression/</guid><description>Baselines, noise, run counts, statistical comparison and per-query thresholds for catching Trino performance regressions in CI.</description><pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate></item><item><title>Reading Trino query plans to find bottlenecks in TPC-H derived queries</title><link>https://iamsatya.com/insights/tpch-trino-query-plans/</link><guid isPermaLink="true">https://iamsatya.com/insights/tpch-trino-query-plans/</guid><description>How to read EXPLAIN and EXPLAIN ANALYZE in Trino, and what fragments, exchanges, join distribution, dynamic filters and statistics tell you.</description><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Testing indirect prompt injection in RAG and agent pipelines</title><link>https://iamsatya.com/insights/indirect-prompt-injection-testing/</link><guid isPermaLink="true">https://iamsatya.com/insights/indirect-prompt-injection-testing/</guid><description>Where untrusted content enters an LLM pipeline, how to design planted-instruction tests with canaries, how to measure attack success, and how to verify defences actually hold.</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Validating CDC pipelines</title><link>https://iamsatya.com/insights/validating-cdc-pipelines/</link><guid isPermaLink="true">https://iamsatya.com/insights/validating-cdc-pipelines/</guid><description>How to test incremental and change-data-capture pipelines for ordering, late and duplicate events, deletes, idempotent MERGE, watermarks and reconciliation windows.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Agent failure handling patterns</title><link>https://iamsatya.com/insights/agent-failure-handling-patterns/</link><guid isPermaLink="true">https://iamsatya.com/insights/agent-failure-handling-patterns/</guid><description>Retries, timeouts, fallbacks and recovery for AI agents, and the cases where a retry makes things worse.</description><pubDate>Sat, 26 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Testing tool-calling reliability in LLM applications</title><link>https://iamsatya.com/insights/tool-calling-reliability-testing/</link><guid isPermaLink="true">https://iamsatya.com/insights/tool-calling-reliability-testing/</guid><description>How to test whether an LLM picks the right tool, passes the right arguments, calls tools in a safe order and recovers from failure, with a pytest-style harness you can adapt.</description><pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Making security findings comparable across tools</title><link>https://iamsatya.com/insights/security-findings-normalisation/</link><guid isPermaLink="true">https://iamsatya.com/insights/security-findings-normalisation/</guid><description>How to turn SAST, SCA, DAST and container scan output into one schema, with stable fingerprints, CWE and CVE mapping, normalised severity and expiring suppressions.</description><pubDate>Tue, 22 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Evals are the new tests</title><link>https://iamsatya.com/insights/evals-are-the-new-tests/</link><guid isPermaLink="true">https://iamsatya.com/insights/evals-are-the-new-tests/</guid><description>Why a demo is not evidence, and what it takes to treat AI evaluation as a release gate rather than a research exercise.</description><pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Security testing AI agents: permissions and trust boundaries</title><link>https://iamsatya.com/insights/agent-permissions-trust-boundaries/</link><guid isPermaLink="true">https://iamsatya.com/insights/agent-permissions-trust-boundaries/</guid><description>How to test what an AI agent is allowed to do, on whose behalf, and with whose data. Least privilege, confused deputies, MCP trust and approval gates, with a test matrix.</description><pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate></item><item><title>From expectations to a production data-quality system</title><link>https://iamsatya.com/insights/data-quality-rules-and-monitoring/</link><guid isPermaLink="true">https://iamsatya.com/insights/data-quality-rules-and-monitoring/</guid><description>Rule catalogues, severity and ownership, where checks run, simple anomaly baselines for volume, freshness and distribution, and alerting that people still read.</description><pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Measuring AI reliability</title><link>https://iamsatya.com/insights/measuring-ai-reliability/</link><guid isPermaLink="true">https://iamsatya.com/insights/measuring-ai-reliability/</guid><description>Why a single pass proves little for non-deterministic systems, and how to measure consistency, set SLOs and error budgets, and choose thresholds you can defend.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>From detection to verified remediation</title><link>https://iamsatya.com/insights/verified-remediation-loop/</link><guid isPermaLink="true">https://iamsatya.com/insights/verified-remediation-loop/</guid><description>Closing the loop on vulnerabilities with ownership routing, severity-based SLAs, targeted re-scans, regression tests and metrics that leadership and auditors can trust.</description><pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Governance is how you go faster</title><link>https://iamsatya.com/insights/governance-is-how-you-go-faster/</link><guid isPermaLink="true">https://iamsatya.com/insights/governance-is-how-you-go-faster/</guid><description>Risk tiers, model inventories and evidence trails are not bureaucracy. Done well, they are what lets an organisation say yes to AI.</description><pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate></item></channel></rss>