Skip to content
Satya Prakash Solanki

Unified application security assurance pipeline

Overview

Most organisations do not lack security tools. They lack a single, defensible answer to the question “is this release safe enough to ship?”. Each scanner produces its own report, in its own format, with its own idea of what “critical” means. The result is a pile of findings rather than a decision.

This reference architecture describes how I would integrate static analysis (SAST), software composition analysis (SCA), container image scanning, infrastructure-as-code (IaC) scanning, dynamic testing (DAST), API security testing and penetration testing into one assurance pipeline. It covers orchestration, multi-tenant isolation, normalisation of findings into a common schema, deduplication, risk-based prioritisation, SBOM generation, release gates written as policy, compliance reporting and verification of fixes.

It draws on my practice running a full application security testing programme (SAST, SCA, DAST, container, IaC, API and penetration testing, aligned to the OWASP Top 10 and NIST guidance) with tools such as Semgrep, Checkmarx, OWASP ZAP, Burp Suite, Nuclei, OpenVAS, Metasploit and nmap, and on designing the multi-tenant security architecture of a security platform (Keycloak identity, tenant isolation, source-control integration). It is a reference design, not a report of a specific deployment. All thresholds below are example values.

Problem statement

A typical secure SDLC grows one tool at a time. SAST arrives first, then dependency scanning, then a container scanner when the platform moves to Kubernetes, then DAST and an annual penetration test. Each addition is reasonable. Together they create four problems.

  1. No common language. A Semgrep rule, a Checkmarx query, a CVE in a transitive dependency and a ZAP alert describe different things in different shapes. Nobody can answer “how many open high-risk issues do we have?” without a spreadsheet.
  2. Duplicates and noise. The same vulnerable library shows up in the SCA report, in the container scan of every image that includes it, and sometimes as a DAST banner finding. Developers see three tickets for one fix and stop trusting the queue.
  3. Severity is not risk. A CVSS 9.8 in a library function the application never calls is less urgent than a CVSS 7.5 that is reachable from the internet and on the CISA Known Exploited Vulnerabilities list. Tool severities alone cannot express that.
  4. No closed loop. Tickets are closed when a developer says “fixed”, not when a scan proves it. Auditors ask for evidence and receive screenshots.

When the platform serves several teams or customers, a fifth problem appears: findings, source code and credentials from one tenant must never be visible to another.

Engineering objectives

The design is held to a short set of objectives, each of which can be tested.

  • One schema. Every finding, from every tool, is stored in a single normalised form with a stable fingerprint.
  • One decision per release. A release gate evaluates versioned policy against normalised findings and returns pass, warn or fail with reasons.
  • Risk-based ordering. Priority combines exploitability, exposure and asset criticality, not tool severity alone.
  • Verified closure. A finding is closed only when a targeted re-scan or re-test no longer reproduces it.
  • Strict tenant isolation. Identity, data and scan execution are isolated per tenant, and isolation is tested, not assumed.
  • Audit-ready evidence. Findings, decisions, exceptions and SBOMs are retained and mapped to the OWASP Top 10, OWASP ASVS and the NIST Secure Software Development Framework (SSDF).

Solution architecture

The pipeline has five logical layers. Scanners are treated as replaceable plug-ins. The value sits in the layers around them.

01 Trigger

  • CI pipeline events
  • Scheduled scans
  • Source-control webhooks
  • Manual pen test intake

02 Scan

  • SAST: Semgrep, Checkmarx
  • SCA and SBOM
  • Container and IaC scanners
  • DAST and API: ZAP, Burp, Nuclei

03 Normalise

  • SARIF ingestion
  • Common finding schema
  • Fingerprint and dedup
  • CWE and CVE mapping

04 Decide

  • Risk scoring: CVSS, EPSS, reachability
  • Policy-as-code gate
  • Exceptions with expiry

05 Assure

  • Ownership routing and SLAs
  • Targeted re-scan
  • Compliance mapping and reports
Figure 1. Logical architecture of the unified assurance pipeline. Scanners are interchangeable; normalisation and policy are the stable core.

Orchestration

Each repository has a scan profile: which scanners run, on which paths, at which depth, and with which credentials. Profiles live in the repository alongside the code, so they are reviewed like any other change. Fast scanners (Semgrep, SCA, IaC, image scanning) run on every pull request. Slower ones (full Checkmarx scans, authenticated DAST, Nuclei templates against staging) run on merge to the main branch and on a schedule, because new CVEs appear against code that has not changed.

Penetration testing does not fit a pipeline, but its output should. Manual findings from Burp Suite, Metasploit-assisted verification or OpenVAS infrastructure scans enter through the same intake as automated ones, so they share a schema, an SLA and a verification step.

Multi-tenant isolation

For a shared platform, isolation is designed in at three levels.

  • Identity. Keycloak provides one realm (or one organisation within a realm) per tenant, with roles scoped to that tenant. Every API call carries a token whose tenant claim is enforced at the service layer, not only in the user interface.
  • Data. Every finding, SBOM and report row carries a tenant identifier, and data access is filtered on it by default. Row-level security in the database is a second line of defence behind application checks.
  • Execution. Scanner jobs run in short-lived, per-tenant workloads with their own network policy and their own secrets. Source-control integration uses per-tenant app installations or tokens with the narrowest scopes needed to clone and comment, never a shared service account.

Technical approach

The flow

  1. 01Commit or schedule
  2. 02Scan
  3. 03Normalise
  4. 04Triage and score
  5. 05Gate
  6. 06Verify fix
Figure 2. Lifecycle of a finding. Normalisation is the step that makes every later step possible.

Normalisation with SARIF as the interchange

SARIF 2.1.0, the OASIS Static Analysis Results Interchange Format, is the natural interchange. Semgrep, many SCA tools and most container scanners can emit it directly. For tools that cannot, such as ZAP or Burp exports, a small adapter converts their JSON or XML into SARIF first. SARIF is then mapped into an internal schema that adds what SARIF does not carry well: tenant, asset, ownership, risk score and lifecycle state.

A normalised finding looks like this:

finding.json
{
"id": "f_01J9Q4K7Z3",
"tenant_id": "t_acme",
"fingerprint": "sha256:9f2c1e...b71a",
"category": "sca",
"source_tools": ["dependency-scanner", "image-scanner"],
"title": "Prototype pollution in transitive dependency",
"cwe": ["CWE-1321"],
"cve": ["CVE-2024-00000"],
"asset": {
"repo": "payments-api",
"component": "pkg:npm/example-lib@4.17.20",
"image": "registry.local/payments-api:1.42.0",
"criticality": "high",
"internet_exposed": true
},
"severity": {
"tool": "critical",
"cvss_base": 9.1,
"epss": 0.12,
"kev": false,
"reachable": "unknown",
"priority": "P2"
},
"status": "open",
"owner": "team-payments",
"first_seen": "2026-09-14T08:12:00Z",
"sla_due": "2026-10-14T08:12:00Z",
"evidence": { "sarif_ref": "scans/8831/results.sarif#/runs/0/results/17" }
}

The CVE identifier above is a placeholder. The important fields are the fingerprint, which drives deduplication, and the severity block, which keeps the raw tool severity alongside the inputs used to compute priority, so the decision can always be explained.

Deduplication and correlation

Fingerprints are computed per category, because “the same issue” means different things to different tools.

Category Fingerprint inputs
SAST rule family, CWE, file path, enclosing function, normalised code snippet
SCA package URL (name and version), vulnerability identifier
Container base image layer digest, package URL, vulnerability identifier
IaC resource address, policy identifier
DAST and API normalised endpoint template, HTTP method, parameter, CWE

Line numbers are deliberately excluded from SAST fingerprints, because an unrelated edit above the finding would otherwise create a “new” issue. SARIF supports this through its partialFingerprints property, and I prefer to use the tool’s own fingerprint when it provides a stable one.

Correlation is a separate step from deduplication. A vulnerable library found by SCA and by the image scanner is merged into one finding with two source_tools. A SAST injection finding and a DAST finding on the same endpoint with the same CWE are linked, not merged, because they carry different evidence. That link is valuable: a static finding confirmed dynamically is strong evidence of real exposure and moves up the queue.

Severity and prioritisation

Priority combines three inputs. The CVSS base score describes technical severity. EPSS estimates the probability of exploitation activity in the near term. Context, meaning reachability, internet exposure and asset criticality, says whether this system and this code path matter. Presence on the CISA KEV catalogue overrides the rest.

Exploit signal Reachable and exposed Reachable, internal only Not reachable or unknown
On KEV, or confirmed by DAST or pen test P1 P1 P2
EPSS high (example: 0.1 or above) P1 P2 P3
CVSS 7.0 or above, low EPSS P2 P3 P3
CVSS below 7.0, low EPSS P3 P4 P4

The EPSS cut-off is an example value and should be tuned to the size of the backlog a team can realistically handle. “Unknown” reachability is treated conservatively for internet-facing assets. Reachability analysis is imperfect, and claiming “not reachable” without evidence is how real exposures get deprioritised.

SBOM generation

Every build produces an SBOM in CycloneDX or SPDX format, attached to the build artefact and stored with the release record. The SBOM serves two purposes. It feeds SCA so that dependency analysis runs against what was actually built, not only against the lockfile. And it allows retrospective queries: when a new CVE is published, the question “which released versions include this component?” becomes a lookup rather than an investigation.

Release gates as policy-as-code

Gate rules are versioned, reviewed and tested like code. The engine can be Open Policy Agent or a small in-house evaluator. What matters is that policy is declarative and separate from pipeline scripts.

release-gate.yaml
policy: release-gate
version: 3
applies_to:
environments: [production]
rules:
- id: no-open-p1
description: No open P1 findings on the release candidate
when: { priority: P1, status: open }
action: fail
- id: p2-within-sla
description: P2 findings must be inside their SLA window
when: { priority: P2, status: open, sla_breached: true }
action: fail
- id: secrets
description: Any detected secret blocks release
when: { category: secret, status: open }
action: fail
- id: sbom-present
description: Release must carry a CycloneDX or SPDX SBOM
require: { artefact: sbom }
action: fail
- id: new-p3
description: New P3 findings introduced by this change
when: { priority: P3, introduced_in: this_change }
action: warn
exceptions:
require: [approver_role: security-lead, reason, expires_at]
max_duration_days: 30

Two details make gates workable. First, gates distinguish findings introduced by this change from inherited backlog, so a team is not blocked by debt it did not create, while the debt stays visible. Second, every exception has an approver, a reason and an expiry date. An exception without an expiry is a permanent policy change made without review.

Reporting and compliance mapping

Because every finding carries a CWE, mapping to frameworks is a lookup table rather than manual work. CWE categories map to the OWASP Top 10. Verification activities map to requirements in the OWASP ASVS. Pipeline practices (SBOMs, automated testing, vulnerability response) map to practices in the NIST SSDF (SP 800-218). The same data then produces a developer view (my open findings, by priority), a leadership view (risk trend and SLA adherence) and an auditor view (evidence per control).

Remediation verification

A finding moves to “fixed” only when a targeted re-scan, scoped to the affected repository, image or endpoint, no longer reproduces the fingerprint. For DAST and pen test findings, the original request is replayed against the fixed build. The detail of this loop, including SLAs and regression tests, is covered in From detection to verified remediation.

Validation strategy

A security pipeline needs its own tests. I validate this design in four ways.

  • Seeded vulnerabilities. Deliberately vulnerable test applications and known-bad dependencies run through the pipeline on a schedule. Each seeded issue must be detected, normalised with the expected CWE, prioritised as expected and blocked by the gate. A missed seed is a pipeline defect.
  • Dedup tests. Fixture SARIF files from different tools describing the same issue must collapse to one finding. Fixtures with a trivial code move must keep the same fingerprint.
  • Policy unit tests. Each gate rule has test cases for pass, warn, fail and expired exception, run in CI whenever policy changes.
  • Tenant isolation tests. Automated tests attempt cross-tenant reads through the API with valid tokens from a different tenant, attempt to access another tenant’s scan workspace, and confirm that source-control tokens cannot reach repositories outside their installation. These run on every platform release, not once.

Metrics and measurement

The pipeline is measured on whether it produces good decisions, not on how many findings it generates. Useful measures:

  • Scan coverage: share of active repositories and deployed images scanned within the expected window.
  • Duplicate ratio: raw findings ingested versus unique normalised findings. A falling ratio after a tool change signals a fingerprinting regression.
  • Triage precision: share of triaged findings confirmed as true positives, per tool and per rule. This tells you which rules to tune or retire.
  • Mean time to remediate by priority, and SLA adherence.
  • Reopen rate: share of verified fixes that reappear. A non-zero rate is normal; a rising one points to fixes that treat symptoms.
  • Gate outcomes: pass, warn and fail counts per release, plus active exceptions and their age.

I do not quote target values here because they depend on the organisation’s risk appetite and backlog. What matters is that each metric has an owner and a trend.

Challenges and trade-offs

Speed versus depth. Full SAST and authenticated DAST are slow. Running everything on every pull request makes developers route around the pipeline. Splitting scans into fast pull-request checks and deeper scheduled runs is a trade-off: some issues are found hours later rather than minutes. I accept that, as long as the gate on the release branch sees the deep results.

Schema ambition. A normalised schema tempts teams to model every tool’s fields. That leads to a schema nobody can maintain. I keep the core small and push tool-specific detail into a linked raw evidence record.

Reachability claims. Reachability analysis reduces noise considerably but can be wrong in both directions, particularly with reflection, dynamic imports and framework magic. I treat “not reachable” as a reason to lower priority, never as a reason to suppress.

False positives and suppression. Every suppression mechanism is also a way to hide real issues. Suppressions are scoped to a fingerprint, carry a reason and an expiry, and are reported as a metric in their own right.

Tenant isolation cost. Per-tenant workloads and credentials cost more than shared runners. For a security platform, that cost is the product. A shared runner that can read another tenant’s source code is a breach waiting for a trigger.

Tool overlap. Two SAST engines find different things, and also many of the same things. Running both raises coverage and noise together. Correlation makes overlap useful instead of costly, but only if fingerprints are good.

Outcome and lessons learned

As a reference design, the outcome is a pipeline in which every finding, regardless of source, has one shape, one owner, one priority and one lifecycle, and in which a release decision can be explained by pointing at a versioned policy and a set of evidence.

The lessons I would carry into any implementation:

  • Normalisation is the foundation. Without it, every later capability is a spreadsheet.
  • Severity is an input, not an answer. Exploit likelihood and context change the order of work more than any tool setting.
  • Gates must separate new issues from inherited debt, or teams will disable them.
  • “Fixed” is a claim until a re-scan confirms it.
  • In a multi-tenant platform, isolation belongs in identity, data and execution, and it must be tested continuously.

Further reading

Try “evaluation”, “red-teaming”, “governance” or “agents”.