Ask a security team how many high-severity issues are open across their applications and watch what happens. Someone opens the SAST console. Someone else exports the dependency scanner. A third person searches the ticketing system for last quarter’s penetration test. An hour later there is a number, and nobody fully believes it.
The problem is not the tools. Semgrep, Checkmarx, OWASP ZAP and container scanners each do their job well. The problem is that their outputs are not comparable. This article describes how I make them comparable: a small common schema, fingerprints that survive code changes, consistent CWE and CVE mapping, a single severity model, and false-positive handling that cannot quietly become permanent.
Why tool output does not compare
Each category of scanner describes a different kind of object.
- SAST reports a code location that matches a rule: file, line, rule identifier, sometimes a data-flow trace.
- SCA reports a package version with a known vulnerability: package name, version, CVE or advisory identifier.
- Container scanning reports the same kind of thing as SCA, but per image and per layer, including operating system packages.
- DAST reports a behaviour observed over HTTP: URL, method, parameter, evidence in a response.
Severity scales differ too. One tool says “error”, another “High”, a third gives a CVSS vector. Some include a CWE, some do not. Counting them together without normalisation is like adding currencies without an exchange rate.
Design the schema small
The temptation is to design a schema that holds every field from every tool. I resist it. A schema should hold what you need to decide and track, and link to raw evidence for everything else.
The core fields I use:
| Field | Purpose |
|---|---|
fingerprint |
Stable identity used for deduplication across scans and tools |
category |
sast, sca, container, iac, dast, api, pentest, secret |
cwe, cve |
Weakness class and specific vulnerability, where known |
location |
One of: file and function, package URL, image and package, endpoint and parameter |
severity |
Raw tool severity, CVSS base score, normalised level |
status |
open, triaged, suppressed, fixed, verified, reopened |
owner |
Team or service that must act |
evidence_ref |
Pointer to the original SARIF result or raw report |
SARIF 2.1.0 is a good interchange format on the way in. Many SAST and SCA tools emit it natively, and it has a place for rule metadata, locations and fingerprints. I still map it into an internal schema, because SARIF was designed for static analysis results and does not natively model ownership, lifecycle state or business context.
- 01Raw tool output
- 02SARIF adapter
- 03Map to schema
- 04Fingerprint
- 05Merge or insert
Fingerprinting for deduplication
A fingerprint answers one question: is this the same issue I have seen before? Get it wrong in one direction and every rebuild creates duplicate tickets. Get it wrong in the other and distinct issues merge, so one fix appears to close two problems.
The inputs depend on the category.
- SAST: rule family, CWE, file path, enclosing function name, and a normalised snippet (whitespace and comments stripped). Not the line number, which moves whenever someone edits code above it.
- SCA: the package URL (purl) without the repository path, plus the vulnerability identifier. The same vulnerable package in two services is two findings, because two teams must act, so I add the repository as a separate key rather than into the hash.
- Container: package URL and vulnerability identifier, plus the digest of the layer that introduced the package. This lets a fix in a shared base image close the finding everywhere it appears.
- DAST and API: HTTP method, a templated path (
/orders/123becomes/orders/{id}), parameter name and CWE.
Where a tool provides a stable fingerprint of its own, for example through SARIF partialFingerprints, I store it and prefer it for matching within that tool. My own fingerprint is used for matching across tools.
Correlation is not deduplication
Two results with the same fingerprint are merged. Two results that are related but not identical are linked. An SQL injection reported by SAST in a handler and a ZAP alert for SQL injection on the endpoint served by that handler are different evidence for the same weakness. Linking them is useful because static plus dynamic confirmation is strong evidence of real exposure. Merging them would lose that.
Mapping CWE and CVE
CWE is the bridge between categories. A CVE identifies a specific flaw in a specific product. A CWE identifies the class of weakness. SAST and DAST findings usually carry a CWE. SCA and container findings carry a CVE and often a CWE in the advisory record.
A few practical rules:
- Map every rule in your SAST and DAST configurations to a CWE once, in a maintained lookup table, rather than trusting each tool’s metadata at runtime.
- Prefer specific CWEs (for example CWE-89, SQL injection) over broad parent categories. Broad categories make framework mapping vague.
- Keep CVE and advisory identifiers as a list. One package version often has several, and aliases between databases are common.
With CWE in place, mapping to the OWASP Top 10 or other frameworks becomes a lookup, not a judgement made per finding.
Normalising severity
I keep three severity values on every finding, because collapsing them loses information.
- Tool severity, exactly as reported, for traceability.
- CVSS base score, from the advisory for SCA and container findings, or estimated from the rule metadata for SAST and DAST. See FIRST CVSS.
- Normalised level (critical, high, medium, low, info), derived by a documented mapping.
The mapping table is code, reviewed like code. For CVSS-based findings, I use the standard qualitative bands from the CVSS specification. For tools that only provide labels, I map labels explicitly per tool and record the mapping version on each finding, so that a change in mapping does not silently rewrite history.
Normalised severity is not priority. Priority adds exploit likelihood (EPSS), known exploitation and asset context. I cover that in the unified assurance pipeline case study.
A worked example in Python
The code below maps a single SARIF result from a SAST tool and a single dependency finding from an SCA tool into the same schema. It is deliberately compact; a production version adds validation and handles more edge cases.
import hashlibimport refrom dataclasses import dataclass, field
LABEL_TO_LEVEL = { # example per-tool mapping, versioned with the code "error": "high", "warning": "medium", "note": "low", "critical": "critical", "high": "high", "moderate": "medium", "low": "low",}MAPPING_VERSION = "2026-09"
def cvss_to_level(score: float | None) -> str | None: if score is None: return None if score >= 9.0: return "critical" if score >= 7.0: return "high" if score >= 4.0: return "medium" return "low" if score > 0 else "info"
def normalise_snippet(text: str) -> str: text = re.sub(r"//.*|#.*", "", text) # strip line comments return re.sub(r"\s+", " ", text).strip()
def fp(*parts: str) -> str: return "sha256:" + hashlib.sha256("|".join(parts).encode()).hexdigest()
@dataclassclass Finding: fingerprint: str category: str title: str location: dict cwe: list[str] = field(default_factory=list) cve: list[str] = field(default_factory=list) tool_severity: str | None = None cvss_base: float | None = None level: str | None = None mapping_version: str = MAPPING_VERSION evidence_ref: str | None = None
def from_sarif_result(run: dict, result: dict, ref: str) -> Finding: rule_id = result["ruleId"] rules = {r["id"]: r for r in run["tool"]["driver"].get("rules", [])} rule = rules.get(rule_id, {}) tags = rule.get("properties", {}).get("tags", []) cwe = sorted({m.upper() for t in tags for m in re.findall(r"cwe-\d+", t, re.I)}) loc = result["locations"][0]["physicalLocation"] path = loc["artifactLocation"]["uri"] snippet = normalise_snippet(loc.get("region", {}).get("snippet", {}).get("text", "")) func = result.get("properties", {}).get("enclosingFunction", "") level_label = result.get("level", "warning") return Finding( fingerprint=fp("sast", rule_id.split(".")[-1], ",".join(cwe), path, func, snippet), category="sast", title=result["message"]["text"][:200], location={"path": path, "function": func}, cwe=cwe, tool_severity=level_label, level=LABEL_TO_LEVEL.get(level_label, "medium"), evidence_ref=ref, )
def from_sca_record(rec: dict, repo: str, ref: str) -> Finding: purl = f"pkg:{rec['ecosystem']}/{rec['package']}@{rec['version']}" vuln_ids = sorted(rec.get("aliases", []) + [rec["id"]]) cvss = rec.get("cvss_score") return Finding( fingerprint=fp("sca", purl, vuln_ids[0]), category="sca", title=rec.get("summary", rec["id"]), location={"repo": repo, "purl": purl}, cwe=rec.get("cwe_ids", []), cve=[v for v in vuln_ids if v.startswith("CVE-")], tool_severity=rec.get("severity"), cvss_base=cvss, level=cvss_to_level(cvss) or LABEL_TO_LEVEL.get(str(rec.get("severity")).lower()), evidence_ref=ref, )Two things are worth noting. The SCA fingerprint uses the first identifier in a sorted list, so the same advisory reported under a GHSA and a CVE alias still produces the same fingerprint. And the SAST fingerprint deliberately excludes the line number, which is the single most common cause of duplicate tickets I see.
False positives and suppression with expiry
Every scanner produces false positives. The question is how you record that judgement without creating a hiding place.
My rules for suppression:
- Scope it to a fingerprint, never to a whole rule or a whole file, unless the rule is being retired for everyone through a reviewed configuration change.
- Record a reason category (false positive, accepted risk, mitigated by control, not applicable) and free text explaining why.
- Require an approver for accepted risk. A developer can mark a false positive; only a security lead can accept a real risk.
- Set an expiry. False positives might expire after twelve months, accepted risks after thirty to ninety days (example values). On expiry the finding reopens and must be re-justified.
- Report suppressions as a metric. A rule with a high suppression rate is a candidate for tuning; a team with a growing pile of accepted risks is a conversation for leadership.
Triage outcomes are also training data for the pipeline. If a particular rule produces mostly false positives in one codebase, the right answer is often to tune or disable the rule, not to suppress each instance.
How to know it works
Normalisation is code, so test it like code. I keep fixture files from each tool and assert three things in CI: the same issue from two tools produces one finding; a trivial edit (moving a function, adding a comment above it) keeps the fingerprint; and two distinct issues in the same file produce two fingerprints. When a scanner is upgraded, the fixtures are re-run before the new version reaches production, because output formats drift.