Exposure Management

Much Ado About Validation

by Brad Hibbert, COO & CSO//17 min read/

See How Brinqa + PlexTrac Close the CTEM LoopSee How Brinqa + PlexTrac Close the CTEM Loop

What's Validation Worth, Does AI Replace the Work, and What Should It Cost?

In the last piece I wrote about validation, I laid out the two moments: pre-remediation and post-remediation, and the methods you'd use for each. There we covered the work itself. This one covers everything that comes after the testing is done.

  • How you measure whether it worked?
  • How you report it to people who don't care about attack paths?
  • How much of your budget does it deserve?

Most programs can tell you they ran the tests; fewer can tell you what came out of them. And almost none can tell you whether they're spending the right amount on the whole effort.

You can usually tell which camp a program is in within the first minute of asking. Ask a validation team how the program is doing, and you'll get an activity answer: tests run, engagements completed, findings confirmed. Ask a CFO the same question, and none of that means anything to them. Activity isn't value. Value means real risk went away, not just a number on a dashboard, and you can explain how you know that in plain terms.

How do you measure the value of security validation?

Start by throwing out the metrics that describe effort instead of outcome. Number of pen tests run, number of BAS simulations fired, hours billed by the red team, these tell you the program is busy, not that it's working. If your reporting leans on any of them, you're reporting motion.

What you actually want are metrics that answer the two questions validation exists to answer.

What pre-remediation validation should tell you

For pre-remediation validation, ask what it eliminated and what it sharpened. The number that matters is your noise reduction rate, the percentage of prioritized exposures that validation determined were unreachable, unexploitable, or already mitigated by a control nobody had credited. A program with a healthy validation practice should be able to say "we took forty percent of the list off the table before anyone touched a ticket." That's not a nice-to-have stat, that's the number that justifies why validation exists at all, because it's directly proportional to the hours of remediation work your teams didn't waste.

What post-remediation validation should tell you

For post-remediation validation, ask a blunter question: how many "closed" tickets are actually closed? Track a fix confirmation rate, the share of remediated exposures that were re-tested and genuinely stopped the attack, versus the share that were closed on paper but still passable when you re-ran the test. Anything short of a hundred percent should bother you a little, because that gap is basically how much of your reported risk reduction isn't real.

Line chart tracking noise reduction rate and fix confirmation rate by quarter, showing both metrics trending upward toward the high 90s.

Sample report snippet: noise reduction rate and fix confirmation rate, tracked by quarter.

In the example above, the noise reduction rate, in yellow, is the share of prioritized exposures that pre-remediation validation eliminated before anyone touched a ticket, climbing here from the high twenties into the low forties. The fix confirmation rate, in green, is the share of closed tickets that were re-tested and actually held, climbing toward the high nineties. Together they answer the two questions validation exists to answer, whether the list is getting cleaner going in, and whether the fixes are actually sticking coming out.

What "risk removed" actually means

Then there's the metric that ties both halves together and is the one boards actually respond to: validated risk removed over time. That means exposure confirmed real before the fix, and confirmed eliminated after it. A ticket count or a patch count can't tell you that, only re-tested, validated risk can. Plot that next to your raw exposure count and you get the chart that actually explains what the security budget bought. A raw finding count going down tells a board nothing, since it might just mean you scanned less. The validated risk curve going down means less to attack, and you can prove it.

Line chart comparing validated risk removed against raw exposure count over six quarters, indexed to Q1, showing validated risk declining steadily while raw exposure count stays flat.

Sample report snippet: validated risk removed (solid) vs. raw exposure count (dashed), indexed to Q1.

In the example above, the dashed line is your raw exposure count, and it stays noisy and roughly flat because new findings show up about as fast as old ones close. The solid line is validated risk removed, and it declines steadily instead. If you only reported the raw count, the program would look stalled. The validated risk line is the one that survives the question of whether you just scanned less. This is the high level chart built for the board

Side Note: Quick detour on what I actually mean by “risk” here, because the word gets thrown around loosely enough that it stops meaning much. I don't mean a CVSS score, and I don't mean a count of criticals. Risk is likelihood times consequence: how likely is it that this specific exposure gets reached and used, and how much would it matter if it did, given what it sits next to and what it protects. A critical vulnerability with no path to anything important is low risk, no matter what the score says. A medium vulnerability three hops from your identity provider is not.

That's also the template for how you communicate it. Don't lead with a score or a count when you're briefing someone outside security. Lead with the asset and the path.

For example: this exposure sat on a route to the customer database, we confirmed it was reachable, we fixed it, and we confirmed the route no longer works.

That sentence means something to a CFO or a board member in a way that “we resolved forty criticals” never will. The risk-removed number you're plotting over time should be built out of exposures described that way, not out of severity totals. If you can't say what asset was protected and what path was closed, you don't actually know what you reduced.

The supporting numbers

A few supporting numbers round this out:

  • mean time to validate, how long it takes from “here's a prioritized exposure” to “confirmed exploitable or confirmed not,” because a validation program that takes six weeks to answer a question isn't keeping pace with anyone.
  • coverage of your crown-jewel attack paths, what fraction of the routes to your most important assets have been tested by anything, recently.

That second one tends to be the most uncomfortable number in the deck, because it's usually the one metric that isn't shaped by what was convenient to test. Validation work naturally drifts toward the exposures that are easy to schedule and quick to confirm, not the ones sitting on your most sensitive systems, so this is often the number that reveals a program has been thorough everywhere except the handful of places that would actually hurt. It's uncomfortable because it's hard to argue with, and it's usually the most useful for exactly the same reason.

Bar chart showing the share of crown-jewel attack paths tested each quarter, sitting below the 50% coverage target line.

Sample report snippet: share of crown-jewel attack paths tested in the trailing quarter.

The example above is the uncomfortable one. It shows the share of routes to your most important assets that have actually been tested in the trailing quarter, sitting well under the fifty percent target line. A program can have a strong noise reduction rate and a strong fix confirmation rate and still have this number sit low, because validation naturally drifts toward what's easy to test rather than what's most critical. It's also the number most likely to turn into a specific, fundable budget ask instead of an abstract one.

How to report this to the board

When you report this upward, lead with the risk-removed curve, support it with the noise reduction rate and the fix confirmation rate, and hold the activity metrics in reserve for anyone who asks how you got there. The board doesn't want your methodology. They want to know if the organization is safer than it was last quarter, and whether they can trust the answer.

Knowing what to measure only gets you halfway there. The other open question is who, or what, should be doing the validating in the first place. Your pen testers still need somewhere real to report and validate that work, which is what PlexTrac's for. What's actually up for debate is who's doing the testing itself.

Better Together: Brinqa + PlexTrac and the Next Era of Exposure Management

Learn MoreArrow RightLearn MoreArrow Right

Does AI replace the pen tester?

No, and the people asking the question usually aren't asking it precisely enough. What's actually changing is which parts of validation scale and which don't.

Coverage and frequency scale. An agentic pen testing tool or a BAS platform can run the same falsifiable test, a test with a clean pass or fail, did this control stop the attack, yes or no, hundreds of times a week, against assets a human team would never get to on any reasonable cadence. That's not a threat to a tester's job, it's a threat to the idea that validation should be rare and expensive. The work that scales is the work that was always mechanical: confirm the exploit, replay the technique, re-test after the patch.

Judgment scales less well, and I wouldn't bet on that changing completely. The newest agentic tools are closing the gap faster than most people expected a year ago. Some now surface business-logic flaws, broken authorization, IDORs, that human testers miss, largely because they can read the whole codebase instead of poking at it from the outside. That's changing what AI exploitability actually means in practice. But the honest read of where things actually stand right now is that the hardest classes are still the exception, not the rule. A widely cited late-2025 academic benchmark found the best autonomous agent missed a critical remote code execution flaw that roughly eighty percent of human testers caught.

A recent HackerOne researcher survey still has a majority of respondents saying AI falls short specifically on business logic, authorization, and the kind of multi-step exploit chaining nobody modeled ahead of time. That gap is real, it's shrinking every quarter, and treating it as if it has already closed, especially on the assets you can least afford to be wrong about, is, in my opinion, still a mistake.

So the honest way to report this internally isn't “how much have we automated,” it's “where is judgment being spent.” Show the split: which validated findings came from automation running at scale, and which came from human-led engagements on the assets that warranted the expense. That split is itself a defensible metric, because it demonstrates you're not paying your senior red team for hours to do what a script could do, and you're not asking a script to do what only a person can.

Once you know what to measure and where judgment is actually being spent, the budget conversation stops being a guess. It becomes a matter of matching spend to where each type of work belongs.

How do you budget for security validation?

Most validation budgets are set the wrong way: as a percentage of the security budget, or a headcount ratio, or last year's number plus ten percent. None of those numbers have anything to do with your actual exposure. Here's a better way to think about it.

Tier your validation spend to the asset, not the tool. Your crown-jewel attack paths, the ones that reach the data stores and identity systems that would actually hurt if compromised, get the expensive, infrequent, human-led validation. These are red team engagements, reserved for the small number of paths where the cost of being wrong dwarfs the cost of the engagement. Your important-but-not-crown-jewel exposures, the ones tied to known exploited vulnerabilities or sitting on plausible entry paths, get automated penetration testing run often enough to keep pace with how fast your environment actually changes. And everything you've already fixed once gets BAS, cheap per run, run continuously, because its whole job is catching drift before it becomes a repeat finding.

Side Note: A BAS platform doesn't hunt for new problems, it replays a known technique on a fixed schedule and checks for a plain pass or fail. A common version of this is exfiltration testing, safely trying to move a dummy sensitive file out through a cloud upload or DNS tunneling, the way an attacker actually would, and checking whether DLP and egress monitoring catch it. Run that same test weekly, and the day it succeeds when it shouldn't, you know something drifted, a proxy rule got rolled back, a DLP policy got misconfigured, and you know it within a day, not six months later at the next pentest.

This approach turns the budget conversation from “how much should we spend on validation” into “what's the cost of the coverage gap.” If you can show that a meaningful share of your crown-jewel paths haven't been tested in the last two quarters, that's not an abstract argument, that's a specific, fundable gap.

Here's what that looks like in practice: Let’s take the example that your crown jewels are the customer database and the identity provider. The red team has hit the database twice this year. However, nobody has actually tried the three routes that reach the identity provider, the vendor VPN with the shared credential, the service account that hasn't rotated in a year, and the backup jump host that was supposed to be decommissioned. That's a coverage gap. Not a mystery, not an unknown unknown, just three known paths that sit there untested because they were never anyone's job to check. Conversely, if your BAS coverage is thin on the assets you've already spent remediation dollars on, you're at risk of quietly losing ground you already paid for. Size the budget to close the largest gap first, not to hit a percentage.

There's another trap worth naming: buying tools instead of buying outcomes. A new scanner license or a second BAS subscription looks like progress on a budget spreadsheet, but a line item doesn't tell you if anything actually got safer. So run a simple test on every validation purchase, tool, hire, or engagement: since we added this:

  1. Our validated risk-removed number go up?
  2. Did our fix confirmation rate go up?

Say you bought a second BAS platform this year because the first one felt limited. If neither number moved after that purchase, you bought more motion, not more safety, and that's the moment to cut it, not defend it in next year's budget.

There's another trap worth naming: buying tools instead of buying outcomes. A new scanner license or a second BAS subscription looks like progress on a budget spreadsheet, but a line item doesn't tell you if anything actually got safer.

So, much ado about validation?

Not really. The activities aren't the hard part anymore, and neither is deciding whether AI belongs in the mix, it does, doing the scale work while your people do the judgment work. The hard part, the part most programs still get wrong, is proving what validation was worth and spending accordingly. Report risk removed, not tests run. Track whether your fixes actually held, not whether the ticket got closed. And size the budget to the coverage gap on the paths that matter, not to a percentage that has nothing to do with your risk.

A validation program you can't measure is a cost center waiting to get cut. A validation program you can measure is the reason the rest of the budget gets approved.

One thing I've glossed over here: all of this, the noise reduction rate, the crown-jewel coverage number, the tiered budget, assumes you already know your attack paths. In practice, that's the harder problem sitting underneath this one. How do you actually discover the routes through your environment, keep that inventory current as identities, assets, and trust relationships change weekly, and use it to decide what validation should work on first, instead of what's easiest to schedule? That's the subject of the next piece. Stay tuned.

The fastest way to know where you stand: put your own environment in front of someone who can show you.

Talk to a Brinqa ExpertArrow RightTalk to a Brinqa ExpertArrow Right

FAQs

Don't measure activity, measure outcome. Track your noise reduction rate (the share of prioritized exposures validation ruled out before remediation), your fix confirmation rate (the share of "closed" tickets that stayed closed when re-tested), and validated risk removed over time. That last one is the metric boards actually respond to, since it reflects exposure confirmed real before the fix and confirmed gone after it, not just a ticket or patch count.

A raw finding count can go down just because you scanned less, so it tells a board nothing on its own. Validated risk removed only counts exposures that were confirmed real, fixed, and confirmed gone on re-test, so a decline in that number means the organization is actually safer, not just reporting less.

No. What's changing is which parts of validation scale. Coverage and frequency scale well with AI and agentic tools, running the same falsifiable pass/fail test hundreds of times a week against assets a human team could never reach on that cadence. Judgment scales less well. The hardest classes of findings, business logic flaws, broken authorization, multi-step exploit chains, are still where human testers catch what AI agents miss.

Report the split, not the automation percentage. Show which validated findings came from automation running at scale and which came from human-led engagements on assets that warranted the expense. That split is itself a defensible metric: it shows you're not paying senior red team hours for what a script could do, and not asking a script to do what only a person can.

Not as a percentage of the security budget or a headcount ratio, those numbers have nothing to do with your actual exposure. Tier spend to the asset instead: crown-jewel attack paths get expensive, infrequent, human-led red team engagements; important-but-not-crown-jewel exposures get automated penetration testing run on a regular cadence; already-remediated findings get cheap, continuous BAS testing to catch drift.

A coverage gap is a specific, known attack path to a critical asset that hasn't actually been tested recently, not an abstract risk. Naming it turns the budget conversation from "how much should we spend" into "what's the cost of leaving this path untested," which is a far easier ask to get funded than a percentage increase.

Lead with the risk-removed curve, support it with the noise reduction rate and fix confirmation rate, and hold activity metrics in reserve for follow-up questions. Boards don't want methodology. They want to know whether the organization is safer than last quarter and whether they can trust the answer.

Run a simple test after any validation purchase: did the validated risk-removed number go up, and did the fix confirmation rate go up? If neither moved, the purchase bought motion, not safety, and that's the signal to cut it rather than defend it in next year's budget.

B
Brad Hibbert
Chief Operating Officer & Chief Strategy Officer
Brad Hibbert brings over 30 years of executive experience in the software industry, with a proven track record of aligning business and technical teams to drive growth and customer success.
See all of Brad's postsArrow Right

Ready to Unify Your Cyber Risk Lifecycle?

Get a DemoArrow RightGet a DemoArrow Right