Where AI Helps in Vulnerability Triage, and Where It Guesses

If you have ever opened a vulnerability scanner export for a client, you know what comes next. There are hundreds of rows, sometimes thousands, most of them are marked high or critical, and somebody has to decide which ones get fixed this week. It is tempting to paste the whole thing into an AI tool and ask it what to fix first, and I understand why, because I have been in IT for over 20 years and the list has never been longer than it is now.

Jerry Gamblin’s annual CVE data review counted 48,185 CVEs published in 2025, which works out to about 132 a day and an increase of about 20 percent over the year before. Nobody reads 132 advisories a day, so it is a fair question, can AI do the triage for you? My answer is that it can do part of the job very well, and the other part it will guess at, and it will sound just as confident while it is guessing.

Where AI helps

These are all language jobs. The information is already on the page and the AI is rewording it or sorting it.

  • Plain English — it turns a scanner finding into a sentence a business owner can read and act on.

  • Grouping — it rolls forty findings into one action, such as “update Chrome on these twelve machines.”

  • First drafts — the executive summary, the ticket for the technician and the email to the client.

  • Reading the description — in a study covered by Help Net Security, six AI models scored more than 31,000 CVEs from the description alone, and the best of them got the attack vector right about 89 percent of the time, because the description usually says whether the flaw can be reached over the network.

Where it guesses

The trouble starts when the answer is not on the page, and these are the places I would not take an AI tool at its word.

  • Whether it is being exploited right now — a model only knows what it was trained on, and unless it is connected to live data it is working from memory that can be months old, while the list of what attackers are using changes every week.

  • The details of a specific CVE — affected versions, the fixed version and the patch number are exactly the kind of detail it can get wrong, and it will not tell you it is unsure.

  • Anything the description leaves out — in that same study all six models got the same 29 percent of CVEs wrong on availability impact, and combining the models did not fix it, because the information was never in the description to begin with.

  • Your client’s environment — it does not know which server faces the internet or which system runs the business. Researchers at Macquarie University and CSIRO’s Data61 tested four models on 384 real vulnerabilities, and the models were weakest on mission impact and tended to rate risk higher than it was.

That last one matters more than it sounds. An AI that calls everything urgent has put you right back where the scanner left you.

What to rank with

The ranking should come from data that somebody maintains and publishes, and most of it is free.

  • CISA KEV — the Known Exploited Vulnerabilities catalog is the list of flaws confirmed to be exploited in the wild. It has a little over 1,700 entries, against more than 48,000 new CVEs last year alone, so if a finding is on that list it goes to the top.

  • EPSS — published daily by FIRST, it is an estimate of the probability that a vulnerability will see exploitation activity in the next 30 days. It tells you how likely, and it does not tell you how bad.

  • CVSS — the severity score, which tells you how bad it would be if it were exploited and says nothing about how likely that is.

  • Your own context — whether the machine is internet facing and what it holds. Only you and your client know that.

A workflow that uses both

  1. Let the data do the ranking. KEV first, then EPSS and CVSS, using a fixed rule so the same scan gives the same answer every time.

  2. Add what you know about the client. Move the internet facing systems and the ones holding client data up the list.

  3. Let AI do the writing. Have it explain the top findings, group the fixes and draft the summary.

  4. Check the facts it states. Any version number or patch it mentions gets checked against the vendor advisory before it goes to the client.

  5. Keep the human sign-off. Your name is on the report, so the final read is yours.

Where to start

This is how I built ClarityOps at OpStacks, the ranking comes from KEV, EPSS and CVSS data and a fixed formula, and the report comes out in plain English for the client. However, you do not need to buy anything to get started. The KEV Triage Checklist on the Free Resources page at OpStacks.net is free and it is a good place to begin. Take a look at how your last client report was ranked, and let me know what you find.

If this was useful, subscribe to the OpStacks Weekly Digest at OpStacks.net, where I share a practical tip like this one every week.

Sources

Jerry Gamblin, 2025 CVE Data Review (January 2026)

Help Net Security, LLMs can assist with vulnerability scoring, but context still matters (December 2025)

Al Haddad, Ikram, Ahmed and Lee, Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization (October 2025)

FIRST, The EPSS Model

CISA, Known Exploited Vulnerabilities Catalog

Next
Next

Shadow AI: The AI Tools Already Running on Your Clients’ Machines