A remediation SLA is the maximum time your organization allows between finding a vulnerability and fixing it. The SLAs that teams actually meet have a few things in common. They are tiered by real risk, meaning whether a flaw is being exploited, whether the system faces the internet and how much the asset matters, rather than by CVSS score alone. They are agreed with the IT staff and system owners who do the work. And they come with clear rules for when the clock starts and stops, an exception process with expiry dates, and reporting that someone reads.
What is a vulnerability remediation SLA?
A remediation SLA is an internal commitment, not a contract with a customer. For each category of vulnerability, it states how many days the responsible team has to remediate or to get a formal exception.
"Remediate" should mean the vulnerability is gone: patched, upgraded, reconfigured, or the affected component removed. Temporary mitigations matter, but they are not the same as a fix.
Why do remediation SLAs get missed?
Most missed SLAs trace back to the policy rather than to the people. The usual causes:
- Too many findings in the top tier. When tiers map straight to CVSS severity, a large share of findings lands in "critical," and a seven-day deadline for hundreds of items is fiction.
- Deadlines set without the people doing the work. Security writes the policy, IT learns about it from an overdue report, and the timeframes ignore change windows and staffing.
- Ambiguous clock rules. If nobody agrees on when a finding became due, every overdue item turns into a debate.
- No legitimate way to say "not yet." Without an exception process, teams quietly ignore findings they can't fix, and the backlog stops meaning anything.
- Reporting nobody acts on. If missed SLAs carry no follow-up, the deadlines become suggestions.
How should you tier vulnerabilities for remediation SLAs?
Tier by risk, using a handful of factors you can actually determine for each finding. Three of them do most of the work.
Is it being exploited in the wild?
Known exploitation is the strongest signal you have. CISA's Known Exploited Vulnerabilities (KEV) catalog lists vulnerabilities with reliable evidence of active exploitation, and it's free to consume as a feed. A known-exploited vulnerability on a relevant system should jump to your fastest tier, whatever its CVSS score.
Is the system exposed to the internet?
An unpatched internet-facing system can be reached by anyone. Think VPN gateways, web applications, remote access portals and file transfer servers. The same flaw on an internal server requires an attacker to get inside first, so exposure should move a finding up a tier.
How critical is the asset?
A vulnerability on a domain controller or a database of customer records matters more than one on a test VM. You need at least a simple criticality rating (high, medium, low) in your asset inventory. If you don't have one, start with the systems the business can't run without and the ones holding sensitive data.
Where do CVSS and EPSS fit?
CVSS base scores describe the technical severity of a flaw in general, not the risk it poses in your environment. Use CVSS as a starting input, then adjust for exploitation, exposure and criticality.
FIRST's Exploit Prediction Scoring System (EPSS) estimates the probability that a vulnerability will be exploited in the next 30 days. It can help separate the likely from the theoretical among findings that aren't yet known to be exploited.
An example tier structure
Here is a starting point. The timeframes reflect common practice rather than any standard, so adjust them to what your teams can sustain.
| Tier | Typical criteria | Starting-point SLA |
|---|---|---|
| Emergency | Known exploited (for example, listed in KEV) and internet-facing or on a critical asset | Mitigate within 48 hours, remediate within 7 days |
| High | Known exploited on other internal systems, or critical/high severity on an internet-facing or critical asset | 14–30 days |
| Medium | High severity on internal, non-critical systems, or medium severity on exposed systems | 60–90 days |
| Low | Everything else | Next scheduled maintenance cycle, up to 180 days |
Keep the number of tiers small. Four is usually enough, and the Emergency tier should fire rarely enough that it gets real attention when it does.
How fast should critical vulnerabilities be patched?
There isn't one right number, but there is a useful public benchmark. CISA's Binding Operational Directive 22-01, issued in November 2021, requires U.S. federal civilian agencies to remediate vulnerabilities in the KEV catalog by the due date CISA assigns to each entry. The directive set default windows of two weeks for vulnerabilities with CVE IDs assigned in 2021 or later, and six months for older ones.
BOD 22-01 doesn't apply to private companies, but it's a fair reference point. A similar or shorter window for known-exploited flaws on your own internet-facing systems is reasonable.
Beyond that, let your own data guide you. If your teams took 45 days on average to fix high-priority findings last quarter, a 7-day SLA won't change behavior on its own. Set a target that stretches the team without being impossible, then tighten it as the process improves.
How do you get IT and system owners to agree to SLAs?
SLAs that security imposes tend to be ignored. SLAs that the patching teams helped write tend to be met. A workable process looks like this:
- Assign an owner to every asset. A finding without an owner has nobody to meet the SLA.
- Share the data first. Show owners the current backlog sorted into the proposed tiers, so they can see the workload each tier would create.
- Align with maintenance windows. If a team patches servers monthly, a 14-day SLA means some findings need out-of-cycle work. Agree up front when that is expected.
- Pilot before enforcing. Run the tiers for a month or two and measure results without escalating. Adjust anything that clearly doesn't fit.
- Get leadership sign-off. Have the final policy approved by someone with authority over both security and IT, so escalations carry weight.
When does the SLA clock start and stop?
Write these rules down. Most arguments about overdue findings are really arguments about the clock.
When the clock starts. The most common and defensible choice is the date the vulnerability was first detected in your environment, whether by a scan, an agent, a penetration test or a vendor notice. Don't start it at ticket creation, which lets triage delays disappear from view.
When a finding changes tier. Vulnerabilities often get added to KEV after you first detect them. A sensible rule is that the new, shorter window runs from the date of the change, but the finding is never due later than its original deadline.
When no fix exists. If the vendor hasn't released a patch, a remediation clock can't reasonably run. Treat it as a mitigation requirement instead: apply a workaround or compensating control within the tier's mitigation window, and start the remediation clock when the fix is published.
When the clock stops. The clock stops when a rescan or agent check verifies the vulnerability is gone. A closed ticket is not proof. If scans run weekly, build that lag into your timeframes rather than counting unverified fixes.
When the clock pauses. Only an approved exception pauses or stops the clock. "Waiting for the next change window" is not a pause.
How should exceptions and risk acceptance work?
Some vulnerabilities can't be fixed on time. The software may be end-of-life, or an upgrade might break a critical application. A formal exception process keeps these findings visible instead of letting them sit silently in the backlog.
Each exception request should record:
- The finding, the affected assets and the current tier
- Why it can't be remediated within the SLA
- Compensating controls in place, such as restricted network access, a disabled feature or extra monitoring
- The business owner who is accepting the risk
- An expiry date
The expiry date is the part teams skip, and it's the most important one. Permanent exceptions become invisible. A common approach is to cap exceptions at around 90 days for higher-risk findings and allow longer for lower tiers, with each renewal requiring a fresh review.
Match approval authority to risk. A low-tier exception might need only the system owner and a security analyst. An exception on an Emergency-tier finding should need senior leadership sign-off, because they are the ones accepting that risk.
NIST SP 800-40 Rev. 4, published in 2022, is a helpful reference here. It describes four risk responses (accept, mitigate, transfer and avoid), which map neatly onto the choices an exception process forces you to make.
What happens when a remediation SLA is missed?
Escalation should be predictable, and owners should know in advance what happens at each stage. A typical ladder:
- Before the deadline: an automated reminder to the owner, for example at 75% of the SLA window.
- At breach: the owner and their manager are notified, and the finding appears on the overdue report.
- A set period after breach (say, 14 days): escalation to the IT director or head of security.
- Persistent breaches: review at whichever leadership or risk forum owns technology risk.
Escalation isn't about blame. Its job is to surface blockers the system owner can't solve alone, such as staffing gaps or vendor dependencies.
How should you report on SLA performance?
A short, consistent monthly report is enough for most small and mid-sized organizations. Useful measures include:
- SLA compliance by tier: the percentage of findings remediated within their SLA during the period.
- Overdue findings: count and age, broken down by owner or team.
- Mean time to remediate by tier: compared against the target for that tier.
- Open exceptions: how many exist, and how many expire in the next 30 days.
- KEV exposure: any KEV-listed vulnerabilities currently open in your environment.
Report by team, not just in aggregate. An 85% overall compliance rate can hide one team sitting at 40% that needs help.
Frequently asked questions
Should remediation SLAs be based on CVSS scores?
CVSS is a reasonable input but a poor sole basis. Base scores ignore exploitation, reachability and what the asset does, and tiering on them alone usually overloads the top tier.
What if a patch isn't available yet?
Apply a mitigation within the tier's mitigation window, such as a vendor workaround, disabling the affected feature or restricting access. Start the remediation clock when the vendor publishes a fix.
Do SLAs apply to misconfigurations as well as software vulnerabilities?
They should. Exposed cloud storage, default credentials and overly permissive firewall rules can be as dangerous as a missing patch. Tier them using the same factors of exposure, asset criticality and ease of abuse.
How often should we review our SLA policy?
At least once a year, and after any major change to your environment or tooling. If compliance sits near 100% for several quarters, consider tightening the timeframes.
Next steps
- Pull the last six months of remediation data and measure how long each severity actually took to fix.
- Add internet exposure and asset criticality to your inventory if they aren't there already.
- Draft four tiers with starting-point timeframes and review them with system owners.
- Write down the clock rules and the exception process, including expiry dates.
- Pilot for a month, adjust, then start reporting and escalating.