Service level language is where most managed IT disputes start. The client reads 'resolution within four hours' and expects their crashed accounting server rebuilt by lunch. The provider meant four hours for a password reset. Both are reading the same sentence.
This post pulls apart the terms: SLA versus SLO, response versus resolution, and the severity matrix that makes 'critical' a defined word rather than a feeling. Get these right and the contract stops being a source of arguments.
SLA and SLO are not the same document
An SLA is a promise in a contract with consequences attached: a service credit, a right to terminate, a penalty. An SLO is an internal target you measure yourself against and report on. A sensible managed IT agreement has a small number of SLAs, mostly around response, and a longer list of SLOs that you publish and track but do not pay penalties on.
This split is honest. You can promise that a human will pick up a critical ticket within a short window at any hour, because you control staffing. You cannot promise that a failed RAID controller will be replaced in four hours, because you do not control the courier. So response is an SLA and resolution is an SLO, with a stated target and a monthly report of how often it was met.
- SLA: contractual, measured, with a remedy when missed
- SLO: internal target, reported to the client, drives improvement
- Response time: SLA candidate, because you control it
- Resolution time: SLO, because the cause is unknown until you look
- Uptime: an SLO per system, stated as a target such as 99.9 percent, with the measurement method written down
Response versus resolution
Response means a qualified technician has acknowledged the ticket, communicated with the requester, and begun work. It does not mean an automated email went out. Say that explicitly, because some providers count the auto-reply and clients feel cheated when they learn it.
Resolution means the requester agrees the issue is fixed, or a workaround is in place and a permanent fix is scheduled. Time to resolution should pause when you are waiting on the client, a vendor, or a part, and the ticket should show that clock stopping. Report both the raw and the paused figure so nobody feels the numbers are being gamed.
- Define response as human acknowledgement plus start of work, not an auto-reply
- Define resolution as requester confirmation or an agreed workaround
- Pause the resolution clock on 'waiting on customer' and 'waiting on vendor' states
- Report raw and paused resolution times side by side
The severity matrix
Severity is a function of two things: how many people are affected and whether there is a workaround. A matrix with those two axes gives you four levels that anyone can apply consistently, including the client's office manager at 7am.
Critical means the business or a whole site cannot operate and there is no workaround: email is down for everyone, the main application is unreachable, a site has no connectivity. High means a department or a business process is blocked, or a single person is blocked and their work is time-critical (payroll on payroll day). Normal means a person is impaired but can work another way. Low is a request, a question, or a nice-to-have. Write examples for each level from the client's own environment so the words mean something to them.
- Critical: whole company or site down, no workaround, 24/7 response with a target under 15 minutes
- High: department or key process blocked, business-hours response within a short window, after-hours by arrangement
- Normal: one person impaired with a workaround, response the same business day
- Low: requests, questions, scheduled work, response within a couple of business days
- The provider may reclassify a ticket after triage, and must say why
Measuring and reporting without games
Pull the numbers from the ticketing system, not from memory. For each severity, report the count of tickets, the percentage that met the response target, the median and worst resolution time, and the reasons for the misses. Show it monthly in a short table and discuss it at the quarterly review.
Resist the urge to make the numbers look better by downgrading severity after the fact or closing tickets before the user confirms. Clients notice. A report that shows two missed critical responses and what changed to prevent the third is worth more than a report that shows perfection nobody believes.
Frequently asked questions
Should there be financial penalties for missing an SLA?
A modest service credit for missed response targets is reasonable and shows confidence. Penalties on resolution time invite arguments about causes. Keep remedies simple and tied to what you control.
Who decides the severity of a ticket?
The requester proposes it, the provider confirms it at triage using the matrix. Disagreements are settled by the written examples, and the account manager can adjust the matrix at the next review if the examples were wrong.
What about uptime guarantees?
State an uptime target per system, say how it is measured and what is excluded (scheduled maintenance, upstream provider outages), and report against it. A blanket uptime promise across everything is not measurable.
Takeaway
Promise what you control and measure what you do not. Response is an SLA, resolution is an SLO, and severity is a matrix of people affected and workaround available, with examples from the client's own business. Report the numbers honestly and the contract becomes a shared scoreboard instead of a weapon.