SLA for critical equipment across 20 branches
Choose an SLA for critical equipment by comparing four hour repair, next business day service, and replacement stock against branch downtime cost.

Critical equipment across a network of 20 branches rarely needs one SLA for every site. It needs a matrix: four hours for nodes whose failure immediately stops work or creates an unacceptable risk, next business day service for sites that can tolerate a pause, and a replacement pool wherever logistics make an engineer's promise unreliable.
The support contract price says little on its own. Compare the full cost of failure: lost margin or productivity, wages paid to idle staff, penalties, manual recovery work, and the likelihood that the stated deadline can actually be met in a particular city. After a few outages, the pattern becomes clear: a cheap SLA can cost a lot, while four hour repair without a nearby spare part often remains a polished line in a contract appendix.
Start with each site's hourly downtime cost
The hourly cost of downtime should set the service level, not the branch manager's seniority or the number of installed computers. Two branches with 40 employees each may need different terms: one takes payments and loses transactions immediately, while the other processes documents that staff can catch up on that evening.
Use a simple calculation for the first estimate:
Стоимость простоя за час =
потерянная валовая маржа
+ стоимость оплаченного, но остановленного труда
+ штрафы и компенсации
+ стоимость ручного обходного процесса
+ ожидаемая стоимость восстановления данных и очередей
Count gross margin rather than all revenue when a sale genuinely disappears. If an order merely moves to tomorrow, the loss is not the order value. It is the extra expense, the delay discount, and the share of orders that customers cancel. A checkout, laboratory system, or access control node can have consequences far beyond revenue: a queue forms, staff switch to paper, and someone later has to enter the accumulated records by hand.
It helps to calculate three values instead of one deceptively precise figure: a normal hour, a peak hour, and a full working day. Suppose a site loses 180,000 tenge in a normal hour, 420,000 tenge during a peak hour, and another 600,000 tenge clearing the backlog after an eight hour outage. The difference between recovery in four hours and recovery the next day then amounts to several million tenge in one incident, not an abstract availability percentage.
Do not add reputational damage to the downtime price with an arbitrary multiplier. Record an observable substitute: the number of complaints, customer compensation, reprocessing expense, and the share of cancelled appointments. Leave the part you cannot measure as a separate risk for management to decide. Otherwise the team will quickly see that someone shaped the model to justify a preselected expensive contract.
For every site, record who owns the estimate and when it will be reviewed. The hourly figure changes after a new process launches, turnover grows, a local warehouse closes, or some operations move to a central system. An annual review suits procurement, but a critical branch deserves another assessment after every material load change.
Criticality belongs to the service, not the box
A server is critical only because a specific service depends on it and has no workable bypass. Identical servers can belong to different classes: one runs a local print queue, while another hosts the system without which the branch cannot serve customers.
Write a short dependency chain for every site: business operation, application, data, network, power, compute node, and peripherals. This exercise often exposes an uncomfortable fact: the company pays handsomely to cover the server, while the actual single point of failure is a router, uninterruptible power supply, or scanner with an unusual interface.
Separate four concepts that procurement discussions constantly blur:
- criticality describes the consequences of an outage;
- availability describes the share of time the service must operate;
- recoverability describes how quickly the team brings the service back;
- supportability describes whether the provider can complete the repair with available people, parts, and site access.
Call all of this "reliability" and the contract becomes impossible to test. Equipment may have high calculated availability, yet one controller failure can stop a branch for two days while a part is shipped. The reverse also happens: a device fails more often than desired, but a local replacement restores work in 30 minutes.
ITIL 4 describes service warranty in terms of availability, capacity, continuity, and security. That is a useful frame, but it covers more than a hardware repair agreement. The equipment provider controls only the agreed part of the chain, so the service owner must connect the hardware SLA separately to recovery of the business operation.
ISO 22301 calls for priorities based on an analysis of how disruptions affect the business. The practical consequence is simple: a server's catalogue price cannot determine its service class. Its priority comes from the process's maximum tolerable downtime and the dependencies that keep that process running.
Four hours must mean restoration, not a response
A four hour SLA makes sense only when the contract states exactly what must happen by the end of the fourth hour. A service desk reply, an engineer's dispatch, part delivery, equipment restoration, and return of the business service are different outcomes.
Contracts often promise a "response within four hours." The provider registers the request after three hours and fifty minutes and meets the promise even though the equipment remains down. A critical site needs a measurable restoration target: a working node, a temporary replacement, or a failover after which the agreed service is available again.
Specify the clock. A 24/7 schedule counts calendar hours at night, on weekends, and on public holidays. An 8x5 schedule normally counts only the working window. Under the second arrangement, a request accepted on Friday evening may legitimately wait until Monday. A branch that opens on Saturdays cannot rely on the head office calendar.
Check what stops the timer. A provider may reasonably wait for site access, configuration details, or approval for a maintenance window, but every pause needs a reason and a timestamp. A general "waiting for customer" status lets the provider stop the clock after any email. List the allowed states instead: no physical access, the customer prohibited a restart, or a required diagnostic log was not supplied.
Define the restoration boundary. If the provider replaces a system board and the equipment passes its self test, but the operating system will not boot because the disk order changed, the branch is still not working. The contract may honestly stop at the hardware layer, but the internal plan must cover the remaining path: booting, checking the network, starting the application, and checking the accumulated queue.
Four hours without geography is not a commitment. The appendix should list all 20 branch addresses, access hours, pass requirements, distance to an engineer, and the storage location of every critical part. If the provider cannot show where the person and spare will come from for a remote city, treat four hour restoration as a risk with an unknown duration.
Next business day is not only for noncritical nodes
Next business day service suits any service that can safely survive one complete working cycle on a reserve or in manual mode. It is a deliberate economic choice, not a label for cheap hardware.
This level is sensible for a workstation when an employee can move to a free computer, for a node with complete automatic redundancy, and for a branch where operations may be postponed without losing data or breaching obligations. It can also be enough for an expensive development server if its outage does not affect current customers.
Check what "next" means. If the provider accepts a request before a stated cutoff, an engineer arrives on the next business day. After that cutoff, the visit may slip by another day. For branches in different time zones, specify the site's local time rather than the service centre's time.
NBD carries a hidden cost: the backup process must remain ready to work. A spare computer needs updates and the necessary permissions, a standby server needs current data, and a paper form must remain valid for later entry. If nobody has tested the bypass for six months, it is a poor reason for buying a cheaper SLA.
The popular advice to assign NBD to all redundant equipment is wrong. Redundancy reduces the chance of a service outage, but after the first failure the system has no margin left. While the team waits until the next day, a second failure, a failover error, or overload on the remaining node can turn a tolerable fault into a complete outage. Set a separate deadline for restoring redundancy in these systems even while the business service remains available.
Assess how long degraded operation is acceptable. A branch may still serve customers on one of two nodes, but do so at half speed. If the queue becomes unacceptable after three hours, formal availability does not justify NBD. Include degraded performance in the cost model, not just the binary states of working and down.
A replacement pool beats logistics but demands discipline
A replacement pool gives the most predictable recovery time wherever shipping a part or sending an engineer takes longer than the permitted outage. A branch employee or local engineer swaps a prepared unit, and the faulty device can be repaired without pressure from the incident clock.
Replacement works especially well for standard workstations, compact servers, network devices, power supplies, drives, and peripherals that can be exchanged as a whole. A complete twin for a large rack server may cost too much, but a pool of failure prone modules, controllers, drives, and power supplies often produces the same result for less.
A pool is not a shelf of sealed boxes. A replacement device needs a compatible configuration, current firmware, tested power, the right mounting hardware, and clear switching instructions. A workstation needs a ready method for restoring the profile and applications. A server needs verified boot, network, storage access, and monitoring settings.
Run the pool as a working system:
- Give every spare device an owner, a storage location, and a list of compatible sites.
- Test power on, completeness, and configuration version on a schedule.
- Seal each kit so a missing cable becomes obvious before an outage.
- Create a replenishment order immediately after issue and set a deadline for returning stock.
- Rehearse a replacement at a representative site, including confirmation of a business operation.
The main mistake is counting one spare twice. One server appears as backup for five branches even though a shared power incident or failed update may disable several devices at once. Size the pool for correlated failures, repair time, and transport distance. A local unit for a remote group of branches is often more useful than two devices in a central warehouse.
Replacement shifts some responsibility to the customer. Decide in advance who may open the chassis, replace a drive, see data on the failed medium, and sign the service record. In a regulated environment, nobody should send a drive by ordinary courier merely because that is convenient for the service process.
Put the 20 branches into service classes
Three classes are usually enough for a 20 branch network when the criteria are measurable and exceptions are documented. More classes complicate the contract, parts inventory, and service desk training, but rarely improve the outcome.
A workable matrix looks like this. Class A covers sites where downtime quickly costs more than premium service and no bypass exists. Their target is restoration within four calendar hours through local parts, an on call engineer, or a ready replacement. Class B contains branches that can work manually or on a reserve for a limited period. NBD with a tested bypass suits them. Class C covers sites where work can move elsewhere or continue on a free standard device while repair proceeds under an agreed plan through the shared pool.
Do not allocate classes on the principle that the capital receives A while regions receive B. A remote site may generate less revenue but perform an irreplaceable operation: issue mandatory documents, run the only warehouse in a region, or support a continuous production area. Geography changes the delivery mechanism, not the acceptable downtime by itself.
For every branch, complete a nine field record: business operation, working hours, normal and peak hourly cost, maximum tolerable downtime, available bypass, time the bypass can operate, critical configuration items, service class, and decision owner. That is enough for an initial justification and a later audit.
Then test borderline cases. If the cost of one average NBD incident exceeds the annual premium for four hour service, Class A has an economic case. If a four hour contract still requires eight hours of shipping, put the money into a local replacement. If a fault does not stop the service but removes redundancy, set a fast deadline for returning fault tolerance without declaring a full outage.
Do not confuse the site's class with every ticket's priority. A Class A site can report a failed noncritical monitor, while a Class B site can lose the only node supporting its bypass. The service desk should set priority from actual impact, using the class as a default rather than a ban on escalation.
The contract must describe a measurable ticket path
A good SLA lets both parties reconstruct an incident timeline without arguing over memories and emails. It needs common timestamps, ticket states, priority rules, and evidence of restoration.
A minimal ticket record looks like this:
{
"site_id": "BR-07",
"asset_id": "SRV-07-01",
"service_impact": "прием клиентов остановлен",
"priority": "P1",
"opened_at": "2026-02-12T09:14:00+05:00",
"access_window": "24x7",
"workaround": "нет",
"timer_state": "running",
"target_restore_at": "2026-02-12T13:14:00+05:00"
}
The site and asset identifiers remove arguments over coverage. The impact description explains priority. A timestamp with its time zone prevents disagreement between the central team and branch. The workaround field shows whether the priority can safely drop, and the timer state makes pauses visible.
Define when a ticket opens. If timing begins only after confirmation by phone, an unreachable support line becomes a loophole. It is safer to use the automatic registration time in an agreed channel and request missing information while the clock is running, apart from a predeclared set of mandatory fields.
Describe the closure test. "Equipment operational" is not enough. For a server, it may include successful hardware diagnostics, boot of the agreed operating environment, available network interfaces, and a test operation confirmed by the owner. If another team maintains the application, record the handover and a separate time for complete service restoration.
Remedies should change behaviour rather than decorate the contract. A small service credit received next quarter does not save a stopped branch. It is more useful to require a breach review, a plan to remove the cause, an increase in local stock after a repeated shortage, and a right to reconsider the particular site's class. Financial compensation still has a place, but it does not replace prevention.
Test the promise before the first outage
Accept an SLA through a simulated failure because a presentation cannot test a building pass, actual part availability, or an engineer's authority to change a configuration. Run the test without risking the live service: use a standby device, an agreed window, or a tabletop exercise for actions that cannot be performed safely.
Choose one site from each class and follow the whole route. An operator opens a ticket with incomplete but acceptable details. The service team assigns an engineer, confirms the part, and reports an estimated time. The branch arranges access. After replacement, the owner completes a test operation. An observer records every timestamp and deviation from the procedure.
Check four outcomes:
- the ticket received the correct priority without a call from an executive;
- the required spare part was physically where the plan said it would be;
- the engineer obtained access and current instructions;
- the business owner, not only the engineer, confirmed restoration.
A test that happens to finish within four hours does not yet prove readiness. The team should explain how much time remained, what happens at night, and who covers for an ill engineer. If success depends on one familiar employee, the contract relies on a personal favour rather than a repeatable process.
Repeat a sample test after changing the service partner, equipment model, branch address, or access rules. Check the replacement pool after every use: teams sometimes discover an empty warehouse slot only at the next failure. In a 20 site network, a rotation that touches every address and every critical equipment type during the planned cycle is enough.
Compare scenarios by expected annual cost
Choose among four hours, NBD, and a replacement pool by comparing the expected annual cost of each option. Include the support price, spare equipment, storage, testing, internal labour, and the remaining damage from outages.
Use this model:
Годовая стоимость варианта =
цена договора
+ годовая стоимость подменного фонда
+ внутренние расходы на готовность
+ сумма по сценариям(
вероятность сценария
x длительность простоя
x стоимость часа площадки
)
Do not derive a spuriously precise probability from two past failures. Build a range for a quiet, baseline, and severe year. Model a single hardware failure, a widespread firmware problem, a region made inaccessible by weather or transport, and a configuration error outside the hardware warranty. These scenarios show where a shared pool stops being shared salvation.
The comparison must account for deliverability. If the provider promises four hours but the required part has historically arrived in nine, use nine until the provider moves stock closer and confirms the change through a test. If a local replacement restores service in 45 minutes only when an administrator is on call, add the on call cost or use the employee's real arrival time.
A mixed answer may be best. Five Class A sites receive local replacements and four hour restoration, nine Class B sites receive NBD with a tested manual mode, and six small sites use a regional pool. These figures are not a universal proportion. They demonstrate that one contract label does not have to cover all 20 addresses.
Run a sensitivity analysis. Double the peak hourly cost, add one day to delivery at a remote branch, and model two simultaneous failures. If a small change in assumptions reverses the choice, it is a borderline decision: record a review trigger or buy limited stock that reduces the costliest risk.
Compare options over the same period. You buy a replacement server once and use it for several years, while paying a service premium annually. Spread the pool's purchase cost over its expected life, then add diagnostics, updates, and secure storage. Residual value can be counted too, but do not overstate it: the spare configuration ages with the main fleet and rarely sells according to the original plan.
Calculate the cost of an SLA breach separately. A service credit reduces the provider's invoice but does not recover the branch's losses. Use actual restoration time and the full damage in the economic model, then show compensation as a separate line. Otherwise a large credit creates the illusion that a prolonged outage became cheap.
Compare budget with avoided risk at each site. A 3 million tenge premium for faster service looks large in isolation, but it makes sense if one probable peak day incident costs 8 million. The same premium is pointless for a site that loses 150,000 tenge in a day and moves reliably to its reserve. This is not a promise of a perfect forecast. It is a transparent decision boundary.
Keep the initial assumptions beside the result: scenario frequency, delivery time, hourly cost, pool life, and availability of an internal specialist. Compare them with tickets a year later. Even a small fleet produces enough evidence to replace debatable estimates with actual timestamps, delay causes, and parts consumption.
Assign ownership across the entire recovery path
A working SLA model ends with named owners for every transition from the first failure signal to a normal business operation, not with a chosen tariff. The equipment provider, internal IT team, application owner, security team, and branch manager own different parts, and the gaps between them usually consume more time than replacing the part itself.
Create one responsibility table for P1 incidents: who confirms impact, who approves a restart, who arranges physical access, who decides on replacement, who restores the configuration, who tests the application, and who tells users that work has resumed. Every action needs a primary owner and a substitute for nights, leave, or illness.
The provider must show how it supports its promise: where engineers are located, which parts are held in the region, how escalation works, and what happens when tickets arrive together. The customer must provide access, a current asset list, configuration backups, and someone able to confirm the result. A one sided SLA does not survive a real incident.
GSE manufactures equipment in Kazakhstan, manages it from design through support, and has a nationwide service network with 24/7 technical support. Test those capabilities for each address: request the delivery plan for the required deadline, the contents of local stock, and the escalation procedure, then put verifiable terms into the contract.
Do not buy four hours for peace of mind or choose NBD solely to save one budget line. Put the downtime cost, the engineer's actual route, replacement readiness, and the time to restore the business operation side by side. The argument over the "best SLA" then becomes a testable decision for each of the 20 sites.
FAQ
How does four hour repair differ from a four hour response?
A response means the provider accepted the ticket or started work. Repair or restoration means the equipment performs the agreed function again or has been replaced with a working device. A critical site needs the second outcome in its contract.
When is next business day service enough?
NBD is enough when the site can operate safely on a reserve or in manual mode throughout the wait. Check the ticket cutoff, the branch calendar, and the actual readiness of the bypass.
How do I calculate a branch's hourly downtime cost?
Add lost gross margin, paid idle time, penalties, manual processing, and expected recovery expense. Calculate a normal hour, peak hour, and backlog effects separately.
Does a branch that only opens by day need a 24/7 SLA?
Not always. It needs one if a night failure must be cleared before opening or if critical batch processes run overnight. Otherwise a working window with a precise restoration deadline may be better value.
What is better for a remote branch, fast dispatch or replacement?
Replacement is usually more predictable when an engineer's journey exceeds tolerable downtime. The device still needs to be configured, tested, and accessible to someone authorised to make the swap.
How many spare devices do 20 branches need?
There is no universal ratio. Account for compatibility, repair duration, transport distance, and simultaneous failures. One central spare cannot be treated as available to every site at the same time.
Can we lower the SLA when servers run in a cluster?
Yes, if the cluster genuinely survives a failure and the team tests it regularly. Still set a short deadline for restoring redundancy because the system carries more risk after the first fault.
Which pauses may be excluded from SLA repair time?
Only predeclared and documented pauses, such as no physical access or the customer's refusal to allow a restart. A generic "waiting for customer" status makes the metric meaningless.
How can we test an SLA before signing a long contract?
Run a simulated ticket at sites in different classes and follow it through a test business operation. Record engineer assignment, part confirmation, access, replacement, and acceptance times.
Who should confirm that an incident is closed?
A technician confirms that the equipment works, while the business service owner checks a test operation. If the application sits outside the contract, record the handover and a separate full service restoration time.