Dedicated Server SLA Checklist: What to Verify Before Signing

SLA & Support

A dedicated server SLA should tell you what happens when infrastructure fails—not merely advertise an uptime percentage. Before signing, verify what the provider measures, when incident clocks start, which events are excluded, how hardware replacement works, and what remedy you receive if the commitment is missed. The strongest SLA is not necessarily the one with the largest percentage; it is the one whose scope matches your workload and whose obligations can be measured and enforced.

Start with your operational requirements

You cannot evaluate an SLA until you know which failures matter to the application. A public website, database node, backup server, and batch-processing host may tolerate very different outage durations.

Document these requirements before comparing providers:

  • Maximum tolerable interruption: How long can the service be unavailable before the business is materially affected?
  • Recovery time objective: How quickly must the application return after a server, network, or storage failure?
  • Recovery point objective: How much recent data can be lost if restoration from backup is required?
  • Support coverage: Do you need a staffed response at all hours, or is business-hours support sufficient?
  • Failure tolerance: Can traffic move to another server or location, or is this machine a single point of failure?
  • Responsibility boundary: Will the provider manage only the hardware and network, or also the operating system and application stack?

An infrastructure SLA does not create application availability by itself. If a workload cannot tolerate a server outage, the design normally needs redundancy, independent backups, tested restoration, and a way to redirect traffic. Service credits do not recover data or keep customers online.

The dedicated server SLA checklist

Clause What to verify Why it matters
Covered service Exact server plan, data center, network and optional SLA tier A general website promise may not apply to your order
Availability Measurement point, formula, reporting period and qualifying outage The same percentage can represent different obligations
Incident clock Start event, pause conditions and definition of restoration A fast response does not guarantee a fast repair
Hardware replacement Diagnosis and replacement targets, component stock and exceptions Replacement may exclude rebuilding the operating system or data
Support Hours, channels, severity levels and escalation path “24/7 infrastructure” does not always mean 24/7 human support
Exclusions Maintenance, attacks, customer error, software and third parties Excluded minutes may disappear from the availability calculation
Remedy Credit formula, claim deadline, cap and termination rights A commitment has limited value if the remedy is impractical

1. Confirm which documents control the service

Collect the complete contract set before approving the order. This may include the order form, master agreement, product schedule, SLA schedule, acceptable use policy, support policy and remote-hands price list.

Check the order of precedence when documents conflict. A sales page may show a headline SLA while the product schedule narrows its scope. Confirm that the applicable legal entity, data center, server family and support tier appear in the signed documents.

Also record whether linked policies can be changed unilaterally. If a provider may update an online SLA during the contract term, ask which version governs your service and how material changes will be communicated.

2. Define what “availability” means

Do not assume that server availability, network availability and application availability are interchangeable. A provider may measure reachability at its network edge even when the operating system is unavailable. Another may count only a loss of network connectivity caused by infrastructure under its control.

Ask the provider to identify:

  • The exact measurement point.
  • Whether availability is measured per server, per port, per network or across the account.
  • Whether partial packet loss, severe latency or degraded bandwidth qualifies as downtime.
  • Whether an unreachable management interface counts.
  • Whose monitoring system determines the official outage duration.
  • How conflicting customer and provider monitoring records are resolved.

The SLA should also state the reporting period. A monthly commitment is evaluated differently from an annual one. Under a simple 30-day calculation, the following percentages produce these theoretical downtime allowances before exclusions:

Availability target Maximum downtime in 30 days
99.9% 43 minutes 12 seconds
99.95% 21 minutes 36 seconds
99.99% 4 minutes 19 seconds

These figures are comparison aids, not predictions. The contractual result may differ because providers define eligible service time, outage events and exclusions differently.

3. Separate response, repair and restoration targets

An SLA may promise an initial response without promising resolution. A technician acknowledging a ticket within 15 minutes does not mean the server will be restored within 15 minutes.

Look for separate definitions of:

  • Acknowledgement time: Time until the provider confirms receipt.
  • Response time: Time until qualified personnel begin investigation.
  • Update interval: Frequency of progress reports during an incident.
  • Repair time: Time to repair or replace the failed component.
  • Restoration time: Time until the covered service becomes usable again.

Verify when each clock starts. It may begin when provider monitoring detects a fault, when you open a correctly classified ticket, or only after the provider confirms an eligible incident. Check when the clock can be paused—for example, while waiting for customer authorization, credentials or diagnostic results.

4. Inspect the hardware replacement commitment

Hardware replacement is a critical dedicated-server clause, but its wording often leaves the customer responsible for substantial recovery work.

Ask whether the target includes diagnosis or starts only after hardware failure has been confirmed. Determine whether it covers the complete server, individual disks, memory, power supplies, RAID controllers and network interfaces. Confirm whether equivalent replacement hardware may have different components or performance characteristics.

The SLA should make clear what happens after replacement. Physical repair may not include:

  • Reinstalling the operating system.
  • Recreating RAID arrays.
  • Restoring data or application configuration.
  • Reassigning custom network settings.
  • Validating application health.

For storage failures, ask about failed-drive retention, secure disposal and access to removed media. If your recovery depends on local disks, establish a separate backup outside the server before deployment.

5. Review every exclusion

Exclusions determine how much protection the SLA actually provides. Read them as carefully as the headline commitment.

Common areas to investigate include:

  • Scheduled and emergency maintenance.
  • Customer operating-system, firewall or routing changes.
  • Software failures and unsupported configurations.
  • Distributed denial-of-service attacks and mitigation events.
  • Traffic exceeding the purchased port or bandwidth limit.
  • Third-party carriers, control panels or software services.
  • Abuse-related suspension or non-payment.
  • Force majeure and events outside the provider’s control.
  • Failure to follow the required incident-reporting procedure.

An exclusion may be reasonable, but it should be specific. Broad language such as “network events beyond our control” needs clarification when the provider selects and manages the upstream connectivity.

6. Verify maintenance rules

Confirm how far in advance planned maintenance is announced, which communication channel is used, and whether there is a defined maintenance window. Check whether maintenance is unlimited or subject to duration and frequency constraints.

Emergency maintenance needs a separate definition. Ask who may classify work as an emergency and whether the provider must send a retrospective incident report. If all maintenance is excluded regardless of notice or duration, the published availability target may not reflect the interruption your users experience.

7. Check network and DDoS boundaries

A network SLA should identify the covered infrastructure. Determine whether it stops at the server’s switch port, includes the provider backbone, or extends to external transit and peering connections.

If DDoS protection is included, verify:

  • Whether mitigation is automatic or requires a support request.
  • Which attack types and traffic volumes are covered.
  • Whether the provider may null-route the server.
  • Whether time under mitigation or null routing counts as downtime.
  • Whether clean-traffic capacity is guaranteed.
  • Whether mitigation generates additional charges.

Do not interpret “DDoS protection included” as a guarantee that every application will remain reachable during every attack.

8. Evaluate support and escalation

Confirm that the support operation matches the urgency of the workload. Record the available channels, staffed hours, supported languages and authentication process for emergency requests.

Review the severity definitions. You should know whether a single unreachable production server qualifies for the highest priority and whether the provider can downgrade a ticket. Obtain an escalation path that extends beyond the first-line queue, including the process for requesting an incident manager during prolonged outages.

Check who is authorized to open and escalate tickets. Configure multiple contacts so that an absent employee does not block emergency support.

9. Calculate the real value of service credits

Service credits are usually a financial remedy, not compensation for the full business impact of an outage. Verify whether credits are automatic or must be claimed, and note the submission deadline.

Check the calculation basis. A percentage may apply only to the monthly recurring charge for the affected server—not setup fees, bandwidth, licenses, support add-ons or the total account invoice. Identify the maximum credit and whether unused credit expires.

Ask whether service credits are the exclusive remedy. For a critical workload or longer commitment, consider negotiating a termination right after repeated or prolonged breaches. A small account credit offers little protection if the service consistently fails operational requirements.

10. Establish evidence before launch

Once the server is delivered, create the evidence needed to detect outages and support a claim.

  1. Enable external monitoring from more than one independent location.
  2. Monitor the public service, server reachability and relevant network ports separately.
  3. Synchronize system clocks and preserve monitoring timestamps.
  4. Test the customer portal, ticket channel and escalation contacts.
  5. Verify out-of-band management or remote console access.
  6. Record the delivered hardware and network configuration.
  7. Test backups and restoration before production traffic is introduced.

Monitoring should reflect the SLA measurement point while also measuring what users experience. The two views answer different questions: whether the provider breached the contract and whether your service met its own availability objective.

11. Create an SLA claim runbook

Do not wait for an outage to learn the claims process. Write a short internal procedure containing:

  • The correct support channel and ticket severity.
  • The people authorized to communicate with the provider.
  • The evidence required for a claim.
  • The claim submission address or portal location.
  • The contractual claim deadline.
  • The internal owner responsible for following the case.

During an incident, preserve monitoring output, ticket numbers, status-page updates, provider messages and a timeline of actions. After restoration, request a root-cause report when the contract permits it and compare the provider’s recorded duration with your own evidence.

Common SLA mistakes

  • Buying the largest percentage: A higher target is not useful if major failure modes are excluded.
  • Confusing response with repair: Ticket acknowledgement does not restore service.
  • Ignoring the service boundary: Hardware, network, operating system and application may have different owners.
  • Assuming credits are automatic: Many agreements require a documented claim within a limited period.
  • Treating an SLA as redundancy: A contractual remedy cannot fail over traffic or restore data.
  • Relying on sales correspondence: Operational promises should appear in the controlling contract documents.
  • Skipping launch tests: Untested console access, backups and escalation contacts often fail when they are first needed.

Final recommendation

Choose the SLA whose measurable commitments cover your most important failure modes. For a replaceable stateless node, basic hardware and network commitments may be sufficient. For a single production server, prioritize clear restoration targets, 24/7 escalation, usable remote access and tested off-server backups. For workloads that cannot tolerate the contractual repair window, deploy redundant servers and treat the SLA as a remedy of last resort.

Before signing, make sure you can answer five questions without relying on a salesperson’s interpretation: what is covered, how downtime is measured, when the clock starts, what is excluded, and what happens when the provider misses the target. If any answer remains ambiguous, resolve it in writing before the server enters production.

Rate article
Add a comment