← All industries

IT SOP Templates

IT runbooks rot fast. A step that worked against last quarter's infrastructure silently stops working, and nobody notices until an incident forces someone to follow it live. These IT SOP templates cover the procedures that carry real risk when they're wrong: access provisioning, incident response, backups, and deployment, each with the approval and rollback steps a junior engineer needs on call at 2 a.m. Record your screen while performing the procedure and ReccordSOP generates the SOP with screenshots of every step. Drift detection flags the moment the documented runbook stops matching what the team actually runs.

Download all 5 templates

Every template on this page in one file, as Markdown or Word. Free, no signup, ready to adapt in Notion, Google Docs, or ReccordSOP.

User access provisioning and deprovisioning

An orphaned account is access a former contractor or terminated employee still has because nobody closed the loop. This procedure closes it in both directions.

  1. 1.Verify manager or ticket-based approval before creating any account
  2. 2.Create the account in the identity provider (Okta, Entra ID, Google Workspace)
  3. 3.Assign role-based access groups, least privilege by default
  4. 4.Enforce MFA enrollment before first login
  5. 5.Send credentials via a secure, expiring channel, never plaintext email
  6. 6.Log the grant in the access register with requester and approver
  7. 7.On termination, revoke access within the same business day and confirm in the register

Incident response

An incident without a defined severity and owner turns into five engineers debugging in a Slack thread while the outage runs. This procedure assigns both immediately.

  1. 1.Acknowledge the alert and assess severity against the defined SEV1 to SEV4 matrix
  2. 2.Page the on-call engineer for SEV1/SEV2, notify the channel for SEV3/SEV4
  3. 3.Declare an incident commander and open a dedicated incident channel
  4. 4.Identify root cause and apply a fix or roll back to the last known-good deploy
  5. 5.Update the status page and notify affected customers per the communication SLA
  6. 6.Monitor for recurrence for the defined window before closing the incident
  7. 7.Write the post-mortem within 48 hours, no blame, with an owner for each follow-up

Backup and disaster recovery

A backup nobody has tested restoring is a backup that might not work. This procedure catches that before the day you actually need it.

  1. 1.Verify the backup schedule and retention policy against the data-loss tolerance
  2. 2.Confirm the last automated backup completed and check the integrity log
  3. 3.Run a test restore to a staging environment on a defined cadence
  4. 4.Time the restore and document actual recovery time against the target RTO
  5. 5.Verify data integrity and record count after the restore
  6. 6.Document any gap between actual and target recovery time
  7. 7.Escalate unresolved gaps to the infrastructure owner

Server deployment

  1. 1.Provision the instance from the approved template or IaC module
  2. 2.Configure networking, security groups, and firewall rules
  3. 3.Install required packages and pin dependency versions
  4. 4.Deploy application code through the CI/CD pipeline, never manually to production
  5. 5.Run health checks and smoke tests before adding to the load balancer
  6. 6.Monitor error rate and latency for the defined bake-in window
  7. 7.Roll back automatically if error rate crosses the alert threshold

Change management approval

  1. 1.Submit a change request with scope, risk level, and rollback plan
  2. 2.Route for approval per risk tier: low-risk to team lead, high-risk to change board
  3. 3.Schedule the change window and notify affected teams
  4. 4.Execute the change with a named implementer and a named observer
  5. 5.Verify success criteria immediately after the change
  6. 6.Roll back if success criteria aren't met within the window
  7. 7.Close the change record with the outcome and any deviation from plan

Frequently asked questions

What is an IT SOP?

An IT SOP is a written procedure for a specific technical task, covering the trigger, the exact commands or systems involved, the approval required, and the rollback plan if something goes wrong. It turns a runbook that lives in one engineer's head into something any on-call engineer can execute correctly under pressure, whether it's provisioning access, responding to an incident, or restoring from backup.

What should an IT SOP include?

The trigger (an alert, a ticket, a hire date), numbered steps naming the exact tool or command, the role with approval authority, and a rollback or escalation path if the procedure fails partway through. For anything touching production, a defined severity or risk tier changes who has to sign off before it runs.

How is an IT SOP different from a runbook?

In practice they overlap. A runbook usually documents diagnosis: the branching logic for what to check when something breaks. An SOP documents execution: the fixed steps for a procedure that should run the same way every time, like provisioning access or restoring a backup. Incident response usually needs both.

How often should IT SOPs be reviewed?

After every infrastructure or tooling change, and at minimum quarterly for anything referencing a specific UI, since vendors ship interface changes that quietly break documented steps. Review immediately after any incident where the runbook didn't match what actually happened, since that gap is exactly what caused the delay.

Related guide

SOP Drift-Risk Checker: Which Lines Go Stale First

A free tool for spotting which of your documented procedures are most likely to have already drifted from what the team actually runs.

Related guide

SOP Templates by Industry

Browse SOP templates for every department: finance, sales, marketing, HR, and support, alongside IT.

Stop writing SOPs manually

Record your screen while performing any process. AI generates the SOP. Drift detection keeps it accurate.

Start for free