Leave us your email address and we'll send you all the new jobs according to your preferences.

Service Manager - Site Reliability Engineering

Posted 1 hour 55 minutes ago by Jobtailor

Permanent
Full Time
Engineering Jobs
Northern Ireland, United Kingdom
Job Description
  • Own the overall health, reliability, availability, and observability of critical business applications and technology services
  • Establish, monitor, and report on SLAs, SLOs, Error Budgets, availability, performance, and operational KPIs
  • Lead Major Incident Management activities, coordinating cross-functional teams during outages and ensuring rapid service restoration and RCA completion
  • Drive Problem Management by identifying recurring issues, analyzing systemic failures, implementing corrective actions, and reducing operational risk
  • Govern Change and Release Management processes, production readiness reviews, maintenance, and deployment activities
  • Ensure effective observability through monitoring, alerting, logging, dashboards, and operational reporting
  • Promote automation and operational efficiency through Infrastructure as Code, DevOps, self-healing, and auto-remediation
  • Partner with engineering, platform, infrastructure, security, and business teams to improve resilience, scalability, stability, and customer experience
  • Act as operational liaison for business stakeholders and vendors, conducting service reviews, communicating risks, and managing escalations
  • Ensure adherence to ITIL-based processes, governance standards, compliance requirements, audit obligations, and operational documentation standards
  • Lead continuous service improvement initiatives to reduce MTTR, increase stability, improve customer satisfaction, and enhance operational maturity
  • Lead operational governance ceremonies, service reviews, incident and problem reviews, change governance forums, readiness assessments, stakeholder communications, and executive service reporting
  • Create, maintain, review, and ensure compliance with runbooks, SOPs, knowledge articles, audit evidence, disaster recovery procedures, and certification artifacts
  • Provide leadership across incident, problem, change, release, and service management disciplines
  • Lead, mentor, coach, and develop a team of Service Analysts
  • Provide training, knowledge sharing, cross-training, and professional development support
  • Support operational risk management by identifying vulnerabilities, assessing impacts, and developing mitigation plans
  • Collaborate on service reliability, scalability, security, production readiness, modernization, and transformation initiatives
  • Drive strategic operational improvements supporting service quality, customer experience, business alignment, and long-term sustainability
Requirements
  • Legal right to work in the UK; Allstate is not providing sponsorship for this vacancy
  • Minimum of 4 years of experience supporting or improving enterprise technology services, infrastructure environments, platform operations, Site Reliability Engineering, IT Operations, or Service Management disciplines (or equivalent)
  • Minimum of 2 years leading and mentoring teams within service reliability, availability, performance, and/or operational governance in a large enterprise environment
  • Experience leading Major Incident response activities, coordinating cross-functional teams, and driving RCA efforts and corrective actions
  • Experience implementing operational improvements, automation initiatives, risk-reduction measures, and continuous service improvement programs
  • Strong understanding of observability practices, including monitoring, alerting, logging, dashboards, and operational reporting
  • Knowledge of ITIL principles and IT Service Management processes
  • Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles
  • Experience working within an SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment
  • Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint, or related identity technologies
  • Experience managing SLAs and operational KPIs (or equivalent)
  • Knowledge of networking technologies, firewalls, DNS, load balancing, and enterprise infrastructure concepts
  • Experience leading or mentoring a team of engineers (or equivalent)
  • Experience with Infrastructure as Code, automation frameworks, cloud-native operational practices, self-healing systems, or auto-remediation
  • Familiarity with API management platforms and enterprise service-integration technologies
  • Experience supporting messaging or event-streaming platforms such as Kafka
  • Knowledge of middleware or integration technologies such as TIBCO or comparable enterprise platforms
  • Experience supporting Microsoft Azure, Amazon Web Services, or Google Cloud Platform
  • Experience using ServiceNow or a comparable ITSM platform
  • Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices
  • Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications
  • Experience leading operational maturity assessments, service governance programs, or production readiness reviews
Core Competencies

Demonstrates expertise in Service Reliability Engineering, Incident Management, and ITIL-based processes, with a strong focus on operational governance, automation, and continuous service improvement. Proven ability to lead cross-functional teams, manage SLAs, and enhance customer experience through effective communication and collaboration.

Highest-signal resume keywords
  • Service Reliability Engineering
  • Incident Management
  • ITIL Principles
  • Infrastructure as Code
  • Operational Governance
Hard Skills
  • Operational Documentation
  • Monitoring
  • Alerting
  • Logging
  • Dashboards
  • Cloud Operations
  • Automation Frameworks
  • Identity and Access Management
  • Networking Technologies
  • Event-Streaming Platforms
Soft Skills
  • Leadership
  • Mentoring
  • Collaboration
  • Communication
  • Problem-Solving
Certifications & Qualifications
  • ITIL
  • SRE
  • Cloud Certifications
  • Kubernetes
  • Security Certifications
Industry Keywords
  • Operational Resilience
  • Business Continuity
  • Disaster Recovery
  • Compliance
  • Risk Management
Tools & Technologies
  • ServiceNow
  • Microsoft Azure
  • Amazon Web Services
  • Google Cloud Platform
  • Kafka
  • TIBCO
Email this Job