SRE Squad Lead
Posted 5 hours 6 minutes ago by Allscreens Nationwide Ltd
The Department for Business, Innovation, Science and Trade (BIST) helps businesses invest, innovate, export and grow across the UK. Digital, Data and Technology (DDaT) helps deliver this by designing, building and running the services, platforms and technology that businesses and colleagues rely on every day.
Our teams work on a wide range of services and technologies, including:
- digital services that help businesses access government support, guidance and opportunities
- data and AI products that support decision making and enable innovation
- internal tools and services that help colleagues work more effectively
- secure technology and platforms that support critical government services
- cyber security services that protect systems, data and users
By joining DDaT, you'll work on services used by businesses, colleagues and citizens across the UK. You'll be part of multidisciplinary teams focused on improving services, solving complex problems and delivering better outcomes for users.
We are committed to creating an inclusive and supportive workplace. In recognition of this, we were named Best Public Sector Employer at the Women in Tech Employer Awards 2025!
Senior Site Reliability Engineer Manager roleIf you would like to find out more about the role, the SRE team and what it's like to work at BIST, we are holding a Hiring Manager Q&A session for this role where you can virtually 'meet the team' on Tuesday 8th September at 12:30pm.
As a Senior Site Reliability Engineer Manager, you will lead the design, delivery and continuous improvement of reliable,scalableand secure platform services that underpin criticalBISTdigital products. Working closely with multidisciplinary, agile teams,you'llensure development teams have the tools and support they needfrom observability andmonitoringthrough to CI/CD pipelines,so services are resilient,performantand centred around user needs.You'llchampion good engineering practices, helping teams adopt service-level thinking using metrics, service-level indicators (SLIs),objectives(SLOs)and error budgets to support informed, collaborative decision making.
This is a people-focused leadership role where you will create an environment in which engineers can do their best work. You will line manage and develop a team of Site Reliability Engineers, supporting their growth and wellbeing, while also acting as a senior technical leader across the wider DDaT community.You'llwork in partnership with product managers, architects and delivery colleagues to shape platform strategy, improve reliability and reduce operational burden. Alongside hands on engineering, you will help build and scale our global platform, support live services through an on call rota, and lead improvements such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes.
Your responsibilities- Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive,collaborativeand high-performing team culture.
- Build strong working relationships with product,deliveryand architecture colleagues to ensure platform services meet business and user needs.
- Provide technical leadership across DevOps/SRE practices, guiding teams to adopt approaches that support reliability,sustainabilityand continuous improvement.
- Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community.
- Design, build andmaintainreliable,secureand scalable cloud-based infrastructure using infrastructure-as-code approaches.
- Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management.
- Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs)and error budgets in a pragmatic and user-focused way.
- Support the development and continual improvement of CI/CD pipelines to enable safe,frequentand low-risk delivery of changes.
- Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture.
- Ensure security, resilience and compliance considerations are understood and embedded into engineering practices.
- AWSandAzure
- GitHub Actions
- AWSCodePipelines/CodeBuild
- Terraform
- Docker
- Elastic Container Service (ECS)
- Elastic Container Registry (ECR)
- ElasticSearch/OpenSearch
- Python and Django framework
- PostgreSQL as a service (Amazon RDS)
- Datadog
- Logstash
- Redis/Elasticache
Company: Government Recruitment Service
Salary: £63,824 - £83,778
Hours: Full-time
Location: Birmingham
Job type: Permanent
Posting date: 25 Aug 2026
Closing date: 14 Sept 2026
Proud member of the Disability Confident employer scheme
About Disability Confident Disability Confident is a government scheme. It encourages employers to recruit and retain disabled people and those with long term health conditions.