SRE Squad Lead
| Company: | Government Recruitment Service |
|---|---|
| Salary: | £63,824 - £83,778 |
| Hours: | Full-time |
| Location: | Edinburgh |
| Job type: | Permanent |
| Posting date: | 25 Aug 2026 |
| Closing date: | 14 Sept 2026 |
Summary
If you would like to find out more about the role, the SRE team and what it’s like to work at BIST, we are holding a Hiring Manager Q&A session for this role where you can virtually 'meet the team' on Tuesday 8th September at 12:30pm. Please click here to book your spot.
About us
The Department for Business, Innovation, Science and Trade (BIST) helps businesses invest, innovate, export and grow across the UK. Digital, Data and Technology (DDaT) helps deliver this by designing, building and running the services, platforms and technology that businesses and colleagues rely on every day.
Our teams work on a wide range of services and technologies, including:
- digital services that help businesses access government support, guidance and opportunities
- data and AI products that support decision making and enable innovation
- internal tools and services that help colleagues work more effectively
- secure technology and platforms that support critical government services
- cyber security services that protect systems, data and users
By joining DDaT, you'll work on services used by businesses, colleagues and citizens across the UK. You'll be part of multidisciplinary teams focused on improving services, solving complex problems and delivering better outcomes for users.
We are committed to creating an inclusive and supportive workplace. In recognition of this, we were named Best Public Sector Employer at the Women in Tech Employer Awards 2025!
As a Senior Site Reliability Engineer Manager, you will lead the design, delivery and continuous improvement of reliable, scalable and secure platform services that underpin critical BIST digital products. Working closely with multidisciplinary, agile teams, you’ll ensure development teams have the tools and support they need from observability and monitoring through to CI/CD pipelines, so services are resilient, performant and centred around user needs. You’ll champion good engineering practices, helping teams adopt service-level thinking using metrics, service-level indicators (SLIs), objectives (SLOs) and error budgets to support informed, collaborative decision making.
This is a people-focused leadership role where you will create an environment in which engineers can do their best work. You will line manage and develop a team of Site Reliability Engineers, supporting their growth and wellbeing, while also acting as a senior technical leader across the wider DDaT community. You’ll work in partnership with product managers, architects and delivery colleagues to shape platform strategy, improve reliability and reduce operational burden. Alongside hands-on engineering, you will help build and scale our global platform, support live services through an on-call rota, and lead improvements such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes.
You will:
- Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive, collaborative and high-performing team culture.
- Build strong working relationships with product, delivery and architecture colleagues to ensure platform services meet business and user needs.
- Provide technical leadership across DevOps/SRE practices, guiding teams to adopt approaches that support reliability, sustainability and continuous improvement.
- Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community.
- Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches.
- Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management.
- Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way.
- Support the development and continual improvement of CI/CD pipelines to enable safe, frequent and low-risk delivery of changes.
- Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture.
- Ensure security, resilience and compliance considerations are understood and embedded into engineering practices.
What tech will you be using?
- AWS and Azure
- GitHub Actions, AWS CodePipelines/CodeBuild
- Terraform
- Docker, Elastic Container Service (ECS) and Elastic Container Registry (ECR)
- ElasticSearch/OpenSearch
- Python and Django framework
- PostgreSQL as a service (Amazon RDS)
- Datadog, Logstash
- Redis/Elasticache
Proud member of the Disability Confident employer scheme
About Disability Confident
Related jobs
Senior Site Reliability Engineer – Python
£63,824 to £83,778 per year
Government Recruitment Service
Edinburgh
PermanentFull timeHead of User Research
£74,639 to £96,980 per year
Government Recruitment Service
Edinburgh
PermanentFull timeOpenShift Engineer (Kubernetes/OpenShift)
£48,987 to £54,430 per year
Lloyds Banking Group
Edinburgh, Edinburgh
HybridPermanentFull time