GPU / Compute Systems Engineer
| Company: | Media Stream AI Limited |
|---|---|
| Salary: | £54,000 - £60,000 |
| Hours: | Full-time |
| Location: | Dundee, DD2 1UR |
| Working pattern: | On-site |
| Job type: | Permanent |
| Posting date: | 19 Aug 2026 |
| Closing date: | 18 Sept 2026 |
Summary
The GPU / Compute Systems Engineer will take ownership of the GPU compute estate at the MSAI Scotland data centre, ensuring the reliability, performance and availability of high-density GPU infrastructure supporting AI, HPC and GPUaaS workloads.
The role will oversee NVIDIA B300 and H200 GPU nodes, server infrastructure, high-performance networking fabrics and cluster orchestration, including Kubernetes and Slurm. The engineer will also support GPUaaS customer deployments, cluster provisioning, diagnostics, burn-in testing, hardware RMA and ongoing client cluster health.
Duties
- Own the day-to-day operation and performance of the GPU compute estate, including NVIDIA B300 and H200 infrastructure.
- Install, configure, test and maintain GPU servers, compute nodes and associated hardware.
- Perform GPU burn-in, validation and acceptance testing following deployment or maintenance.
- Monitor GPU, CPU, memory, storage, networking and system health across the compute estate.
- Diagnose and resolve GPU, server, driver, firmware and operating-system issues.
- Manage hardware failures, including RMA processes, component replacement and vendor escalation.
- Maintain GPU node firmware, drivers, operating systems and supporting software.
- Manage and troubleshoot the RoCE and/or InfiniBand high-performance compute fabric.
- Monitor network performance, bandwidth, latency and GPU-to-GPU communication.
- Support configuration and optimisation of RDMA, RoCE and InfiniBand environments.
- Deploy, configure and maintain Kubernetes clusters supporting GPU workloads.
- Deploy, configure and maintain Slurm/HPC workload management environments.
- Manage GPU scheduling, resource allocation and cluster utilisation.
- Support GPUaaS client deployments, including provisioning, configuration and cluster onboarding.
- Troubleshoot customer workloads and infrastructure issues within agreed support and escalation procedures.
- Monitor client cluster health, GPU utilisation and system performance.
- Support the deployment and configuration of containerised AI/HPC workloads.
- Maintain system images, configurations, automation scripts and deployment tooling.
- Automate repetitive operational tasks using Python, Bash or similar scripting tools.
- Support capacity planning for GPU compute, networking, storage and associated infrastructure.
- Work closely with the NOC, Network Operations, BMS/DCIM, Critical Facilities and Data Centre teams.
- Coordinate with NVIDIA, server manufacturers and specialist technology partners on technical issues and hardware support.
- Maintain accurate asset registers, cluster documentation, network diagrams and operating procedures.
- Develop and maintain runbooks, SOPs and troubleshooting guides.
- Support incident management, root-cause analysis and post-incident reviews.
- Participate in planned maintenance, upgrades and cluster expansion activities.
- Ensure changes to production infrastructure are managed through the appropriate change-control process.
- Participate in an on-call rota for critical GPU, compute and cluster incidents.
Related jobs
Network Operations (NOC) Engineer
£41,000 to £47,000 per year
Media Stream AI Limited
Dundee
On-sitePermanentFull timeData Centre Technician
£33,000 to £40,000 per year
Media Stream AI Limited
Dundee
On-sitePermanentFull timeData Centre Manager (Site Lead)
£70,000 to £81,000 per year
Media Stream AI Limited
Dundee
On-sitePermanentFull time