This is an experimental service. Give feedback (opens in new tab)

English |

GPU / Compute Systems Engineer

Company:Media Stream AI Limited
Salary:£54,000 - £60,000
Hours:Full-time
Location:Dundee, DD2 1UR
Working pattern:On-site
Job type:Permanent
Posting date:19 Aug 2026
Closing date:18 Sept 2026
Apply for this job

Summary

The GPU / Compute Systems Engineer will take ownership of the GPU compute estate at the MSAI Scotland data centre, ensuring the reliability, performance and availability of high-density GPU infrastructure supporting AI, HPC and GPUaaS workloads.

The role will oversee NVIDIA B300 and H200 GPU nodes, server infrastructure, high-performance networking fabrics and cluster orchestration, including Kubernetes and Slurm. The engineer will also support GPUaaS customer deployments, cluster provisioning, diagnostics, burn-in testing, hardware RMA and ongoing client cluster health.

Duties

- Own the day-to-day operation and performance of the GPU compute estate, including NVIDIA B300 and H200 infrastructure.

- Install, configure, test and maintain GPU servers, compute nodes and associated hardware.

- Perform GPU burn-in, validation and acceptance testing following deployment or maintenance.

- Monitor GPU, CPU, memory, storage, networking and system health across the compute estate.

- Diagnose and resolve GPU, server, driver, firmware and operating-system issues.

- Manage hardware failures, including RMA processes, component replacement and vendor escalation.

- Maintain GPU node firmware, drivers, operating systems and supporting software.

- Manage and troubleshoot the RoCE and/or InfiniBand high-performance compute fabric.

- Monitor network performance, bandwidth, latency and GPU-to-GPU communication.

- Support configuration and optimisation of RDMA, RoCE and InfiniBand environments.

- Deploy, configure and maintain Kubernetes clusters supporting GPU workloads.

- Deploy, configure and maintain Slurm/HPC workload management environments.

- Manage GPU scheduling, resource allocation and cluster utilisation.

- Support GPUaaS client deployments, including provisioning, configuration and cluster onboarding.

- Troubleshoot customer workloads and infrastructure issues within agreed support and escalation procedures.

- Monitor client cluster health, GPU utilisation and system performance.

- Support the deployment and configuration of containerised AI/HPC workloads.

- Maintain system images, configurations, automation scripts and deployment tooling.

- Automate repetitive operational tasks using Python, Bash or similar scripting tools.

- Support capacity planning for GPU compute, networking, storage and associated infrastructure.

- Work closely with the NOC, Network Operations, BMS/DCIM, Critical Facilities and Data Centre teams.

- Coordinate with NVIDIA, server manufacturers and specialist technology partners on technical issues and hardware support.

- Maintain accurate asset registers, cluster documentation, network diagrams and operating procedures.

- Develop and maintain runbooks, SOPs and troubleshooting guides.

- Support incident management, root-cause analysis and post-incident reviews.

- Participate in planned maintenance, upgrades and cluster expansion activities.

- Ensure changes to production infrastructure are managed through the appropriate change-control process.

- Participate in an on-call rota for critical GPU, compute and cluster incidents.

Apply for this job

Related jobs

Network Operations (NOC) Engineer

£41,000 to £47,000 per year

Media Stream AI Limited

Dundee

On-sitePermanentFull time

Data Centre Technician

£33,000 to £40,000 per year

Media Stream AI Limited

Dundee

On-sitePermanentFull time

Data Centre Manager (Site Lead)

£70,000 to £81,000 per year

Media Stream AI Limited

Dundee

On-sitePermanentFull time
Browse more jobs