| Location | |
| Website | https://www.nscale.com |
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations — acting as the operational bridge between Support, DC Operations, and Engineering.
You will:
Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates.
6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.
Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.
Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.
Hands-on experience with RDMA fabrics — InfiniBand and/or RoCE — including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.
Slurm operations for large multi-GPU jobs — containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.
Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.
Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.
Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.
Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.
Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.
Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.
Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar).
Understanding of how data centres operate — servers, networks, storage, power, and cooling — ideally gained through an operational support background.
Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.
Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.
Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues.
OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar).
Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role.
Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production.
Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus.
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$120,000 - $170,000USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.