it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

6 jobs found

Email me jobs like this
Refine Search
Current Search
senior site reliability engineer hpc
Senior HPC/GPU Systems Engineer
Nscale San Francisco, California
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
HPC/GPU Systems Engineer
Nscale Seattle, Washington
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details
HPC/GPU Systems Engineer
Nscale San Francisco, California
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details
Senior HPC/GPU Systems Engineer
Nscale Seattle, Washington
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
Senior HPC/GPU Systems Engineer
Nscale Houston, Texas
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
HPC/GPU Systems Engineer
Nscale Houston, Texas
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets - tickets, alerts, hardware faults, and customer issues - across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world. You will: Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope - with clean, evidence-rich handovers. Communicate technical detail clearly, specifically, and concisely - in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill. Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast. Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through. Seek feedback and invest in learning - this role is a deliberate pathway to Senior. Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting. What You'll Be Doing Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it. Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA. Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch. Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis. Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review. Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels. Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover. Participate in changes under peer review, learning risk assessment and backout practices in live customer environments. Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns). Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes. Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope. Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed. Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability. Take part in incident reviews as a contributor and help track preventative follow-ups in your scope. Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed. Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required. About You Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services). Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately. GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal. Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate. Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role. Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation. Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review. Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control. Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background. Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior. Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed. Nice to Have These are growth areas, not prerequisites: High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts. High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage-network troubleshooting. OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar). Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role. Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault. Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time. What We Can Offer You At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work. Equal Opportunities Statement We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities . click apply for full job details

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board