it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

2 jobs found

Email me jobs like this
Refine Search
Current Search
sr technical program manager networking capacity delivery
Senior Technical Product Manager, GPU Infrastructure
Nscale New York, New York
About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do. About the role Technical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform. You partner closely with engineering, design, research, and go-to-market teams to translate customer problems and operational realities into shippable product outcomes. As a Senior Technical Product Manager for Fleet Operations, you own the product strategy for the day 0-2+ operational software that runs our global GPU fleet - the systems that bring capacity online, keep it healthy, and restore it fast when things go wrong. You partner daily with Fleet Software engineering teams, SRE, and Support to turn operational pain into durable product: provisioning and bringup (day 0), testing and deployment (day 1), and the full lifecycle of monitoring, incident response, repair, RMA, firmware, and decommissioning (day 2+). You operate at team scope, owning a major product area and driving multi-quarter initiatives that directly move fleet availability, utilisation, and time-to-recover. Senior Technical Product Manager, Fleet Operations 1 What you'll be doing Own the strategy and roadmap for a significant Fleet Operations product area - e.g. provisioning and bring-up, fleet health and telemetry, incident and repair workflows, firmware and lifecycle management, or capacity and inventory. Lead multi-sprint, cross-functional initiatives from problem framing through rollout across live GPU clusters, working hand-in-hand with Fleet Software, SRE, data centre operations, and Support. Turn operational ambiguity into product: shadow on-call rotations, ride along with support and repair workflows, and translate recurring toil into tooling, automation, and platform capabilities. Define the metrics that matter for a GPU fleet - availability, utilisation, MTTR, time-to-bring-up, hardware failure rates, support ticket deflection - and drive the roadmap against them. Partner with engineering on architecture and trade-offs for systems that span bare metal, orchestration, observability, and control planes. Drive incident reviews and postmortems into product commitments; close the loop so the same class of issue doesn't recur. Mentor junior product managers and raise the quality bar for PRDs, reviews, and product decisions across the team. Represent Fleet Operations in planning, reviews, and leadership updates. What you need 5-8 years of product management experience in software or technology, with a track record of owning significant product areas in infrastructure, platform, or operations-facing products. Strong technical fluency in large-scale systems: you can lead discussions with engineering on architecture, trade-offs, and feasibility across provisioning, orchestration, observability, and control-plane design Experience building products for operators - SREs, NOC/support teams, data centre technicians, or similar - and a genuine appetite for understanding their workflows. Demonstrated ability to move from an ambiguous operational problem space to shipped product outcomes that measurably improve reliability, efficiency, or time-to-recover. Experience mentoring or informally leading peers. Excellent written and verbal communication; you can make complex product decisions legible to engineers, operators, and executives alike. Experience with data centre networking technologies, including high-performance GPU interconnects such as InfiniBand and RoCE (RDMA over Converged Ethernet), and an understanding of how backend (east-west/compute) and frontend (north-south/storage and management) network fabrics are designed and operated at scale. Familiarity with WAN, edge, and global backbone architectures - including how multi-site connectivity, peering, and traffic engineering support a globally distributed GPU fleet. Experience partnering with network engineering teams on fabric health, congestion monitoring, and link-level failure workflows, ideally in environments where network performance directly impacts training or inference workloads. Nice to haves Degree in computer science, engineering, or a related field, or prior experience as an engineer or SRE. Hands-on background in cloud infrastructure, bare-metal provisioning, fleet or hardware lifecycle management, observability/monitoring platforms, or incident management tooling. Experience with bare-metal provisioning systems such as OpenStack Ironic (or equivalents like MAAS, Tinkerbell, or in-house provisioning stacks). Experience with DCIM tools such as NetBox (or equivalents like Device42 or Nautobot) for inventory, cabling, and rack/asset management. Experience with ITSM and ticketing platforms such as Jira Service Management (or equivalents like ServiceNow, Zendesk, or Freshservice) for support, incident, and RMA workflows. Experience with observability and monitoring platforms such as Grafana, Prometheus, Datadog, or equivalents - ideally including defining SLOs, dashboards, and alerting for large fleets. Familiarity with GPU or accelerated compute environments, data centre operations, or hyperscaler-style fleet management. Experience operating in high-growth or early-stage environments where the product is being built alongside the fleet itself Join Nscale as we build a world-class AI cloud platform. If you're excited about owning the software that keeps a global GPU fleet running - and raising the bar for the team around you - we'd love to hear from you! At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $200,000-$280,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.
09/23/2026
Full time
About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do. About the role Technical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform. You partner closely with engineering, design, research, and go-to-market teams to translate customer problems and operational realities into shippable product outcomes. As a Senior Technical Product Manager for Fleet Operations, you own the product strategy for the day 0-2+ operational software that runs our global GPU fleet - the systems that bring capacity online, keep it healthy, and restore it fast when things go wrong. You partner daily with Fleet Software engineering teams, SRE, and Support to turn operational pain into durable product: provisioning and bringup (day 0), testing and deployment (day 1), and the full lifecycle of monitoring, incident response, repair, RMA, firmware, and decommissioning (day 2+). You operate at team scope, owning a major product area and driving multi-quarter initiatives that directly move fleet availability, utilisation, and time-to-recover. Senior Technical Product Manager, Fleet Operations 1 What you'll be doing Own the strategy and roadmap for a significant Fleet Operations product area - e.g. provisioning and bring-up, fleet health and telemetry, incident and repair workflows, firmware and lifecycle management, or capacity and inventory. Lead multi-sprint, cross-functional initiatives from problem framing through rollout across live GPU clusters, working hand-in-hand with Fleet Software, SRE, data centre operations, and Support. Turn operational ambiguity into product: shadow on-call rotations, ride along with support and repair workflows, and translate recurring toil into tooling, automation, and platform capabilities. Define the metrics that matter for a GPU fleet - availability, utilisation, MTTR, time-to-bring-up, hardware failure rates, support ticket deflection - and drive the roadmap against them. Partner with engineering on architecture and trade-offs for systems that span bare metal, orchestration, observability, and control planes. Drive incident reviews and postmortems into product commitments; close the loop so the same class of issue doesn't recur. Mentor junior product managers and raise the quality bar for PRDs, reviews, and product decisions across the team. Represent Fleet Operations in planning, reviews, and leadership updates. What you need 5-8 years of product management experience in software or technology, with a track record of owning significant product areas in infrastructure, platform, or operations-facing products. Strong technical fluency in large-scale systems: you can lead discussions with engineering on architecture, trade-offs, and feasibility across provisioning, orchestration, observability, and control-plane design Experience building products for operators - SREs, NOC/support teams, data centre technicians, or similar - and a genuine appetite for understanding their workflows. Demonstrated ability to move from an ambiguous operational problem space to shipped product outcomes that measurably improve reliability, efficiency, or time-to-recover. Experience mentoring or informally leading peers. Excellent written and verbal communication; you can make complex product decisions legible to engineers, operators, and executives alike. Experience with data centre networking technologies, including high-performance GPU interconnects such as InfiniBand and RoCE (RDMA over Converged Ethernet), and an understanding of how backend (east-west/compute) and frontend (north-south/storage and management) network fabrics are designed and operated at scale. Familiarity with WAN, edge, and global backbone architectures - including how multi-site connectivity, peering, and traffic engineering support a globally distributed GPU fleet. Experience partnering with network engineering teams on fabric health, congestion monitoring, and link-level failure workflows, ideally in environments where network performance directly impacts training or inference workloads. Nice to haves Degree in computer science, engineering, or a related field, or prior experience as an engineer or SRE. Hands-on background in cloud infrastructure, bare-metal provisioning, fleet or hardware lifecycle management, observability/monitoring platforms, or incident management tooling. Experience with bare-metal provisioning systems such as OpenStack Ironic (or equivalents like MAAS, Tinkerbell, or in-house provisioning stacks). Experience with DCIM tools such as NetBox (or equivalents like Device42 or Nautobot) for inventory, cabling, and rack/asset management. Experience with ITSM and ticketing platforms such as Jira Service Management (or equivalents like ServiceNow, Zendesk, or Freshservice) for support, incident, and RMA workflows. Experience with observability and monitoring platforms such as Grafana, Prometheus, Datadog, or equivalents - ideally including defining SLOs, dashboards, and alerting for large fleets. Familiarity with GPU or accelerated compute environments, data centre operations, or hyperscaler-style fleet management. Experience operating in high-growth or early-stage environments where the product is being built alongside the fleet itself Join Nscale as we build a world-class AI cloud platform. If you're excited about owning the software that keeps a global GPU fleet running - and raising the bar for the team around you - we'd love to hear from you! At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $200,000-$280,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.
Senior Technical Product Manager, Observability
Nscale New York, New York
About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do. About the role Technical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform, partnering with engineering, design, and go-to-market to turn customer and operational problems into shippable outcomes. As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline that scrapes data from physical infrastructure, the aggregation and storage layer, and the observability surfaces (logs, metrics, and traces) that enable fleet management, incident response, and alerting at scale. You partner daily with Fleet Software, Network Engineering, Data Centre Operations, and customer teams to make fleet health visible, actionable, and reliable as Nscale scales from a handful of deployments to a globally distributed fleet. What you'll be doing Own the roadmap for Nscale's observability platform: the telemetry pipeline, log and metrics aggregation, trace collection, and customer facing APIs and dashboards that surface fleet health to customers and operators. Define how logs, metrics, and traces are captured from physical infrastructure, aggregated, and surfaced through the observability platform to enable customers to manage their fleet and handle incidents. Own alerting strategy and optimisation: define what matters, reduce noise, and ensure the right signal reaches the right person at the right time. Capture and prioritise new telemetry requirements as the fleet scales, working with engineering to extend coverage across new hardware, sites, and deployment types. Shadow incident reviews and site operations to turn recurring manual effort and visibility gaps into platform capabilities. Define and drive the metrics that matter: alert signal-to-noise ratio, time-to-detect, time-to-resolve, telemetry coverage, and platform reliability. Mentor junior PMs and raise the bar for PRDs, reviews, and product decisions across the team. What you need 5-8 years in product management, with a track record owning significant areas in observability, infrastructure, or operations-facing products. Demonstrated experience building observability stacks: you have owned a product that captures and surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved. Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry. Experience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch deployment pipelines. Experience building for operators and delivery teams (design engineers, project controllers, PMs, SREs, DC technicians) and a genuine appetite for their workflows. Strong technical fluency: you can lead architecture and trade-off discussions across telemetry pipelines, time-series storage, alerting systems, and observability integrations. A record of moving ambiguous operational problems to shipped outcomes that measurably improve visibility, incident response, or fleet reliability. Excellent written and verbal communication across engineers, operators, and executives. Nice to haves Broader observability problem domain experience across different toolsets beyond the above stack. Familiarity with bare-metal provisioning tools (OpenStack Ironic, MAAS, or similar) or network automation tooling (NetBox, Nautobot, or similar).Degree in CS or engineering, or prior experience as an engineer, SRE, or infrastructure operator. Familiarity with GPU or accelerated compute infrastructure, data centre operations, or hyperscaler-style deployment at scale. ITSM: Jira Service Management, ServiceNow, Zendesk, or Freshservice. Experience in high-growth environments where the product is being built alongside the fleet it monitors. Join Nscale as we build a world-class AI cloud platform. If you're excited about owning the software that turns contracts into live GPU capacity, we'd love to hear from you! At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $220,000-$260,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.
09/23/2026
Full time
About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do. About the role Technical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform, partnering with engineering, design, and go-to-market to turn customer and operational problems into shippable outcomes. As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline that scrapes data from physical infrastructure, the aggregation and storage layer, and the observability surfaces (logs, metrics, and traces) that enable fleet management, incident response, and alerting at scale. You partner daily with Fleet Software, Network Engineering, Data Centre Operations, and customer teams to make fleet health visible, actionable, and reliable as Nscale scales from a handful of deployments to a globally distributed fleet. What you'll be doing Own the roadmap for Nscale's observability platform: the telemetry pipeline, log and metrics aggregation, trace collection, and customer facing APIs and dashboards that surface fleet health to customers and operators. Define how logs, metrics, and traces are captured from physical infrastructure, aggregated, and surfaced through the observability platform to enable customers to manage their fleet and handle incidents. Own alerting strategy and optimisation: define what matters, reduce noise, and ensure the right signal reaches the right person at the right time. Capture and prioritise new telemetry requirements as the fleet scales, working with engineering to extend coverage across new hardware, sites, and deployment types. Shadow incident reviews and site operations to turn recurring manual effort and visibility gaps into platform capabilities. Define and drive the metrics that matter: alert signal-to-noise ratio, time-to-detect, time-to-resolve, telemetry coverage, and platform reliability. Mentor junior PMs and raise the bar for PRDs, reviews, and product decisions across the team. What you need 5-8 years in product management, with a track record owning significant areas in observability, infrastructure, or operations-facing products. Demonstrated experience building observability stacks: you have owned a product that captures and surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved. Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry. Experience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch deployment pipelines. Experience building for operators and delivery teams (design engineers, project controllers, PMs, SREs, DC technicians) and a genuine appetite for their workflows. Strong technical fluency: you can lead architecture and trade-off discussions across telemetry pipelines, time-series storage, alerting systems, and observability integrations. A record of moving ambiguous operational problems to shipped outcomes that measurably improve visibility, incident response, or fleet reliability. Excellent written and verbal communication across engineers, operators, and executives. Nice to haves Broader observability problem domain experience across different toolsets beyond the above stack. Familiarity with bare-metal provisioning tools (OpenStack Ironic, MAAS, or similar) or network automation tooling (NetBox, Nautobot, or similar).Degree in CS or engineering, or prior experience as an engineer, SRE, or infrastructure operator. Familiarity with GPU or accelerated compute infrastructure, data centre operations, or hyperscaler-style deployment at scale. ITSM: Jira Service Management, ServiceNow, Zendesk, or Freshservice. Experience in high-growth environments where the product is being built alongside the fleet it monitors. Join Nscale as we build a world-class AI cloud platform. If you're excited about owning the software that turns contracts into live GPU capacity, we'd love to hear from you! At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $220,000-$260,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board