Req ID: 383115 NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Senior Principal Cloud Architect - AWS & Service Activation Hybrid Irving, TX or Charlotte, NC to join our team in Irving, Texas (US-TX), United States (US). Please Note: This hybrid role requires onsite client support 2-3 days weekly in either Irving, TX or Charlotte, NC. Candidates must be located within commuting distance of one of these offices. Applications from candidates outside these areas may not be considered. Position Summary We are seeking a Principal Cloud Engineer to architect design, build, and operate a secure, scalable, and highly automated multi-cloud platform leveraging Google Cloud Platform (GCP), Amazon Web Services (AWS), and Microsoft Azure. As a key technical leader, you will drive the evolution of our multi-cloud engineering, enabling resilient, cloud-native services that support mission-critical enterprise workloads. This role requires deep expertise in cloud architecture, infrastructure as code (IaC), and distributed systems engineering, with a strong focus on Terraform-driven automation and principal engineering standardization. The ideal candidate brings strong Kubernetes, IAM, networking, and policy management experience, along with a disciplined approach to security, reliability, and cost optimization. Key Responsibilities: Principal Engineering, Design & Architecture Architect, engineer, and implement scalable, secure, and resilient cloud solutions across GCP and AWS, with Azure as a plus. Engineer multi-region, highly available architectures aligned with enterprise standards and best practices. Infrastructure Automation & IaC Lead engineering, adoption, and development of Terraform-based infrastructure as code, ensuring consistency, compliance, and reusability. Build and maintain modular Terraform frameworks for provisioning cloud infrastructure. Integrate IaC into CI/CD pipelines to enable fully automated provisioning and deployment workflows. Containerization & Kubernetes Platforms Direct knowledge and engineering leveraging Kubernetes platforms (EKS, GKE) and associated ecosystems. Implementation of container orchestration best practices including scaling, observability, and security. Identity, Access & Security Governance Experienced working knowledge architecting and designing robust IAM strategies across GCP and AWS including least-privilege access models. Implement policy frameworks for access control, resource governance, and regulatory compliance. Networking & Connectivity Direct knowledge and experience architecting, designing, and engineering cloud networking solutions including VPC/VNet design, routing, peering, VPNs, and private connectivity. Ensure secure, high-performance connectivity between cloud environments and on-premises systems. Policy Management & Governance Deep understanding of policy management, policy-as-code frameworks to ensure consistent governance of resources, tagging strategies, and compliance controls. Collaboration with governance, policy, and security teams to ensure cloud environments meet enterprise and regulatory standards. Modernization & Mentorship Champion cloud-native and principal engineering engineering best practices across the organization. Drive the evolution of principal engineering practice by incorporating infrastructure-as-code (IaC) principles. Mentor junior engineers and provide technical leadership for complex cloud initiatives. Serve as an escalation point for critical cloud infrastructure and principal engineering issues. Ideal experience to have for this role: Cloud Platforms & Architecture Extensive firsthand experience with AWS and GCP in large-scale enterprise environments. Strong expertise in cloud architecture, including high availability, disaster recovery, and multi-region design. Deep understanding of cloud-native services and distributed systems principles. AWSControl tower and landing Zones. Infrastructure as Code & Automation Advanced proficiency in Terraform for infrastructure provisioning and lifecycle management. Experience designing reusable modules and managing Terraform state at scale. Familiarity with CI/CD tooling and Git-based workflows. Kubernetes & Container Platforms Strong experience with Kubernetes (EKS, GKE) and container ecosystems. Knowledge of Helm, service meshes, and container security best practices. Identity & Access Management (IAM) Deep understanding of IAM concepts in AWS and GCP, including roles, policies, and federation. Experience implementing RBAC and least-privilege access models at scale. Networking & Core Infrastructure Strong knowledge of cloud networking concepts: VPCs, subnets, routing, load balancing, DNS, and hybrid connectivity. Experience designing secure network architectures across multi-cloud and hybrid environments. Policy Management & Governance Direct experience with policy enforcement frameworks and governance models in AWS and GCP. Familiarity with policy-as-code tools and compliance automation practices. Required Qualifications: 8+ years of experience in cloud infrastructure, principal engineering, and/or cloud architecture roles. Hands-on expertise with AWS and GCP in enterprise environments, including high availability, disaster recovery, and multi-region architectures. Strong experience with Terraform, Infrastructure as Code (IaC), reusable modules, and automated CI/CD deployments. Proven experience designing and supporting Kubernetes platforms (EKS, GKE) and containerized environments. Deep knowledge of IAM, RBAC, least-privilege access models, and cloud security best practices. Strong understanding of cloud networking, including VPCs, routing, load balancing, DNS, VPNs, and hybrid connectivity. Experience implementing cloud governance, policy enforcement, compliance controls, and policy-as-code frameworks. Experience with AWS Control Tower and Landing Zone architectures. Highly Preferred Qualifications: 8-10+ years of experience engineering infrastructure in multi-cloud environments. Experience with cloud-native technologies and container platforms such as Kubernetes, Red Hat OpenShift, EKS, and GKE. Experience operating in highly regulated environments with knowledge of data privacy, compliance frameworks, and disaster recovery. Experience with principal engineering engineering, Internal Developer Platforms (IDPs), and self-service cloud models. Knowledge of FinOps practices and cloud cost optimization strategies. Experience with Azure and enterprise multi-cloud architectures. Familiarity with Helm, service meshes, and container security best practices. Relevant certifications such as AWS Certified Solutions Architect, Google Professional Cloud Architect, CKA, or CKAD. NTT DATA provides a reasonable range of compensation for U.S.-based positions. The starting pay range for this role will depend on the location of the successful candidate. For candidates in either of these locations, the starting pay range is $150,450k - $200k. Actual compensation will depend on a number of factors, including the candidate's actual work location, relevant experience, technical skills, and other qualifications. This position may also be eligible for incentive compensation based on individual and/or company performance. If the position offered in temporary, the position will not be eligible for incentive compensation. This position is eligible for company benefits including medical, dental, and vision insurance with an employer contribution, flexible spending or health savings account, life and AD&D insurance, short and long term disability coverage, paid time off, employee assistance, participation in a 401k program with company match, and additional voluntary or legally-required benefits. About NTT DATA NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries . click apply for full job details
09/23/2026
Full time
Req ID: 383115 NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Senior Principal Cloud Architect - AWS & Service Activation Hybrid Irving, TX or Charlotte, NC to join our team in Irving, Texas (US-TX), United States (US). Please Note: This hybrid role requires onsite client support 2-3 days weekly in either Irving, TX or Charlotte, NC. Candidates must be located within commuting distance of one of these offices. Applications from candidates outside these areas may not be considered. Position Summary We are seeking a Principal Cloud Engineer to architect design, build, and operate a secure, scalable, and highly automated multi-cloud platform leveraging Google Cloud Platform (GCP), Amazon Web Services (AWS), and Microsoft Azure. As a key technical leader, you will drive the evolution of our multi-cloud engineering, enabling resilient, cloud-native services that support mission-critical enterprise workloads. This role requires deep expertise in cloud architecture, infrastructure as code (IaC), and distributed systems engineering, with a strong focus on Terraform-driven automation and principal engineering standardization. The ideal candidate brings strong Kubernetes, IAM, networking, and policy management experience, along with a disciplined approach to security, reliability, and cost optimization. Key Responsibilities: Principal Engineering, Design & Architecture Architect, engineer, and implement scalable, secure, and resilient cloud solutions across GCP and AWS, with Azure as a plus. Engineer multi-region, highly available architectures aligned with enterprise standards and best practices. Infrastructure Automation & IaC Lead engineering, adoption, and development of Terraform-based infrastructure as code, ensuring consistency, compliance, and reusability. Build and maintain modular Terraform frameworks for provisioning cloud infrastructure. Integrate IaC into CI/CD pipelines to enable fully automated provisioning and deployment workflows. Containerization & Kubernetes Platforms Direct knowledge and engineering leveraging Kubernetes platforms (EKS, GKE) and associated ecosystems. Implementation of container orchestration best practices including scaling, observability, and security. Identity, Access & Security Governance Experienced working knowledge architecting and designing robust IAM strategies across GCP and AWS including least-privilege access models. Implement policy frameworks for access control, resource governance, and regulatory compliance. Networking & Connectivity Direct knowledge and experience architecting, designing, and engineering cloud networking solutions including VPC/VNet design, routing, peering, VPNs, and private connectivity. Ensure secure, high-performance connectivity between cloud environments and on-premises systems. Policy Management & Governance Deep understanding of policy management, policy-as-code frameworks to ensure consistent governance of resources, tagging strategies, and compliance controls. Collaboration with governance, policy, and security teams to ensure cloud environments meet enterprise and regulatory standards. Modernization & Mentorship Champion cloud-native and principal engineering engineering best practices across the organization. Drive the evolution of principal engineering practice by incorporating infrastructure-as-code (IaC) principles. Mentor junior engineers and provide technical leadership for complex cloud initiatives. Serve as an escalation point for critical cloud infrastructure and principal engineering issues. Ideal experience to have for this role: Cloud Platforms & Architecture Extensive firsthand experience with AWS and GCP in large-scale enterprise environments. Strong expertise in cloud architecture, including high availability, disaster recovery, and multi-region design. Deep understanding of cloud-native services and distributed systems principles. AWSControl tower and landing Zones. Infrastructure as Code & Automation Advanced proficiency in Terraform for infrastructure provisioning and lifecycle management. Experience designing reusable modules and managing Terraform state at scale. Familiarity with CI/CD tooling and Git-based workflows. Kubernetes & Container Platforms Strong experience with Kubernetes (EKS, GKE) and container ecosystems. Knowledge of Helm, service meshes, and container security best practices. Identity & Access Management (IAM) Deep understanding of IAM concepts in AWS and GCP, including roles, policies, and federation. Experience implementing RBAC and least-privilege access models at scale. Networking & Core Infrastructure Strong knowledge of cloud networking concepts: VPCs, subnets, routing, load balancing, DNS, and hybrid connectivity. Experience designing secure network architectures across multi-cloud and hybrid environments. Policy Management & Governance Direct experience with policy enforcement frameworks and governance models in AWS and GCP. Familiarity with policy-as-code tools and compliance automation practices. Required Qualifications: 8+ years of experience in cloud infrastructure, principal engineering, and/or cloud architecture roles. Hands-on expertise with AWS and GCP in enterprise environments, including high availability, disaster recovery, and multi-region architectures. Strong experience with Terraform, Infrastructure as Code (IaC), reusable modules, and automated CI/CD deployments. Proven experience designing and supporting Kubernetes platforms (EKS, GKE) and containerized environments. Deep knowledge of IAM, RBAC, least-privilege access models, and cloud security best practices. Strong understanding of cloud networking, including VPCs, routing, load balancing, DNS, VPNs, and hybrid connectivity. Experience implementing cloud governance, policy enforcement, compliance controls, and policy-as-code frameworks. Experience with AWS Control Tower and Landing Zone architectures. Highly Preferred Qualifications: 8-10+ years of experience engineering infrastructure in multi-cloud environments. Experience with cloud-native technologies and container platforms such as Kubernetes, Red Hat OpenShift, EKS, and GKE. Experience operating in highly regulated environments with knowledge of data privacy, compliance frameworks, and disaster recovery. Experience with principal engineering engineering, Internal Developer Platforms (IDPs), and self-service cloud models. Knowledge of FinOps practices and cloud cost optimization strategies. Experience with Azure and enterprise multi-cloud architectures. Familiarity with Helm, service meshes, and container security best practices. Relevant certifications such as AWS Certified Solutions Architect, Google Professional Cloud Architect, CKA, or CKAD. NTT DATA provides a reasonable range of compensation for U.S.-based positions. The starting pay range for this role will depend on the location of the successful candidate. For candidates in either of these locations, the starting pay range is $150,450k - $200k. Actual compensation will depend on a number of factors, including the candidate's actual work location, relevant experience, technical skills, and other qualifications. This position may also be eligible for incentive compensation based on individual and/or company performance. If the position offered in temporary, the position will not be eligible for incentive compensation. This position is eligible for company benefits including medical, dental, and vision insurance with an employer contribution, flexible spending or health savings account, life and AD&D insurance, short and long term disability coverage, paid time off, employee assistance, participation in a 401k program with company match, and additional voluntary or legally-required benefits. About NTT DATA NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries . click apply for full job details
The Global Program Controls organization provides governance, reporting, controls processes, cost management support, and portfolio visibility to help Oracle Cloud Infrastructure deliver global data center programs with greater consistency, transparency, and execution discipline. We are seeking a highly experienced individual contributor to help transform complex operational and technical information into accurate, decision-ready insights for senior executives. This role is designed for someone who is relentless about data integrity, intellectually curious, highly responsive, and comfortable operating in a complex, fast-moving, technical environment. The successful candidate will be able to move fluidly between detailed operational data and CEO-level communication-digging into the underlying facts, identifying inconsistencies, resolving ambiguity, and distilling the results into clear, concise, and credible executive reporting. This is not a traditional data analyst, project manager, or communications role. It combines elements of all three. The ideal candidate will bring strong analytical judgment, executive communication skills, technical aptitude, and a delivery-oriented mindset. Prior data center experience is not required. More important is a genuine interest in learning technical subject matter and the ability to become conversant in areas such as data center operations, low-voltage infrastructure, rack and compute deployment, hardware testing, and server turn-up activities. Key Responsibilities Data Integrity and Analytical Rigor Establish confidence in the accuracy, consistency, completeness, and traceability of critical business and operational data. Investigate discrepancies across systems, reports, teams, and data sources, and drive issues through resolution. Ask probing questions, challenge unsupported assumptions, and ensure that reported information can withstand executive scrutiny. Develop practical processes and controls that improve data quality and reporting reliability over time. Identify gaps in available information and work across teams to establish a clear, defensible source of truth. Balance speed with precision, particularly during time-sensitive executive requests and operational escalations. Executive Reporting and Communication Translate highly technical, tactical, or operational information into clear and compelling narratives for the CEO and other senior executives. Produce executive reports, briefings, presentations, summaries, and decision-support materials. Distill large volumes of information into the few insights, risks, decisions, and actions that matter most. Tailor the depth, language, and format of communications to the intended executive audience. Ensure that executive materials are accurate, concise, logically structured, and supported by verifiable data. Anticipate likely executive questions and prepare the supporting analysis needed to answer them. Business and Technical Learning Develop a deep understanding of the company's business model, operations, technology, customers, systems, and key performance drivers. Become comfortable with technical concepts related to data center operations, low-voltage work, rack and compute deployment, hardware validation and testing, server installation, and server turn-up. Learn unfamiliar technical and operational subject matter quickly, even without prior domain experience. Engage confidently with engineers, operators, vendors, and subject-matter experts to understand technical processes, terminology, dependencies, and risks. Ask thoughtful questions and continue learning until able to accurately explain technical topics to both operational and executive audiences. Build strong working relationships with subject-matter experts and learn enough to challenge, clarify, and synthesize their input. Connect information across functions to identify broader implications, dependencies, risks, and opportunities. Approach technical subject matter with curiosity, humility, and genuine enthusiasm rather than hesitation or avoidance. Program Execution and Problem Solving Bring a project manager's discipline to ambiguous, cross-functional work. Define problems, organize workstreams, clarify owners, establish milestones, and drive deliverables to completion. Coordinate across technical, operational, finance, strategy, and leadership teams to obtain required inputs and resolve blockers. Step into urgent or unstructured situations, create order quickly, and drive toward a practical solution. Manage multiple high-priority requests while maintaining accuracy, responsiveness, and sound judgment. Follow through persistently until commitments are completed and issues are resolved. Executive and Operational Support Serve as a trusted, high-leverage partner to senior leadership on critical analyses, executive deliverables, and time-sensitive business issues. Independently take ownership of complex requests with limited initial direction. Help diagnose and resolve urgent business or operational "fire drills." Exercise discretion when handling sensitive, confidential, or executive-level information. Provide flexibility during periods of elevated business activity, including occasional support outside standard working hours when urgent circumstances require it. Responsibilities Candidate Profile The ideal candidate is a senior, hands-on operator who combines intellectual curiosity with an exceptional standard for accuracy. This person is comfortable going deep into the details without losing sight of the larger business narrative. They are not satisfied with numbers simply because they appear in a system or report. They want to understand where the data came from, how it was calculated, whether it is complete, and whether it tells the full story. At the same time, they understand that analysis creates value only when it leads to clarity, decisions, and action. They are energized-not intimidated-by technical complexity and unfamiliar subject matter. They do not need to arrive with deep expertise in data center infrastructure, low-voltage systems, rack and compute deployment, testing, or server turn-up. They do need to enjoy learning how these areas work, be willing to engage deeply with technical experts, and develop enough fluency to identify risks, ask credible questions, and communicate clearly about the work. Required Qualifications Approximately 10-15 years of relevant experience in business operations, strategy and operations, analytics, executive reporting, program management, management consulting, finance, corporate strategy, or a related field. Demonstrated success translating complex technical or operational information into executive-level reports and presentations. Exceptional written communication skills, including the ability to draft concise, polished language suitable for a CEO or board-level audience. Strong analytical judgment and a demonstrated commitment to data accuracy and integrity. Experience investigating inconsistent or incomplete data and coordinating across teams to resolve underlying issues. Proven ability to lead complex, cross-functional initiatives as a senior individual contributor. Strong project and program management capabilities, including planning, prioritization, dependency management, and follow-through. Ability to learn unfamiliar businesses, technologies, infrastructure, and operating models quickly. Demonstrated comfort engaging with detailed technical information and working closely with engineering or operations teams. Comfort operating in an environment with ambiguity, changing priorities, and urgent requests. High degree of ownership, responsiveness, discretion, and professional judgment. This is an onsite position in Nashville, TN. Preferred Qualifications Experience in management consulting, corporate strategy, finance, operations, transformation, technical program management, or an executive office. Experience preparing materials for C-suite executives, boards of directors, investors, or other senior stakeholders. Familiarity with business intelligence, reporting, data governance, operational metrics, or enterprise systems. Experience working in a technically complex, infrastructure-intensive, or operationally demanding environment. Exposure to data center operations, physical infrastructure, low-voltage systems, hardware deployment, testing, commissioning, or related technical environments. Evidence of succeeding across industries or learning a new industry quickly. Advanced proficiency with presentation, spreadsheet, reporting, and data-visualization tools. What Success Looks Like Within the first year, the successful candidate will: Become a trusted source of accurate, well-supported information for senior leadership. Develop a strong working understanding of the company's technical and operational environment. Build practical fluency in data center operations, infrastructure deployment, testing, and server turn-up processes. Improve the clarity, consistency, and credibility of executive reporting. Reduce time spent reconciling conflicting data and resolving recurring reporting issues click apply for full job details
09/23/2026
Full time
The Global Program Controls organization provides governance, reporting, controls processes, cost management support, and portfolio visibility to help Oracle Cloud Infrastructure deliver global data center programs with greater consistency, transparency, and execution discipline. We are seeking a highly experienced individual contributor to help transform complex operational and technical information into accurate, decision-ready insights for senior executives. This role is designed for someone who is relentless about data integrity, intellectually curious, highly responsive, and comfortable operating in a complex, fast-moving, technical environment. The successful candidate will be able to move fluidly between detailed operational data and CEO-level communication-digging into the underlying facts, identifying inconsistencies, resolving ambiguity, and distilling the results into clear, concise, and credible executive reporting. This is not a traditional data analyst, project manager, or communications role. It combines elements of all three. The ideal candidate will bring strong analytical judgment, executive communication skills, technical aptitude, and a delivery-oriented mindset. Prior data center experience is not required. More important is a genuine interest in learning technical subject matter and the ability to become conversant in areas such as data center operations, low-voltage infrastructure, rack and compute deployment, hardware testing, and server turn-up activities. Key Responsibilities Data Integrity and Analytical Rigor Establish confidence in the accuracy, consistency, completeness, and traceability of critical business and operational data. Investigate discrepancies across systems, reports, teams, and data sources, and drive issues through resolution. Ask probing questions, challenge unsupported assumptions, and ensure that reported information can withstand executive scrutiny. Develop practical processes and controls that improve data quality and reporting reliability over time. Identify gaps in available information and work across teams to establish a clear, defensible source of truth. Balance speed with precision, particularly during time-sensitive executive requests and operational escalations. Executive Reporting and Communication Translate highly technical, tactical, or operational information into clear and compelling narratives for the CEO and other senior executives. Produce executive reports, briefings, presentations, summaries, and decision-support materials. Distill large volumes of information into the few insights, risks, decisions, and actions that matter most. Tailor the depth, language, and format of communications to the intended executive audience. Ensure that executive materials are accurate, concise, logically structured, and supported by verifiable data. Anticipate likely executive questions and prepare the supporting analysis needed to answer them. Business and Technical Learning Develop a deep understanding of the company's business model, operations, technology, customers, systems, and key performance drivers. Become comfortable with technical concepts related to data center operations, low-voltage work, rack and compute deployment, hardware validation and testing, server installation, and server turn-up. Learn unfamiliar technical and operational subject matter quickly, even without prior domain experience. Engage confidently with engineers, operators, vendors, and subject-matter experts to understand technical processes, terminology, dependencies, and risks. Ask thoughtful questions and continue learning until able to accurately explain technical topics to both operational and executive audiences. Build strong working relationships with subject-matter experts and learn enough to challenge, clarify, and synthesize their input. Connect information across functions to identify broader implications, dependencies, risks, and opportunities. Approach technical subject matter with curiosity, humility, and genuine enthusiasm rather than hesitation or avoidance. Program Execution and Problem Solving Bring a project manager's discipline to ambiguous, cross-functional work. Define problems, organize workstreams, clarify owners, establish milestones, and drive deliverables to completion. Coordinate across technical, operational, finance, strategy, and leadership teams to obtain required inputs and resolve blockers. Step into urgent or unstructured situations, create order quickly, and drive toward a practical solution. Manage multiple high-priority requests while maintaining accuracy, responsiveness, and sound judgment. Follow through persistently until commitments are completed and issues are resolved. Executive and Operational Support Serve as a trusted, high-leverage partner to senior leadership on critical analyses, executive deliverables, and time-sensitive business issues. Independently take ownership of complex requests with limited initial direction. Help diagnose and resolve urgent business or operational "fire drills." Exercise discretion when handling sensitive, confidential, or executive-level information. Provide flexibility during periods of elevated business activity, including occasional support outside standard working hours when urgent circumstances require it. Responsibilities Candidate Profile The ideal candidate is a senior, hands-on operator who combines intellectual curiosity with an exceptional standard for accuracy. This person is comfortable going deep into the details without losing sight of the larger business narrative. They are not satisfied with numbers simply because they appear in a system or report. They want to understand where the data came from, how it was calculated, whether it is complete, and whether it tells the full story. At the same time, they understand that analysis creates value only when it leads to clarity, decisions, and action. They are energized-not intimidated-by technical complexity and unfamiliar subject matter. They do not need to arrive with deep expertise in data center infrastructure, low-voltage systems, rack and compute deployment, testing, or server turn-up. They do need to enjoy learning how these areas work, be willing to engage deeply with technical experts, and develop enough fluency to identify risks, ask credible questions, and communicate clearly about the work. Required Qualifications Approximately 10-15 years of relevant experience in business operations, strategy and operations, analytics, executive reporting, program management, management consulting, finance, corporate strategy, or a related field. Demonstrated success translating complex technical or operational information into executive-level reports and presentations. Exceptional written communication skills, including the ability to draft concise, polished language suitable for a CEO or board-level audience. Strong analytical judgment and a demonstrated commitment to data accuracy and integrity. Experience investigating inconsistent or incomplete data and coordinating across teams to resolve underlying issues. Proven ability to lead complex, cross-functional initiatives as a senior individual contributor. Strong project and program management capabilities, including planning, prioritization, dependency management, and follow-through. Ability to learn unfamiliar businesses, technologies, infrastructure, and operating models quickly. Demonstrated comfort engaging with detailed technical information and working closely with engineering or operations teams. Comfort operating in an environment with ambiguity, changing priorities, and urgent requests. High degree of ownership, responsiveness, discretion, and professional judgment. This is an onsite position in Nashville, TN. Preferred Qualifications Experience in management consulting, corporate strategy, finance, operations, transformation, technical program management, or an executive office. Experience preparing materials for C-suite executives, boards of directors, investors, or other senior stakeholders. Familiarity with business intelligence, reporting, data governance, operational metrics, or enterprise systems. Experience working in a technically complex, infrastructure-intensive, or operationally demanding environment. Exposure to data center operations, physical infrastructure, low-voltage systems, hardware deployment, testing, commissioning, or related technical environments. Evidence of succeeding across industries or learning a new industry quickly. Advanced proficiency with presentation, spreadsheet, reporting, and data-visualization tools. What Success Looks Like Within the first year, the successful candidate will: Become a trusted source of accurate, well-supported information for senior leadership. Develop a strong working understanding of the company's technical and operational environment. Build practical fluency in data center operations, infrastructure deployment, testing, and server turn-up processes. Improve the clarity, consistency, and credibility of executive reporting. Reduce time spent reconciling conflicting data and resolving recurring reporting issues click apply for full job details
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
09/23/2026
Full time
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
09/23/2026
Full time
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
09/23/2026
Full time
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The RO Performance Team is accountable for the largest metrics system deployment in the AV industry. The team manages data pipelines that monitor Waymo's AV fleet in the real world at various latencies: from minutes to hours; supporting critical business decisions for deployments in new markets and identifying emergent business critical challenges. In addition, the same pipeline technologies are deployed for the virtual world to measure key outcomes from Waymo's massive scale simulations, generating insights to support new Driver releases. Important objectives for the next couple of years include: Reliably scaling as the business scales, Waymo is in its fast expansion phase and systems have to be built to support 100x growth. Timeliness: As Waymo's deployments accelerate across the globe, time to first detection needs to consistently improve. Comprehensive coverage: Expanding into new weather conditions and new markets brings new challenges. Increasing the metric and detection offerings to meet these challenges is key to achieving our business goals. The team works closely with many other teams such as: metrics development teams including large AI deployments, System engineers and Data scientists that produce definitions and quality controls, infrastructure and UI teams that support underlying capabilities, product teams to stay sensitive to changing business needs, and site reliability engineering (SRE) teams to maintain high system uptime. You Will: Be a part of the Event Insights sub-team, Once metric events, e.g., a hard brake occurred, are minted from heavy logs processing pipelines, an event lifecycle emerges where such events may get sent to other systems for further refinement, including but not limited to: human triage, VLM inference, clustering. Key workstreams that the TLM will be accountable for include: Clustering: Work on creation of a dynamic event clustering product to produce event clusters at low latency for rapid understanding of emerging issues during new deployments of the Waymo Driver in simulation or in the real world. Storage & APIs: Events need to be stored and queried by different downstream systems. Consumption APIs include pull or push based paradigms. Labels received from human triage or VLMs update existing events. Low Latency Orchestration: Gathering further event insights requires interaction with other systems at Waymo such as for human triage or large VLM inference. Orchestrating these interactions reliably and at low latency is key for meeting Waymo's challenges. Visualization: Work closely with partner UI teams to ensure that event insights can be delivered to end consumers and provide value for understanding the Waymo Driver's performance. You have: 5+ years of full-time software engineering experience, or a quantitative PhD with 2+ years of professional software engineering experience C++ proficiency Python familiarity SQL familiarity Excited about autonomous driving, Simulation + Eval We prefer: 7+ years of industry experience or 3+ years post-doc experience in a quantitative- or quality- focused engineering role in which you had experience as a technical lead for a team in which you were performing tasks like developing hypotheses, designing and running experiments, processing data from experiments, synthesizing conclusions, and ensuring the long term stability and health of the production system. B.Sc. in Computer Science Experience working with large FAANG scale distributed systems. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $213,000-$263,000 USD
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/23/2026
Full time
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/23/2026
Full time
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/23/2026
Full time
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while working on the technologies that are redefining computing. You'll harness the power of AI to deliver groundbreaking solutions to some of the world's toughest problems-and see your work have real, lasting impact! What you'll be doing: Own SRE solutions end to end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. Use IaC(Infrastructure as Code) and config management to standardize and automate provisioning everywhere. Deliver solutions in a globally distributed, multi cloud hybrid environment - On prem, AWS, GCP, and OCI. Design for failure with redundancy, failure domains, progressive delivery, and strict change control. Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. Conduct capacity management and planning to meet ongoing operational needs. Detects performance issues and recommends solutions to maintain world class service quality. Collaborate with various teams in a fast paced environment to ensure seamless project completion. Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. Experience supporting large scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. Proficient in monitoring, metrics, container management, and log collection tools. 5+ years of coding/scripting experience in at least two high level programming languages such as Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: Published technical write ups or talks (conference presentations, meetups, engineering blogs) that deep dive into real world reliability, observability, or large scale HPC/SRE problems and their solutions. Maintainer or co maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense. Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense. Kodiak is recruiting a Staff Software Engineer to lead design and development of mission planning, routing, and human-assisted autonomy for Kodiak driverless trucks. The role provides scope and visibility to influence the roadmap for the entire motion planning stack, including behavior, prediction, and trajectory generation. This is an opportunity to collaborate with Product, Perception, Controls, and other cross-functional teams in order to improve overall autonomy performance. Customers interact with Kodiak trucks through mission assignments, making this a key engineering position to achieve product excellence. The role spans a range of motion planning responsibilities, from graph traversal and routing, to mission planning and product development, to human-assisted autonomy and operations. In this role, you will: Lead the design and development of mission planning, routing, and human-assisted autonomy for driverless operation on public roads Develop the architecture and long-term roadmap for motion planning Identify metrics to orient engineering teams toward product excellence Study the product for opportunities to delight our customers Mentor and provide technical leadership to junior and senior engineers, raising the bar for engineering excellence across the engineering organization Integrate onboard and offboard driving logic Collaborate with human-assisted autonomy and operations specialists to facilitate safe, efficient missions Evaluate the effectiveness of the Kodiak Driver Automate and expedite truck launches and landings Partner cross-functionally with Perception, Controls, Systems, Simulation, and Product teams to define interfaces, requirements, and system behavior Solve difficult real-world problems involving route generation, mission updates, responsive driving, real-time and cloud computing, and more! What you'll bring: 6+ years of experience developing production-level robotics software Strong facility with modern C++ Expertise in motion planning, graph theory, distributed computing, and related domains Delivery of complex autonomous systems Leadership to solve ambiguous, real-world challenges Bonus points if you have: Managed or mentored robotics engineers Deployed software on autonomous vehicles Built a safety case for driverless autonomy Improved system performance (latency, throughput, reliability) in production environments Built on simulation frameworks Based decisions on data analysis What we offer: Competitive compensation package including equity and annual bonuses Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits) MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, & Critical Illness Insurance Flexible PTO, 10 paid holidays, and generous parental leave policies Our office is centrally located in Mountain View, CA Office perks: dog-friendly, free catered lunch, a fully stocked kitchen, and free EV charging Long Term Disability, Short Term Disability, Life Insurance Wellbeing Benefits - Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation) Fidelity 401(k) Commuter, FSA, Dependent Care FSA, HSA Various incentive programs (referral bonuses, patent bonuses, etc.) The pay range listed below reflects the base salary in our SF/Silicon Valley location, across several internal levels. Actual starting pay will be based on job-related factors including: work location, experience, relevant training, education, skill level and performance during interview. Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits package California Pay Range $200,000-$250,000 USD At Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited. Kodiak is committed to equal opportunity employment regardless of race, ethnicity, religion, gender identity, sexual orientation, age, disability, or veteran status, or any other basis protected by applicable law. In alignment with its business operations, Kodiak adheres to all relevant statutes, regulations, and administrative prerequisites. Accordingly, roles that carry more sensitive requirements may be limited to candidates that can satisfy additional scrutiny and eligibility for such positions may hinge on verification of a candidate's residence, U.S. person status, and/or citizenship status. Should the position require, and Kodiak determines that a candidate's residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak's sole discretion, or, as an alternative, opt not to proceed with the candidate's application. If applicable, Kodiak may provide visa sponsorship for eligible candidates. We use a third-party AI tool (Endorsed) to assist in the initial screening of applications. As part of the evaluation process, we provide Endorsed with job requirements and candidate-submitted applications. Final hiring decisions are made by our human recruitment team, and no automated system makes the ultimate decision regarding hiring. Certain features of the platform may qualify it as an Automated Employment Decision Tool (AEDT) under applicable regulations. We began using Endorsed on January 1, 2026. You can review the independent bias audit report covering our use of Endorsed here ( ). By submitting your application, you acknowledge that your application may be processed by AI systems as part of the screening and selection process. If you have any questions or would like to request a separate review of your application, please contact with "Separate Review Request" in the email subject line.
09/23/2026
Full time
Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense. Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense. Kodiak is recruiting a Staff Software Engineer to lead design and development of mission planning, routing, and human-assisted autonomy for Kodiak driverless trucks. The role provides scope and visibility to influence the roadmap for the entire motion planning stack, including behavior, prediction, and trajectory generation. This is an opportunity to collaborate with Product, Perception, Controls, and other cross-functional teams in order to improve overall autonomy performance. Customers interact with Kodiak trucks through mission assignments, making this a key engineering position to achieve product excellence. The role spans a range of motion planning responsibilities, from graph traversal and routing, to mission planning and product development, to human-assisted autonomy and operations. In this role, you will: Lead the design and development of mission planning, routing, and human-assisted autonomy for driverless operation on public roads Develop the architecture and long-term roadmap for motion planning Identify metrics to orient engineering teams toward product excellence Study the product for opportunities to delight our customers Mentor and provide technical leadership to junior and senior engineers, raising the bar for engineering excellence across the engineering organization Integrate onboard and offboard driving logic Collaborate with human-assisted autonomy and operations specialists to facilitate safe, efficient missions Evaluate the effectiveness of the Kodiak Driver Automate and expedite truck launches and landings Partner cross-functionally with Perception, Controls, Systems, Simulation, and Product teams to define interfaces, requirements, and system behavior Solve difficult real-world problems involving route generation, mission updates, responsive driving, real-time and cloud computing, and more! What you'll bring: 6+ years of experience developing production-level robotics software Strong facility with modern C++ Expertise in motion planning, graph theory, distributed computing, and related domains Delivery of complex autonomous systems Leadership to solve ambiguous, real-world challenges Bonus points if you have: Managed or mentored robotics engineers Deployed software on autonomous vehicles Built a safety case for driverless autonomy Improved system performance (latency, throughput, reliability) in production environments Built on simulation frameworks Based decisions on data analysis What we offer: Competitive compensation package including equity and annual bonuses Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits) MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, & Critical Illness Insurance Flexible PTO, 10 paid holidays, and generous parental leave policies Our office is centrally located in Mountain View, CA Office perks: dog-friendly, free catered lunch, a fully stocked kitchen, and free EV charging Long Term Disability, Short Term Disability, Life Insurance Wellbeing Benefits - Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation) Fidelity 401(k) Commuter, FSA, Dependent Care FSA, HSA Various incentive programs (referral bonuses, patent bonuses, etc.) The pay range listed below reflects the base salary in our SF/Silicon Valley location, across several internal levels. Actual starting pay will be based on job-related factors including: work location, experience, relevant training, education, skill level and performance during interview. Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits package California Pay Range $200,000-$250,000 USD At Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited. Kodiak is committed to equal opportunity employment regardless of race, ethnicity, religion, gender identity, sexual orientation, age, disability, or veteran status, or any other basis protected by applicable law. In alignment with its business operations, Kodiak adheres to all relevant statutes, regulations, and administrative prerequisites. Accordingly, roles that carry more sensitive requirements may be limited to candidates that can satisfy additional scrutiny and eligibility for such positions may hinge on verification of a candidate's residence, U.S. person status, and/or citizenship status. Should the position require, and Kodiak determines that a candidate's residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak's sole discretion, or, as an alternative, opt not to proceed with the candidate's application. If applicable, Kodiak may provide visa sponsorship for eligible candidates. We use a third-party AI tool (Endorsed) to assist in the initial screening of applications. As part of the evaluation process, we provide Endorsed with job requirements and candidate-submitted applications. Final hiring decisions are made by our human recruitment team, and no automated system makes the ultimate decision regarding hiring. Certain features of the platform may qualify it as an Automated Employment Decision Tool (AEDT) under applicable regulations. We began using Endorsed on January 1, 2026. You can review the independent bias audit report covering our use of Endorsed here ( ). By submitting your application, you acknowledge that your application may be processed by AI systems as part of the screening and selection process. If you have any questions or would like to request a separate review of your application, please contact with "Separate Review Request" in the email subject line.
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Network Engineer position at the Director level, which is part of the job family responsible for maintaining the stability and reliability of the organization's infrastructure systems, ensuring optimal performance and availability to support business operations. Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals. Interested in joining a team that's eager to create, innovate and make an impact on the world? Read on. We are looking for a Networks Engineer (SME) to join our team, The broad knowledge of all major topics is critical. The ideal candidate will have 8+ years of experience in the enterprise environment (preferably in a Financial Services Firm), Technology vendor or systems integrator in the role of designing or implementing and providing operational support to the large scale networks. Ability to work both independently and as a team member is essential. Strong communication and presentation skills are imperative to be able explain network design, articulate risk, and facilitate decision making. What you'll bring to the role: Ethernet technologies: STP, 802.1Q, VPC, Multilayer Switching; Leaf-Spine IP Fabric IP Routing: Multicast PIM, OSPF, BGP, MPBGP, TCP/IP VxLAN Platform knowledge on Cisco/Arista Switching & Routing Platforms Experience working with low latency market data & Internet Service Providers environment Knowledge with IP NAT, PAT, Multicast Routing - IGMP, IPSec VPN, GRE tunnels Skills Desired: Familiarity with JIRA for project and task management, specifically work in Agile. Ability to use Microsoft applications with strong presentation skills using Canva, Powerpoint Proficient in UNIX or Linux Experience with scripting / automation using Python / Ansible is a strong plus. Wireshark, Infoblox, HPNA, Splunk, Sevone, Ansible is a plus Good to have some cloud and wireless knowledge Strong sense of security concept include defense in depth, zero trust networking, least privilege principle, risks and controls. Additionally, the ideal candidate should be Flexible and adaptable to meet the team's target Honest, hardworking, and reliable Possess a strong sense of ownership and accountability Must be able to work on the weekend and help with the executions and deployments. Must be able to drive enterprise level initiatives working with senior stakeholders across different regions. The resource must be willing to be onsite from day 1. With manager's approval, some flexibility to work in a hybrid environment. WHAT YOU CAN EXPECT FROM MORGAN STANLEY: At Morgan Stanley, we raise, manage and allocate capital for our clients - helping them reach their goals. We do it in a way that's differentiated - and we've done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren't just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you'll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There's also ample opportunity to move about the business for those who show passion and grit in their work. To learn more about our offices across the globe, please copy and paste into your browser. Expected base pay rates for the role will be between $120,000 and $165,000 per year at the commencement of employment. However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs. Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents. Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences. For more information, please visit:
09/23/2026
Full time
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Network Engineer position at the Director level, which is part of the job family responsible for maintaining the stability and reliability of the organization's infrastructure systems, ensuring optimal performance and availability to support business operations. Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals. Interested in joining a team that's eager to create, innovate and make an impact on the world? Read on. We are looking for a Networks Engineer (SME) to join our team, The broad knowledge of all major topics is critical. The ideal candidate will have 8+ years of experience in the enterprise environment (preferably in a Financial Services Firm), Technology vendor or systems integrator in the role of designing or implementing and providing operational support to the large scale networks. Ability to work both independently and as a team member is essential. Strong communication and presentation skills are imperative to be able explain network design, articulate risk, and facilitate decision making. What you'll bring to the role: Ethernet technologies: STP, 802.1Q, VPC, Multilayer Switching; Leaf-Spine IP Fabric IP Routing: Multicast PIM, OSPF, BGP, MPBGP, TCP/IP VxLAN Platform knowledge on Cisco/Arista Switching & Routing Platforms Experience working with low latency market data & Internet Service Providers environment Knowledge with IP NAT, PAT, Multicast Routing - IGMP, IPSec VPN, GRE tunnels Skills Desired: Familiarity with JIRA for project and task management, specifically work in Agile. Ability to use Microsoft applications with strong presentation skills using Canva, Powerpoint Proficient in UNIX or Linux Experience with scripting / automation using Python / Ansible is a strong plus. Wireshark, Infoblox, HPNA, Splunk, Sevone, Ansible is a plus Good to have some cloud and wireless knowledge Strong sense of security concept include defense in depth, zero trust networking, least privilege principle, risks and controls. Additionally, the ideal candidate should be Flexible and adaptable to meet the team's target Honest, hardworking, and reliable Possess a strong sense of ownership and accountability Must be able to work on the weekend and help with the executions and deployments. Must be able to drive enterprise level initiatives working with senior stakeholders across different regions. The resource must be willing to be onsite from day 1. With manager's approval, some flexibility to work in a hybrid environment. WHAT YOU CAN EXPECT FROM MORGAN STANLEY: At Morgan Stanley, we raise, manage and allocate capital for our clients - helping them reach their goals. We do it in a way that's differentiated - and we've done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren't just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you'll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There's also ample opportunity to move about the business for those who show passion and grit in their work. To learn more about our offices across the globe, please copy and paste into your browser. Expected base pay rates for the role will be between $120,000 and $165,000 per year at the commencement of employment. However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs. Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents. Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences. For more information, please visit:
As a Senior DevOps Engineer, you will be a key technical contributor responsible for designing, implementing, and operating scalable, resilient infrastructure and CI/CD pipelines that support the full software development lifecycle. You will work closely with Wolters Kluwer Product Teams to embed DevOps best practices into agile workflows, enabling continuous integration, automated testing, and reliable deployment across environments. In this role, you will focus on hands on engineering excellence-building, automating, and operating cloud native platforms that improve system reliability, performance, and maintainability. You will partner closely with application development teams to bridge development and operations, supporting microservices, containerized workloads, and cloud native architectures. While not a formal people manager, you will act as a senior technical mentor and role model, promoting engineering rigor, code as infrastructure, and continuous improvement. Responsibilities Design, engineer, and automate secure, scalable cloud infrastructure in Azure and AWS using Infrastructure as Code (IaC) tools such as Terraform, Ansible, and Jenkins, applying software engineering best practices including modular design, version control, and automated testing. Implement and maintain Infrastructure as Code and CI/CD pipelines, contributing reusable modules, templates, and patterns that improve consistency and reliability across teams. Design and implement modern compute platforms, including containerized and serverless solutions (AKS, EKS, Docker, Azure Functions), with an emphasis on scalability, maintainability, and performance. Build and maintain CI/CD pipelines as software products, ensuring strong test coverage, artifact management, promotion workflows, and deployment automation across multiple environments. Support and evolve cloud native architectures, applying engineering principles such as abstraction, decoupling, fault isolation, and observability. Implement observability solutions using metrics, logging, and tracing to enable proactive issue detection, faster troubleshooting, and root cause analysis. Ensure infrastructure and automation solutions comply with enterprise DevOps, security, and compliance standards, contributing to architectural reviews and governance processes. Serve as a senior technical mentor, providing guidance through code reviews, design discussions, and knowledge sharing-without direct people management responsibilities. Evaluate and prototype emerging tools and technologies, applying engineering rigor to assess value, performance, and integration feasibility. Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity planning, and incident response to improve system reliability and reduce operational toil. Deploy, operate, and support business critical applications, ensuring high availability, fault tolerance, and performance optimization. Participate in modernization initiatives, supporting the re architecture and cloud native transformation of legacy platforms. Identify and remediate engineering inefficiencies by proposing and implementing automation and architectural improvements. Participate in post incident reviews, contributing to blameless root cause analysis and long term corrective actions. Participate in on call rotations, continuously improving alert quality, reducing noise, and automating remediation where possible. Qualifications Bachelor's degree in Engineering, Computer Science, or a related field (Master's degree preferred). 5+ years of experience in DevOps, Site Reliability Engineering, Release Engineering, or related roles, with strong hands on software engineering experience. Strong background in software development, with experience in languages such as Python, .NET, or Java. Proven experience working with cloud platforms (Azure and AWS). Proficiency in scripting languages such as PowerShell and Bash. Solid understanding of core Azure and AWS services (PaaS, IaaS, SaaS). Strong experience with source control and automation tools, including Git. Hands on experience with Infrastructure as Code tools such as Terraform or CloudFormation. Experience building and operating CI/CD pipelines using tools such as Azure DevOps, Jenkins, or similar platforms. Strong problem solving skills with attention to detail and operational excellence. Ability to clearly communicate technical concepts to engineers and non engineering stakeholders. Demonstrated commitment to DevOps culture, including continuous integration, automated testing, deployment automation, and full lifecycle ownership. Experience building or supporting observability platforms and defining operational best practices. Experience troubleshooting and automating diagnostics across Linux and Windows environments. Our Interview Practices To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we're getting to know you-not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process. Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process. Compensation: $92,700.00 - $161,850.00 USDThis role is eligible for Bonus. Compensation range listed is based on primary location of the position. Actual base salary offer is influenced by a wide array of factors including but not limited to skills, experience and actual hiring location. Your recruiter can share more information about the specific offer for the job location during the hiring process. Additional Information: Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life, including but not limited to: Medical, Dental, & Vision Plans, 401(k), FSA/HSA, Commuter Benefits, Tuition Assistance Plan, Vacation and Sick Time, and Paid Parental Leave. Full details of our benefits are available upon request.
09/23/2026
Full time
As a Senior DevOps Engineer, you will be a key technical contributor responsible for designing, implementing, and operating scalable, resilient infrastructure and CI/CD pipelines that support the full software development lifecycle. You will work closely with Wolters Kluwer Product Teams to embed DevOps best practices into agile workflows, enabling continuous integration, automated testing, and reliable deployment across environments. In this role, you will focus on hands on engineering excellence-building, automating, and operating cloud native platforms that improve system reliability, performance, and maintainability. You will partner closely with application development teams to bridge development and operations, supporting microservices, containerized workloads, and cloud native architectures. While not a formal people manager, you will act as a senior technical mentor and role model, promoting engineering rigor, code as infrastructure, and continuous improvement. Responsibilities Design, engineer, and automate secure, scalable cloud infrastructure in Azure and AWS using Infrastructure as Code (IaC) tools such as Terraform, Ansible, and Jenkins, applying software engineering best practices including modular design, version control, and automated testing. Implement and maintain Infrastructure as Code and CI/CD pipelines, contributing reusable modules, templates, and patterns that improve consistency and reliability across teams. Design and implement modern compute platforms, including containerized and serverless solutions (AKS, EKS, Docker, Azure Functions), with an emphasis on scalability, maintainability, and performance. Build and maintain CI/CD pipelines as software products, ensuring strong test coverage, artifact management, promotion workflows, and deployment automation across multiple environments. Support and evolve cloud native architectures, applying engineering principles such as abstraction, decoupling, fault isolation, and observability. Implement observability solutions using metrics, logging, and tracing to enable proactive issue detection, faster troubleshooting, and root cause analysis. Ensure infrastructure and automation solutions comply with enterprise DevOps, security, and compliance standards, contributing to architectural reviews and governance processes. Serve as a senior technical mentor, providing guidance through code reviews, design discussions, and knowledge sharing-without direct people management responsibilities. Evaluate and prototype emerging tools and technologies, applying engineering rigor to assess value, performance, and integration feasibility. Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity planning, and incident response to improve system reliability and reduce operational toil. Deploy, operate, and support business critical applications, ensuring high availability, fault tolerance, and performance optimization. Participate in modernization initiatives, supporting the re architecture and cloud native transformation of legacy platforms. Identify and remediate engineering inefficiencies by proposing and implementing automation and architectural improvements. Participate in post incident reviews, contributing to blameless root cause analysis and long term corrective actions. Participate in on call rotations, continuously improving alert quality, reducing noise, and automating remediation where possible. Qualifications Bachelor's degree in Engineering, Computer Science, or a related field (Master's degree preferred). 5+ years of experience in DevOps, Site Reliability Engineering, Release Engineering, or related roles, with strong hands on software engineering experience. Strong background in software development, with experience in languages such as Python, .NET, or Java. Proven experience working with cloud platforms (Azure and AWS). Proficiency in scripting languages such as PowerShell and Bash. Solid understanding of core Azure and AWS services (PaaS, IaaS, SaaS). Strong experience with source control and automation tools, including Git. Hands on experience with Infrastructure as Code tools such as Terraform or CloudFormation. Experience building and operating CI/CD pipelines using tools such as Azure DevOps, Jenkins, or similar platforms. Strong problem solving skills with attention to detail and operational excellence. Ability to clearly communicate technical concepts to engineers and non engineering stakeholders. Demonstrated commitment to DevOps culture, including continuous integration, automated testing, deployment automation, and full lifecycle ownership. Experience building or supporting observability platforms and defining operational best practices. Experience troubleshooting and automating diagnostics across Linux and Windows environments. Our Interview Practices To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we're getting to know you-not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process. Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process. Compensation: $92,700.00 - $161,850.00 USDThis role is eligible for Bonus. Compensation range listed is based on primary location of the position. Actual base salary offer is influenced by a wide array of factors including but not limited to skills, experience and actual hiring location. Your recruiter can share more information about the specific offer for the job location during the hiring process. Additional Information: Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life, including but not limited to: Medical, Dental, & Vision Plans, 401(k), FSA/HSA, Commuter Benefits, Tuition Assistance Plan, Vacation and Sick Time, and Paid Parental Leave. Full details of our benefits are available upon request.