Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Job Description Job Description Overview Configuration Manager WORK LOCATION: Hanscom AFB, MA (In Office/On Base 3 days a week) SALARY RANGE: $115,000 - $135,000 depending on experience, certifications, and qualifications JOB STATUS: Full-time; salaried SECURITY CLEARANCE: MUST have an active Secret security clearance Astrion has an exciting opportunity for a Configuration Manager to support the Survivable Airborne Operations Center (SAOC) Division at Hanscom AFB, MA. The SAOC program is a Major Defense Acquisition Program tasked with leading the development and delivery of the replacement fleet for the E-4B National Airborne Operations (NAOC). Responsibilities Responsible for maintaining technical baselines, especially software baselines, of USAF systems in accordance with Dept of Defense and USAF guidance. Design, develop, and establish configuration management documentation based on prescribed standards, program requirements, and conventional best practices. Prepare documentation for, coordinate, and lead Configuration Control Boards. Conduct closeout actions and archive all CM documentation. Report CCB status to leadership. Assist in preparing and reviewing program/project documentation, e.g. Requests for Proposals (RFPs), Statement Of Objectives (SOOs), Statements-of-Work (SOWs), Performance Works Statements (PWS), System/Technical Requirements Documents (SRD/TRD), Contract Data Requirements Lists (CDRL), AF Form 1067s, etc. and assist in source selections for systems. Review and update CM plans periodically in accordance with USAF guidance. Develop process, then maintain and standardize technical performance matrices to include capturing update steps, configuration control considerations, and approval process. Develop process documents. Maintain configuration control of assessments developed by the technical team. Qualifications Experience on Dept of Defense or Air Force acquisition programs is highly desired. Experience with Configuration Management tools. Demonstrated characteristics of a self-starter who is able to work independently and effectively interact with program team members and senior leadership.
09/23/2026
Full time
Job Description Job Description Overview Configuration Manager WORK LOCATION: Hanscom AFB, MA (In Office/On Base 3 days a week) SALARY RANGE: $115,000 - $135,000 depending on experience, certifications, and qualifications JOB STATUS: Full-time; salaried SECURITY CLEARANCE: MUST have an active Secret security clearance Astrion has an exciting opportunity for a Configuration Manager to support the Survivable Airborne Operations Center (SAOC) Division at Hanscom AFB, MA. The SAOC program is a Major Defense Acquisition Program tasked with leading the development and delivery of the replacement fleet for the E-4B National Airborne Operations (NAOC). Responsibilities Responsible for maintaining technical baselines, especially software baselines, of USAF systems in accordance with Dept of Defense and USAF guidance. Design, develop, and establish configuration management documentation based on prescribed standards, program requirements, and conventional best practices. Prepare documentation for, coordinate, and lead Configuration Control Boards. Conduct closeout actions and archive all CM documentation. Report CCB status to leadership. Assist in preparing and reviewing program/project documentation, e.g. Requests for Proposals (RFPs), Statement Of Objectives (SOOs), Statements-of-Work (SOWs), Performance Works Statements (PWS), System/Technical Requirements Documents (SRD/TRD), Contract Data Requirements Lists (CDRL), AF Form 1067s, etc. and assist in source selections for systems. Review and update CM plans periodically in accordance with USAF guidance. Develop process, then maintain and standardize technical performance matrices to include capturing update steps, configuration control considerations, and approval process. Develop process documents. Maintain configuration control of assessments developed by the technical team. Qualifications Experience on Dept of Defense or Air Force acquisition programs is highly desired. Experience with Configuration Management tools. Demonstrated characteristics of a self-starter who is able to work independently and effectively interact with program team members and senior leadership.
Gravitee is a 2025 Gartner Magic Quadrant Leader , on a mission to govern the world's intelligence . We deliver the industry's most advanced platform for Any API, Any Event, and Any AI Agent , trusted by global leaders like Michelin, Roche, and Blue Yonder. Why join us? The Mission : We are the first to bridge traditional API Management with the new frontier of AI Agent Security The Momentum : A high-growth Leader - combining market credibility with startup speed The DNA : We hire people who Hold Nothing Back - passionate builders who want to redefine digital infrastructure Don't just watch the AI revolution. Build the infrastructure that controls and secures it. The Role We are looking for a Senior Software Engineer to build and maintain the identity and authorization features of Gravitee Access Management (AM) - across the AM runtime and the access-management experience in Gamma, Gravitee's next generation product surface. This is a new role. Today, AM engineering is based entirely in Europe. This hire establishes US-hours ownership of Level 3 and Level 4 authentication and authorization incidents, and adds delivery capacity toward AM parity in Gamma - part of building sustainable L3/L4 engineering capability in the US. You will split your time roughly 80% feature delivery and 20% L3/L4 support and bug fixing (it varies week to week), working as an embedded member of the AM team, which is based in Europe. What You Will Be Doing In this role, you will: Design and deliver features end to end, from discovery and technical design through implementation, testing, release, and iteration. Build and maintain identity and authorization features of Gravitee Access Management, across the AM runtime and the AM experience in Gamma. Implement and support OAuth 2.0 and OIDC flows (authorization code + PKCE, client credentials, token exchange), SAML 2.0 as both IdP and SP, SCIM, and FAPI/CIBA/UMA profiles. Work with token and session semantics - JWT, JWKS, key rotation, revocation, introspection, MFA and step-up, WebAuthn/FIDO2, and IdP federation and social login. Keep security behavior and upgrades safe: standards compliance, secure defaults, certificate and secret handling, consent, audit logs, and defenses against token replay, SSRF, and account takeover. Own safe migrations and backward compatibility across MongoDB and JDBC, and support multi-domain, multi-region deployments and login/token endpoint performance. Own US-hours Level 3 and Level 4 escalations for AM customers as part of the L3 pager duty rotation. Use LLMs and AI-assisted development tools thoughtfully for prototyping, implementation, testing, debugging, and exploration, applying sound engineering judgment to validate AI-generated work. Write meaningful automated tests and contribute to reliable delivery practices. • Collaborate with product managers, designers, engineers, and technical leaders - including the AM team based in Europe - to discover effective solutions and improve them through code and design reviews. Share what you learn and help the team make practical choices as identity standards and protocols evolve. Essential Skills We are looking for evidence that you can succeed in the role, whether gained through employment, open-source work, or equivalent practical experience: 5+ years building and running production backend software, on a team that ships and supports its own product; you have personally resolved production incidents. Strong Java experience (C# accepted if the object-oriented depth is there), with Maven and a reactive stack such as Vert.x/RxJava. Deep working knowledge of identity standards: OAuth 2.0 and OIDC flows (authorization code + PKCE, client credentials, token exchange), SAML 2.0, SCIM, and FAPI/CIBA/UMA profiles. • Solid grasp of token and session semantics: JWT, JWKS, rotation, revocation, introspection, MFA/step-up, WebAuthn/FIDO2, and IdP federation. A security-first mindset: secure defaults, certificate and secret handling, audit logging, and awareness of token replay, SSRF, and account-takeover risks. Experience with safe migrations and backward compatibility across persistent data stores such as MongoDB or JDBCbacked relational databases. Git-based workflow, code review, and writing your own automated tests. Hands-on experience using LLMs or AI coding assistants as part of an engineering workflow, combined with the judgment to review and improve their output. Clear communication, collaborative problem-solving, and the ability to take an ambiguous problem through to production. Desired Skills You do not need to match every item. We would be especially interested in experience with: • Experience at an API gateway, proxy, or service-mesh vendor, or on the API platform team of a large company (e.g., Kong, Google Apigee, MuleSoft, Tyk, Solo.io, Traefik, WSO2). Kubernetes operators and CRDs; OpenAPI tooling; service mesh or Envoy experience. Docker, Kubernetes, and cloud-native application delivery. Model Context Protocol (MCP), Agent2Agent (A2A), tool calling, LLM proxies, or other emerging AI protocols and standards. Prior production experience is not required. Building or operating LLM-powered applications, RAG systems, or agentic workflows - especially their security, governance, and observability needs. Open-source software or enterprise developer platforms. Who Thrives at Gravitee Our growth is powered by people who bring passion to what they build, professionalism to how they work, and a commitment to doing things well. You will thrive here if you: • Bring energy and a constructive attitude to the team. • Adapt quickly and enjoy learning unfamiliar technologies and domains. • Take ownership, communicate clearly, and follow through with urgency. • Balance delivery speed with thoughtful engineering judgment. • Start with the customer problem and care about the quality of the experience you create. • Enjoy working in a fast-moving, collaborative, international environment. Life at Gravitee At Gravitee, we invest in humans, not just roles. You'll get: • Salary of $160,000 • Competitive medical coverage. • Pension / 401(k) program options. • Stock options - you build it, you own it. • 25 days of holiday plus in-country national holidays. • Three mental health days and a wellness allowance. • Your birthday off. • A professional development budget to support your growth. • A hybrid work culture with hubs across regions. • Quarterly team events and an annual company offsite. • A collaborative, international company culture. • Opportunities to grow your scope and career as Gravitee grows. At Gravitee, we believe diverse perspectives make better products and stronger teams. No employee or applicant will be treated less favorably on the grounds of sex, marital status, race, color, nationality, ethnic or national origin, disability, gender, sexual orientation, gender identity, age, pregnancy or maternity, marital or civil partner status, religion, or belief. By applying, you consent to Gravitee storing and processing the personal information you submit as part of the recruitment process.
09/23/2026
Full time
Gravitee is a 2025 Gartner Magic Quadrant Leader , on a mission to govern the world's intelligence . We deliver the industry's most advanced platform for Any API, Any Event, and Any AI Agent , trusted by global leaders like Michelin, Roche, and Blue Yonder. Why join us? The Mission : We are the first to bridge traditional API Management with the new frontier of AI Agent Security The Momentum : A high-growth Leader - combining market credibility with startup speed The DNA : We hire people who Hold Nothing Back - passionate builders who want to redefine digital infrastructure Don't just watch the AI revolution. Build the infrastructure that controls and secures it. The Role We are looking for a Senior Software Engineer to build and maintain the identity and authorization features of Gravitee Access Management (AM) - across the AM runtime and the access-management experience in Gamma, Gravitee's next generation product surface. This is a new role. Today, AM engineering is based entirely in Europe. This hire establishes US-hours ownership of Level 3 and Level 4 authentication and authorization incidents, and adds delivery capacity toward AM parity in Gamma - part of building sustainable L3/L4 engineering capability in the US. You will split your time roughly 80% feature delivery and 20% L3/L4 support and bug fixing (it varies week to week), working as an embedded member of the AM team, which is based in Europe. What You Will Be Doing In this role, you will: Design and deliver features end to end, from discovery and technical design through implementation, testing, release, and iteration. Build and maintain identity and authorization features of Gravitee Access Management, across the AM runtime and the AM experience in Gamma. Implement and support OAuth 2.0 and OIDC flows (authorization code + PKCE, client credentials, token exchange), SAML 2.0 as both IdP and SP, SCIM, and FAPI/CIBA/UMA profiles. Work with token and session semantics - JWT, JWKS, key rotation, revocation, introspection, MFA and step-up, WebAuthn/FIDO2, and IdP federation and social login. Keep security behavior and upgrades safe: standards compliance, secure defaults, certificate and secret handling, consent, audit logs, and defenses against token replay, SSRF, and account takeover. Own safe migrations and backward compatibility across MongoDB and JDBC, and support multi-domain, multi-region deployments and login/token endpoint performance. Own US-hours Level 3 and Level 4 escalations for AM customers as part of the L3 pager duty rotation. Use LLMs and AI-assisted development tools thoughtfully for prototyping, implementation, testing, debugging, and exploration, applying sound engineering judgment to validate AI-generated work. Write meaningful automated tests and contribute to reliable delivery practices. • Collaborate with product managers, designers, engineers, and technical leaders - including the AM team based in Europe - to discover effective solutions and improve them through code and design reviews. Share what you learn and help the team make practical choices as identity standards and protocols evolve. Essential Skills We are looking for evidence that you can succeed in the role, whether gained through employment, open-source work, or equivalent practical experience: 5+ years building and running production backend software, on a team that ships and supports its own product; you have personally resolved production incidents. Strong Java experience (C# accepted if the object-oriented depth is there), with Maven and a reactive stack such as Vert.x/RxJava. Deep working knowledge of identity standards: OAuth 2.0 and OIDC flows (authorization code + PKCE, client credentials, token exchange), SAML 2.0, SCIM, and FAPI/CIBA/UMA profiles. • Solid grasp of token and session semantics: JWT, JWKS, rotation, revocation, introspection, MFA/step-up, WebAuthn/FIDO2, and IdP federation. A security-first mindset: secure defaults, certificate and secret handling, audit logging, and awareness of token replay, SSRF, and account-takeover risks. Experience with safe migrations and backward compatibility across persistent data stores such as MongoDB or JDBCbacked relational databases. Git-based workflow, code review, and writing your own automated tests. Hands-on experience using LLMs or AI coding assistants as part of an engineering workflow, combined with the judgment to review and improve their output. Clear communication, collaborative problem-solving, and the ability to take an ambiguous problem through to production. Desired Skills You do not need to match every item. We would be especially interested in experience with: • Experience at an API gateway, proxy, or service-mesh vendor, or on the API platform team of a large company (e.g., Kong, Google Apigee, MuleSoft, Tyk, Solo.io, Traefik, WSO2). Kubernetes operators and CRDs; OpenAPI tooling; service mesh or Envoy experience. Docker, Kubernetes, and cloud-native application delivery. Model Context Protocol (MCP), Agent2Agent (A2A), tool calling, LLM proxies, or other emerging AI protocols and standards. Prior production experience is not required. Building or operating LLM-powered applications, RAG systems, or agentic workflows - especially their security, governance, and observability needs. Open-source software or enterprise developer platforms. Who Thrives at Gravitee Our growth is powered by people who bring passion to what they build, professionalism to how they work, and a commitment to doing things well. You will thrive here if you: • Bring energy and a constructive attitude to the team. • Adapt quickly and enjoy learning unfamiliar technologies and domains. • Take ownership, communicate clearly, and follow through with urgency. • Balance delivery speed with thoughtful engineering judgment. • Start with the customer problem and care about the quality of the experience you create. • Enjoy working in a fast-moving, collaborative, international environment. Life at Gravitee At Gravitee, we invest in humans, not just roles. You'll get: • Salary of $160,000 • Competitive medical coverage. • Pension / 401(k) program options. • Stock options - you build it, you own it. • 25 days of holiday plus in-country national holidays. • Three mental health days and a wellness allowance. • Your birthday off. • A professional development budget to support your growth. • A hybrid work culture with hubs across regions. • Quarterly team events and an annual company offsite. • A collaborative, international company culture. • Opportunities to grow your scope and career as Gravitee grows. At Gravitee, we believe diverse perspectives make better products and stronger teams. No employee or applicant will be treated less favorably on the grounds of sex, marital status, race, color, nationality, ethnic or national origin, disability, gender, sexual orientation, gender identity, age, pregnancy or maternity, marital or civil partner status, religion, or belief. By applying, you consent to Gravitee storing and processing the personal information you submit as part of the recruitment process.
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE JOB Anduril is seeking a Chief Engineer to lead the design, development, and technical integration for a product line of critical network capabilities. This role is not just about leading; it's about re-envisioning the technical future of secure, resilient, and high-bandwidth communication for all of Anduril and the DoW's autonomous platforms and integrated defense solutions. A successful candidate will not only have a strong technical background, but must also possess a vision for leadership that embraces Anduril's mission to disrupt the defense technology ecosystem across technical and programmatic domains. As Chief Engineer, you will be responsible for the end-to-end technical success and architectural coherence of complex, multi-disciplinary network systems. You will operate at the intersection of complex engineering challenges in embedded systems, networking protocols, mechanical integration, and strategic program execution. You will work directly with our most senior technical leaders, program managers, and customers to define, build, and deliver cutting-edge networking capabilities that impact national security. WHAT YOU'LL DO Technical Vision & Architecture: Own and define the comprehensive technical architecture, design principles, and engineering roadmap for complex, distributed network systems. Ensure technical consistency, scalability, performance, and cyber-resilience across all technical domains (embedded software, UI/UX design, network protocols, hardware integration, ruggedized mechanical designs). System Integrity & Risk Management: Serve as the ultimate technical arbiter, ensuring the integrity, reliability, security, and survivability of our network systems in contested and high security environments. Proactively identify and mitigate technical risks related to performance, security, and physical robustness. Engineering Leadership & Mentorship: Provide decisive technical leadership and guidance to large, multi-disciplinary engineering teams (Mechanical, Electrical, Software) to deliver high quality products that can be manufactured at scale. Foster a culture of technical excellence, rigorous design, rapid iteration, and continuous improvement, ensuring seamless collaboration across disciplines. Strategic Technical Guidance: Translate strategic objectives and customer needs into actionable technical requirements and solutions. Drive trade studies, technology selections, and critical design decisions related to network topology, protocol selection, embedded hardware integration, power management, thermal management, and software-defined networking. Stakeholder Engagement: Act as the primary technical interface with internal stakeholders (Product, Program Management, Executive Leadership) and external stakeholders (customers, partners, regulatory bodies), clearly articulating complex technical concepts, challenges, and solutions for network capabilities. Rapid Prototyping, Fielding, and Scaled Production: Drive aggressive timelines for design, development, integration, and testing of network systems, supporting rapid prototyping, fielding, and scaled production of operational capabilities for defense applications. REQUIRED QUALIFICATIONS 15+ years of progressive experience in complex network systems engineering, R&D, and product development, with a significant portion in a leadership or Chief Engineer capacity for multi-disciplinary systems. Deep and broad technical expertise across network system design and implementation, encompassing significant experience in at least two of the following areas, with a strong understanding of the others: Software Engineering: Network protocols (e.g., IP, routing, transport layer), distributed systems, embedded networking, real-time operating systems, cybersecurity for networks, data link layer implementation. Electrical Engineering: Digital communications, power systems for ruggedized network nodes, high-speed digital design for embedded processors, PCB design for communication systems. Mechanical Engineering: Ruggedized enclosure design, thermal management for radios and processors, vibration/shock mitigation, integration of network hardware onto various platforms (airborne, ground, maritime, soldier-worn). Cross-disciplinary systems integration experience, particularly bridging the gap between hardware and software for robust network performance in demanding environments. Proven ability to define, architect, and lead the development of highly complex, integrated network systems from concept through deployment and sustainment. Exceptional leadership skills with a track record of building, mentoring, and inspiring high-performing multi-disciplinary technical teams in fast-paced, challenging environments. Demonstrated experience in strategic technical planning, roadmapping, and managing technical debt while delivering on aggressive schedules for networking products. Strong problem-solving acumen, with the ability to break down highly ambiguous technical challenges related to network performance and resilience, and drive to practical, innovative solutions. Excellent communication, presentation, and interpersonal skills to effectively convey complex technical concepts to diverse audiences, including executives, customers, and technical teams across different engineering disciplines. Bachelor's degree in Computer Science, Electrical Engineering, Mechanical Engineering, Aerospace Engineering, Robotics, or a related technical field. Master's or Ph.D. preferred. Minimum Secret Security Clearance required. (Active TS/SCI strongly preferred). PREFERRED QUALIFICATIONS Background in cybersecurity for network infrastructure and resilience against near-peer nation state cyber risks. Experience with network modeling, simulation, and analysis tools for predicting performance in diverse environments. Demonstrated ability to drive technical decisions involving significant financial and programmatic impact for network system development. Familiarity with military communication standards and requirements (e.g., Joint All-Domain Command and Control (JADC2) concepts). US Salary Range $254,000-$336,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: . click apply for full job details
09/23/2026
Full time
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE JOB Anduril is seeking a Chief Engineer to lead the design, development, and technical integration for a product line of critical network capabilities. This role is not just about leading; it's about re-envisioning the technical future of secure, resilient, and high-bandwidth communication for all of Anduril and the DoW's autonomous platforms and integrated defense solutions. A successful candidate will not only have a strong technical background, but must also possess a vision for leadership that embraces Anduril's mission to disrupt the defense technology ecosystem across technical and programmatic domains. As Chief Engineer, you will be responsible for the end-to-end technical success and architectural coherence of complex, multi-disciplinary network systems. You will operate at the intersection of complex engineering challenges in embedded systems, networking protocols, mechanical integration, and strategic program execution. You will work directly with our most senior technical leaders, program managers, and customers to define, build, and deliver cutting-edge networking capabilities that impact national security. WHAT YOU'LL DO Technical Vision & Architecture: Own and define the comprehensive technical architecture, design principles, and engineering roadmap for complex, distributed network systems. Ensure technical consistency, scalability, performance, and cyber-resilience across all technical domains (embedded software, UI/UX design, network protocols, hardware integration, ruggedized mechanical designs). System Integrity & Risk Management: Serve as the ultimate technical arbiter, ensuring the integrity, reliability, security, and survivability of our network systems in contested and high security environments. Proactively identify and mitigate technical risks related to performance, security, and physical robustness. Engineering Leadership & Mentorship: Provide decisive technical leadership and guidance to large, multi-disciplinary engineering teams (Mechanical, Electrical, Software) to deliver high quality products that can be manufactured at scale. Foster a culture of technical excellence, rigorous design, rapid iteration, and continuous improvement, ensuring seamless collaboration across disciplines. Strategic Technical Guidance: Translate strategic objectives and customer needs into actionable technical requirements and solutions. Drive trade studies, technology selections, and critical design decisions related to network topology, protocol selection, embedded hardware integration, power management, thermal management, and software-defined networking. Stakeholder Engagement: Act as the primary technical interface with internal stakeholders (Product, Program Management, Executive Leadership) and external stakeholders (customers, partners, regulatory bodies), clearly articulating complex technical concepts, challenges, and solutions for network capabilities. Rapid Prototyping, Fielding, and Scaled Production: Drive aggressive timelines for design, development, integration, and testing of network systems, supporting rapid prototyping, fielding, and scaled production of operational capabilities for defense applications. REQUIRED QUALIFICATIONS 15+ years of progressive experience in complex network systems engineering, R&D, and product development, with a significant portion in a leadership or Chief Engineer capacity for multi-disciplinary systems. Deep and broad technical expertise across network system design and implementation, encompassing significant experience in at least two of the following areas, with a strong understanding of the others: Software Engineering: Network protocols (e.g., IP, routing, transport layer), distributed systems, embedded networking, real-time operating systems, cybersecurity for networks, data link layer implementation. Electrical Engineering: Digital communications, power systems for ruggedized network nodes, high-speed digital design for embedded processors, PCB design for communication systems. Mechanical Engineering: Ruggedized enclosure design, thermal management for radios and processors, vibration/shock mitigation, integration of network hardware onto various platforms (airborne, ground, maritime, soldier-worn). Cross-disciplinary systems integration experience, particularly bridging the gap between hardware and software for robust network performance in demanding environments. Proven ability to define, architect, and lead the development of highly complex, integrated network systems from concept through deployment and sustainment. Exceptional leadership skills with a track record of building, mentoring, and inspiring high-performing multi-disciplinary technical teams in fast-paced, challenging environments. Demonstrated experience in strategic technical planning, roadmapping, and managing technical debt while delivering on aggressive schedules for networking products. Strong problem-solving acumen, with the ability to break down highly ambiguous technical challenges related to network performance and resilience, and drive to practical, innovative solutions. Excellent communication, presentation, and interpersonal skills to effectively convey complex technical concepts to diverse audiences, including executives, customers, and technical teams across different engineering disciplines. Bachelor's degree in Computer Science, Electrical Engineering, Mechanical Engineering, Aerospace Engineering, Robotics, or a related technical field. Master's or Ph.D. preferred. Minimum Secret Security Clearance required. (Active TS/SCI strongly preferred). PREFERRED QUALIFICATIONS Background in cybersecurity for network infrastructure and resilience against near-peer nation state cyber risks. Experience with network modeling, simulation, and analysis tools for predicting performance in diverse environments. Demonstrated ability to drive technical decisions involving significant financial and programmatic impact for network system development. Familiarity with military communication standards and requirements (e.g., Joint All-Domain Command and Control (JADC2) concepts). US Salary Range $254,000-$336,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: . click apply for full job details
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
We exist to wow our customers. We know we're doing the right thing when we hear our customers say, "How did I ever live without Coupang?" Born out of an obsession to make shopping, eating, and living easier than ever, we're collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds - a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what's possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Search and Discovery Product The Search and Discovery Product Team is responsible for creating a world class search and discovery shopping experience for our customers, from "intent" based shopping to delighting the customer with serendipitous product discovery on every visit. We are obsessed with optimizing the entire search customer journey, from query understanding and search relevance to personalized and category/mission aware search and recommendation experiences. We believe the experience of finding the right fresh groceries delivered to your door the next morning, should be personal and different from shopping for a new laptop or your next pair of jeans. Search and Recommendations drive over 70% of product discovery for our millions of daily customers, so the scale and impact of our work is massive! Role Overview With the latest advances in AI/ML, we have the opportunity to rethink the traditional search stack to deliver unique high-quality experiences to our customers that were traditionally outside the realm of possibility. This Search Product leader will be responsible for a key area of Search within the Search & Discovery org as a senior IC or Tech lead who may mentor some junior PMs along with their own IC scope. What You Will Do Strategic Impact: Define and execute on the vision, strategy, and roadmap for Search - driving quarterly planning and driving daily execution to deliver unique customer experiences. Goal Ownership & Execution: Partner with business and technology stakeholders to define OKRs and identify key problems that unlock targeted growth. Own and deliver OKRs by defining product roadmaps and driving day-to-day execution, ensuring high-quality and timely delivery. Metrics Definition & Data-Driven Insights: Define input and output metrics, regularly reviewing them to proactively gain insights and improve search. Conduct deep data analysis to generate actionable insights for defining and solving complex business problems. Competitive Analysis & Benchmarking: Perform competitive analysis to understand how leading eCommerce businesses solve problems and benchmark Coupang's products and practices against global best-in-class standards, identifying opportunities for improvement. Cross-Functional Collaboration: Effectively communicate, collaborate, and drive alignment with diverse stakeholders, including senior executives across business and technology partners, ensuring seamless execution. Team Development & Culture: Conduct product manager interviews and participate in hiring decisions to attract top talent. Provide performance feedback for peers and partners, fostering a culture of excellence and continuous improvement. Basic Qualification Experience with Search/Ads or large scale ML systems Hands on experience vibe coding and building prototypes with LLMs MBA with a Bachelor's degree in Computer Science, Engineering, or equivalent experience. 10+ years of experience in ML product management, including developing, delivering, and executing on a strong product vision. AI first analytical thinker who works well in a global, fast-paced, data-driven, cross-functional, and highly ambiguous entrepreneurial culture. Bias for action and willingness to dive deep and get your hands dirty to get the job done. Sound business judgment and proven ability to effectively communicate and influence others, including executive level reviews. Ability to think strategically and at the same time stay on top of tactical execution. Preferred Qualification Background in Search/Ads ML Experience in launching and scaling consumer facing products, with ecommerce experience a definite plus Experience working with multiple teams across different geographies Pay & Benefits Our compensation reflects the cost of living across several US geographic markets. At Coupang, your base pay is one part of your total compensation. The base pay for this position ranges from $232,000/year to $290,000/year. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience. General Description of All Benefits Medical/Dental/Vision/Life, AD&D insurance Flexible Spending Accounts (FSA) & Health Savings Account (HSA) Long-term/Short-term Disability Employee Assistance Program (EAP) program 401K Plan with Company Match 18-21 days of the Paid Time Off (PTO) a year based on the tenure 12 Paid Holidays Up to 6 weeks of Paid Parental leave Pre-tax commuter benefits MTV - Free Electric Car Charging Station General Description of Other Compensation "Other Compensation" includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the hired applicant in addition to their established salary range or wage scale. Recruitment Process and Others Recruitment Process Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview - Offer The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage Details to Consider This job posting may be closed prior to the stated end date for application if all openings are filled. Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. Those eligible for employment protection (recipients of veteran's benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below: Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to actual or perceived race (including traits historically associated with race, including but not limited to hair texture and protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender (including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status (including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate. Coupang is also committed to providing a safe work environment for its employees and its consumers. If you need assistance and/or a reasonable accommodation in the application of recruiting process due to a disability, please contact us at . Job Requisition ID : R
09/23/2026
Full time
We exist to wow our customers. We know we're doing the right thing when we hear our customers say, "How did I ever live without Coupang?" Born out of an obsession to make shopping, eating, and living easier than ever, we're collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds - a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what's possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Search and Discovery Product The Search and Discovery Product Team is responsible for creating a world class search and discovery shopping experience for our customers, from "intent" based shopping to delighting the customer with serendipitous product discovery on every visit. We are obsessed with optimizing the entire search customer journey, from query understanding and search relevance to personalized and category/mission aware search and recommendation experiences. We believe the experience of finding the right fresh groceries delivered to your door the next morning, should be personal and different from shopping for a new laptop or your next pair of jeans. Search and Recommendations drive over 70% of product discovery for our millions of daily customers, so the scale and impact of our work is massive! Role Overview With the latest advances in AI/ML, we have the opportunity to rethink the traditional search stack to deliver unique high-quality experiences to our customers that were traditionally outside the realm of possibility. This Search Product leader will be responsible for a key area of Search within the Search & Discovery org as a senior IC or Tech lead who may mentor some junior PMs along with their own IC scope. What You Will Do Strategic Impact: Define and execute on the vision, strategy, and roadmap for Search - driving quarterly planning and driving daily execution to deliver unique customer experiences. Goal Ownership & Execution: Partner with business and technology stakeholders to define OKRs and identify key problems that unlock targeted growth. Own and deliver OKRs by defining product roadmaps and driving day-to-day execution, ensuring high-quality and timely delivery. Metrics Definition & Data-Driven Insights: Define input and output metrics, regularly reviewing them to proactively gain insights and improve search. Conduct deep data analysis to generate actionable insights for defining and solving complex business problems. Competitive Analysis & Benchmarking: Perform competitive analysis to understand how leading eCommerce businesses solve problems and benchmark Coupang's products and practices against global best-in-class standards, identifying opportunities for improvement. Cross-Functional Collaboration: Effectively communicate, collaborate, and drive alignment with diverse stakeholders, including senior executives across business and technology partners, ensuring seamless execution. Team Development & Culture: Conduct product manager interviews and participate in hiring decisions to attract top talent. Provide performance feedback for peers and partners, fostering a culture of excellence and continuous improvement. Basic Qualification Experience with Search/Ads or large scale ML systems Hands on experience vibe coding and building prototypes with LLMs MBA with a Bachelor's degree in Computer Science, Engineering, or equivalent experience. 10+ years of experience in ML product management, including developing, delivering, and executing on a strong product vision. AI first analytical thinker who works well in a global, fast-paced, data-driven, cross-functional, and highly ambiguous entrepreneurial culture. Bias for action and willingness to dive deep and get your hands dirty to get the job done. Sound business judgment and proven ability to effectively communicate and influence others, including executive level reviews. Ability to think strategically and at the same time stay on top of tactical execution. Preferred Qualification Background in Search/Ads ML Experience in launching and scaling consumer facing products, with ecommerce experience a definite plus Experience working with multiple teams across different geographies Pay & Benefits Our compensation reflects the cost of living across several US geographic markets. At Coupang, your base pay is one part of your total compensation. The base pay for this position ranges from $232,000/year to $290,000/year. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience. General Description of All Benefits Medical/Dental/Vision/Life, AD&D insurance Flexible Spending Accounts (FSA) & Health Savings Account (HSA) Long-term/Short-term Disability Employee Assistance Program (EAP) program 401K Plan with Company Match 18-21 days of the Paid Time Off (PTO) a year based on the tenure 12 Paid Holidays Up to 6 weeks of Paid Parental leave Pre-tax commuter benefits MTV - Free Electric Car Charging Station General Description of Other Compensation "Other Compensation" includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the hired applicant in addition to their established salary range or wage scale. Recruitment Process and Others Recruitment Process Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview - Offer The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage Details to Consider This job posting may be closed prior to the stated end date for application if all openings are filled. Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. Those eligible for employment protection (recipients of veteran's benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below: Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to actual or perceived race (including traits historically associated with race, including but not limited to hair texture and protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender (including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status (including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate. Coupang is also committed to providing a safe work environment for its employees and its consumers. If you need assistance and/or a reasonable accommodation in the application of recruiting process due to a disability, please contact us at . Job Requisition ID : R
Sr. Director of Product Managment - Search & Discovery We exist to wow our customers. We know we're doing the right thing when we hear ourcustomers say, "How did we ever live without Coupang?" Born out of an obsession to make shopping, eating, and living easier than ever, we're collectively disrupting the multi-billion- dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds - a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what's possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Search and Discovery Product The Search and Discovery Product Team is responsible for creating a world class search and discovery shopping experience for our customers, from "intent" based shopping to delighting the customer with serendipitous product discovery on every visit. We are obsessed with optimizing the entire search customer journey, from query understanding and search relevance to personalized and category/mission aware search and recommendation experiences. We believe the experience for finding the right fresh groceries delivered to your door the next morning, should be personal and different from shopping for a new laptop or your next pair of jeans. Search and Recommendations drive over 70% of product discovery for our millions of daily customers, so the scale and impact of our work is massive! About the Role With the latest advances in AI/ML, we have the opportunity to rethink the traditional search stack to deliver unique high-quality experiences to our customers that were traditionally outside the realm of possibility. This Search Product Leader role will be a hand on role with the potential of building out a team of product managers focusing on Search ML stack. Responsibilities will include the following: Define the vision, strategy, and roadmap to re-imagine our search and discovery stack to deliver unique customer experiences while directing day-to-day execution towards that vision. Hire and develop a high performing team of Product Managers with deep expertise in ML based systems. Background in Search ML is a definite plus. Regularly review metrics and proactively seek out new and improved mechanisms for visibility on performance and greater optimization of the Search and Discovery product portfolio. Collaborate and drive alignment with a variety of stakeholders including senior executives. Key Experience Bachelor's degree in Computer Science, Engineering, or equivalent experience. 10+ years of experience in ML product management, including developing, delivering, and executing on a strong product vision. Experience working on Search and Discovery customer experiences/ML stack. Analytical thinker who works well in a global, fast-paced, data-driven, cross-functional, and highly ambiguous entrepreneurial culture. Bias for action and willingness to get your hands dirty to get the job done. Sound business judgment and proven ability to effectively communicate and influence others, including executive level reviews. Ability to think strategically and at the same time stay on top of tactical execution. Preferred e-Commerce experience Experience in launching and scaling consumer facing products Experience working with multiple teams across different geographies Pay & Benefits Our compensation reflects the cost of labor across several US geographic markets. At Coupang, your base pay is one part of your total compensation. The base pay for this position ranges from $207,900/year in our lowest geographic market to $386,100/year in our highest geographic market. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience. General Description of All Benefits Medical/Dental/Vision/Life, AD&D insurance Flexible Spending Accounts (FSA) & Health Savings Account (HSA) Long-term/Short-term Disability Employee Assistance Program (EAP) program 401K Plan with Company Match 18-21 days of the Paid Time Off (PTO) a year based on the tenure 12 Public Holidays Paid Parental leave Pre-tax commuter benefits MTV - Free Electric Car Charging Station General Description of Other Compensation "Other Compensation" includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the hired applicant in addition to their established salary range or wage scale. Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to actual or perceived race (including traits historically associated with race, including but not limited to hair texture and protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender (including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status (including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate. Coupang is also committed to providing a safe work environment for its employees and its consumers If you need assistance and/or a reasonable accommodation in the application of recruiting process due to a disability, please contact us at .
09/23/2026
Full time
Sr. Director of Product Managment - Search & Discovery We exist to wow our customers. We know we're doing the right thing when we hear ourcustomers say, "How did we ever live without Coupang?" Born out of an obsession to make shopping, eating, and living easier than ever, we're collectively disrupting the multi-billion- dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds - a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what's possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Search and Discovery Product The Search and Discovery Product Team is responsible for creating a world class search and discovery shopping experience for our customers, from "intent" based shopping to delighting the customer with serendipitous product discovery on every visit. We are obsessed with optimizing the entire search customer journey, from query understanding and search relevance to personalized and category/mission aware search and recommendation experiences. We believe the experience for finding the right fresh groceries delivered to your door the next morning, should be personal and different from shopping for a new laptop or your next pair of jeans. Search and Recommendations drive over 70% of product discovery for our millions of daily customers, so the scale and impact of our work is massive! About the Role With the latest advances in AI/ML, we have the opportunity to rethink the traditional search stack to deliver unique high-quality experiences to our customers that were traditionally outside the realm of possibility. This Search Product Leader role will be a hand on role with the potential of building out a team of product managers focusing on Search ML stack. Responsibilities will include the following: Define the vision, strategy, and roadmap to re-imagine our search and discovery stack to deliver unique customer experiences while directing day-to-day execution towards that vision. Hire and develop a high performing team of Product Managers with deep expertise in ML based systems. Background in Search ML is a definite plus. Regularly review metrics and proactively seek out new and improved mechanisms for visibility on performance and greater optimization of the Search and Discovery product portfolio. Collaborate and drive alignment with a variety of stakeholders including senior executives. Key Experience Bachelor's degree in Computer Science, Engineering, or equivalent experience. 10+ years of experience in ML product management, including developing, delivering, and executing on a strong product vision. Experience working on Search and Discovery customer experiences/ML stack. Analytical thinker who works well in a global, fast-paced, data-driven, cross-functional, and highly ambiguous entrepreneurial culture. Bias for action and willingness to get your hands dirty to get the job done. Sound business judgment and proven ability to effectively communicate and influence others, including executive level reviews. Ability to think strategically and at the same time stay on top of tactical execution. Preferred e-Commerce experience Experience in launching and scaling consumer facing products Experience working with multiple teams across different geographies Pay & Benefits Our compensation reflects the cost of labor across several US geographic markets. At Coupang, your base pay is one part of your total compensation. The base pay for this position ranges from $207,900/year in our lowest geographic market to $386,100/year in our highest geographic market. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience. General Description of All Benefits Medical/Dental/Vision/Life, AD&D insurance Flexible Spending Accounts (FSA) & Health Savings Account (HSA) Long-term/Short-term Disability Employee Assistance Program (EAP) program 401K Plan with Company Match 18-21 days of the Paid Time Off (PTO) a year based on the tenure 12 Public Holidays Paid Parental leave Pre-tax commuter benefits MTV - Free Electric Car Charging Station General Description of Other Compensation "Other Compensation" includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the hired applicant in addition to their established salary range or wage scale. Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to actual or perceived race (including traits historically associated with race, including but not limited to hair texture and protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender (including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status (including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate. Coupang is also committed to providing a safe work environment for its employees and its consumers If you need assistance and/or a reasonable accommodation in the application of recruiting process due to a disability, please contact us at .
As a Senior DevOps Engineer, you will be a key technical contributor responsible for designing, implementing, and operating scalable, resilient infrastructure and CI/CD pipelines that support the full software development lifecycle. You will work closely with Wolters Kluwer Product Teams to embed DevOps best practices into agile workflows, enabling continuous integration, automated testing, and reliable deployment across environments. In this role, you will focus on hands on engineering excellence-building, automating, and operating cloud native platforms that improve system reliability, performance, and maintainability. You will partner closely with application development teams to bridge development and operations, supporting microservices, containerized workloads, and cloud native architectures. While not a formal people manager, you will act as a senior technical mentor and role model, promoting engineering rigor, code as infrastructure, and continuous improvement. Responsibilities Design, engineer, and automate secure, scalable cloud infrastructure in Azure and AWS using Infrastructure as Code (IaC) tools such as Terraform, Ansible, and Jenkins, applying software engineering best practices including modular design, version control, and automated testing. Implement and maintain Infrastructure as Code and CI/CD pipelines, contributing reusable modules, templates, and patterns that improve consistency and reliability across teams. Design and implement modern compute platforms, including containerized and serverless solutions (AKS, EKS, Docker, Azure Functions), with an emphasis on scalability, maintainability, and performance. Build and maintain CI/CD pipelines as software products, ensuring strong test coverage, artifact management, promotion workflows, and deployment automation across multiple environments. Support and evolve cloud native architectures, applying engineering principles such as abstraction, decoupling, fault isolation, and observability. Implement observability solutions using metrics, logging, and tracing to enable proactive issue detection, faster troubleshooting, and root cause analysis. Ensure infrastructure and automation solutions comply with enterprise DevOps, security, and compliance standards, contributing to architectural reviews and governance processes. Serve as a senior technical mentor, providing guidance through code reviews, design discussions, and knowledge sharing-without direct people management responsibilities. Evaluate and prototype emerging tools and technologies, applying engineering rigor to assess value, performance, and integration feasibility. Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity planning, and incident response to improve system reliability and reduce operational toil. Deploy, operate, and support business critical applications, ensuring high availability, fault tolerance, and performance optimization. Participate in modernization initiatives, supporting the re architecture and cloud native transformation of legacy platforms. Identify and remediate engineering inefficiencies by proposing and implementing automation and architectural improvements. Participate in post incident reviews, contributing to blameless root cause analysis and long term corrective actions. Participate in on call rotations, continuously improving alert quality, reducing noise, and automating remediation where possible. Qualifications Bachelor's degree in Engineering, Computer Science, or a related field (Master's degree preferred). 5+ years of experience in DevOps, Site Reliability Engineering, Release Engineering, or related roles, with strong hands on software engineering experience. Strong background in software development, with experience in languages such as Python, .NET, or Java. Proven experience working with cloud platforms (Azure and AWS). Proficiency in scripting languages such as PowerShell and Bash. Solid understanding of core Azure and AWS services (PaaS, IaaS, SaaS). Strong experience with source control and automation tools, including Git. Hands on experience with Infrastructure as Code tools such as Terraform or CloudFormation. Experience building and operating CI/CD pipelines using tools such as Azure DevOps, Jenkins, or similar platforms. Strong problem solving skills with attention to detail and operational excellence. Ability to clearly communicate technical concepts to engineers and non engineering stakeholders. Demonstrated commitment to DevOps culture, including continuous integration, automated testing, deployment automation, and full lifecycle ownership. Experience building or supporting observability platforms and defining operational best practices. Experience troubleshooting and automating diagnostics across Linux and Windows environments. Our Interview Practices To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we're getting to know you-not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process. Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process. Compensation: $92,700.00 - $161,850.00 USDThis role is eligible for Bonus. Compensation range listed is based on primary location of the position. Actual base salary offer is influenced by a wide array of factors including but not limited to skills, experience and actual hiring location. Your recruiter can share more information about the specific offer for the job location during the hiring process. Additional Information: Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life, including but not limited to: Medical, Dental, & Vision Plans, 401(k), FSA/HSA, Commuter Benefits, Tuition Assistance Plan, Vacation and Sick Time, and Paid Parental Leave. Full details of our benefits are available upon request.
09/23/2026
Full time
As a Senior DevOps Engineer, you will be a key technical contributor responsible for designing, implementing, and operating scalable, resilient infrastructure and CI/CD pipelines that support the full software development lifecycle. You will work closely with Wolters Kluwer Product Teams to embed DevOps best practices into agile workflows, enabling continuous integration, automated testing, and reliable deployment across environments. In this role, you will focus on hands on engineering excellence-building, automating, and operating cloud native platforms that improve system reliability, performance, and maintainability. You will partner closely with application development teams to bridge development and operations, supporting microservices, containerized workloads, and cloud native architectures. While not a formal people manager, you will act as a senior technical mentor and role model, promoting engineering rigor, code as infrastructure, and continuous improvement. Responsibilities Design, engineer, and automate secure, scalable cloud infrastructure in Azure and AWS using Infrastructure as Code (IaC) tools such as Terraform, Ansible, and Jenkins, applying software engineering best practices including modular design, version control, and automated testing. Implement and maintain Infrastructure as Code and CI/CD pipelines, contributing reusable modules, templates, and patterns that improve consistency and reliability across teams. Design and implement modern compute platforms, including containerized and serverless solutions (AKS, EKS, Docker, Azure Functions), with an emphasis on scalability, maintainability, and performance. Build and maintain CI/CD pipelines as software products, ensuring strong test coverage, artifact management, promotion workflows, and deployment automation across multiple environments. Support and evolve cloud native architectures, applying engineering principles such as abstraction, decoupling, fault isolation, and observability. Implement observability solutions using metrics, logging, and tracing to enable proactive issue detection, faster troubleshooting, and root cause analysis. Ensure infrastructure and automation solutions comply with enterprise DevOps, security, and compliance standards, contributing to architectural reviews and governance processes. Serve as a senior technical mentor, providing guidance through code reviews, design discussions, and knowledge sharing-without direct people management responsibilities. Evaluate and prototype emerging tools and technologies, applying engineering rigor to assess value, performance, and integration feasibility. Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity planning, and incident response to improve system reliability and reduce operational toil. Deploy, operate, and support business critical applications, ensuring high availability, fault tolerance, and performance optimization. Participate in modernization initiatives, supporting the re architecture and cloud native transformation of legacy platforms. Identify and remediate engineering inefficiencies by proposing and implementing automation and architectural improvements. Participate in post incident reviews, contributing to blameless root cause analysis and long term corrective actions. Participate in on call rotations, continuously improving alert quality, reducing noise, and automating remediation where possible. Qualifications Bachelor's degree in Engineering, Computer Science, or a related field (Master's degree preferred). 5+ years of experience in DevOps, Site Reliability Engineering, Release Engineering, or related roles, with strong hands on software engineering experience. Strong background in software development, with experience in languages such as Python, .NET, or Java. Proven experience working with cloud platforms (Azure and AWS). Proficiency in scripting languages such as PowerShell and Bash. Solid understanding of core Azure and AWS services (PaaS, IaaS, SaaS). Strong experience with source control and automation tools, including Git. Hands on experience with Infrastructure as Code tools such as Terraform or CloudFormation. Experience building and operating CI/CD pipelines using tools such as Azure DevOps, Jenkins, or similar platforms. Strong problem solving skills with attention to detail and operational excellence. Ability to clearly communicate technical concepts to engineers and non engineering stakeholders. Demonstrated commitment to DevOps culture, including continuous integration, automated testing, deployment automation, and full lifecycle ownership. Experience building or supporting observability platforms and defining operational best practices. Experience troubleshooting and automating diagnostics across Linux and Windows environments. Our Interview Practices To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we're getting to know you-not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process. Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process. Compensation: $92,700.00 - $161,850.00 USDThis role is eligible for Bonus. Compensation range listed is based on primary location of the position. Actual base salary offer is influenced by a wide array of factors including but not limited to skills, experience and actual hiring location. Your recruiter can share more information about the specific offer for the job location during the hiring process. Additional Information: Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life, including but not limited to: Medical, Dental, & Vision Plans, 401(k), FSA/HSA, Commuter Benefits, Tuition Assistance Plan, Vacation and Sick Time, and Paid Parental Leave. Full details of our benefits are available upon request.