RELOCATION ASSISTANCE: Relocation assistance may be available CLEARANCE REQUIRED FOR START: Yes CLEARANCE TYPE: Secret TRAVEL: Yes, 10% of the Time Description At Northrop Grumman, our employees have incredible opportunities to work on revolutionary systems that impact people's lives around the world today, and for generations to come. Our pioneering and inventive spirit has enabled us to be at the forefront of many technological advancements in our nation's history - from the first flight across the Atlantic Ocean, to stealth bombers, to landing on the moon. We look for people who have bold new ideas, courage and a pioneering spirit to join forces to invent the future, and have fun along the way. Our culture thrives on intellectual curiosity, cognitive diversity and bringing your whole self to work - and we have an insatiable drive to do what others think is impossible. Our employees are not only part of history, they're making history. Northrop Grumman Space Systems is seeking a Staff Software Signal Processing Engineer to join our team supporting a radar program in Colorado Springs, Colorado. In this staff-level role, you will provide technical leadership for the development of advanced radar digital signal processing software in support of identification and tracking of objects in geosynchronous orbit. You will act as a senior technical authority for radar DSP software-guiding architecture, design decisions, and implementation approaches across the team-while remaining hands-on with code and technical problem solving. Job responsibilities will include, but are not limited to, the following: Serve as a technical leader and subject-matter expert for radar digital signal processing software, providing guidance on system architecture, algorithm selection, implementation strategies, and performance optimization. Design, develop, document, test, and debug complex applications software and software systems that contain logical and mathematical solutions, with particular emphasis on high performance radar DSP pipelines. Lead multidisciplinary research and collaborate closely with systems, hardware, and equipment designers to plan, design, develop, and utilize electronic data processing systems for mission and product software. Work with stakeholders to determine computer user and mission needs; analyze system capabilities to resolve problems on program intent, output requirements, input data acquisition, programming techniques, and controls; prepare operating instructions and technical documentation. Guide the design and development of compilers, utility programs, and/or supporting infrastructure where needed to enable scalable signal processing workflows. Provide mentorship and technical direction to other engineers, including code reviews, design reviews, and coaching on best practices for performance, reliability, and maintainability. Demonstrated experience developing and optimizing radar signal processing algorithms and pipelines (DSP), including use of common engineering toolchains such as DSP System Toolbox / Signal Processing Toolbox / Phased Array System toolsets. Demonstrated experience with GPGPU acceleration (e.g., CUDA) for performance critical signal processing and/or analytics workloads, including profiling, optimization, and balancing CPU/GPU resources. Experience using AI enabled engineering tooling (e.g., an AI coding assistant) to accelerate development tasks such as code explanation, refactoring, and test generation-while maintaining clear engineer accountability for reviewed code and technical decisions. Basic Qualifications: Must have an active U.S. Government Secret security clearance at time of application, current and within scope. Demonstrated proficiency with C++, Java, Python, and CUDA, including application to performance critical workloads. Experience with Linux development environments. Experience with RF signal processing and radar DSP concepts. Experience with GPU optimization for real time or near real time processing. Experience with NVIDIA MatX or similar high performance numerical libraries. Experience leveraging AI tooling for software development workflows. Experience with high-side development and processes, including secure environments and appropriate handling of classified data. Preferred Qualifications: Able to explain how to process a radar signal to generate a Range-Doppler Map. Describe the process of using a sparse array to accomplish beamforming. Discuss the concept of loss budgets and system tradeoffs. Explain the application of calibration for a distributed system. C++ expertise: Strong knowledge of real time processing, current C++ standards, and modern C++ techniques and patterns for high performance, maintainable code. CUDA expertise: Deep understanding of the benefits and constraints of running signal processing code on GPGPUs; familiarity with cuFFT and/or cuSignal; experience with MatX preferred but not required. Functional familiarity with Matlab and Python for signal processing, prototyping, and data analysis. Self starter and technical leader: Able to take minimal direction, quickly learn complex systems, and independently drive solutions. Approaches every story with the broader system in mind-identifying defects, optimizations, and architectural implications along the way. Comfortable influencing designs across teams and defending technical decisions to stakeholders. Experience writing and implementing signal processing algorithms is expected; at the staff level, candidates should additionally demonstrate the ability to own system level DSP architectures, make trade studies, and lead others in implementing robust, scalable solutions. Primary Level Salary Range: $161,000.00 - $241,400.00 The above salary range represents a general guideline; however, Northrop Grumman considers a number of factors when determining base salary offers such as the scope and responsibilities of the position and the candidate's experience, education, skills and current market conditions. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay. Annual bonuses are designed to reward individual contributions as well as allow employees to share in company results. Employees in Vice President or Director positions may be eligible for Long Term Incentives. In addition, Northrop Grumman provides a variety of benefits including health insurance coverage, life and disability insurance, savings plan, Company paid holidays and paid time off (PTO) for vacation and/or personal business. The application period for the job is estimated to be 20 days from the job posting date. However, this timeline may be shortened or extended depending on business needs and the availability of qualified candidates. Northrop Grumman is an Equal Opportunity Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class. For our complete EEO and pay transparency statement, please visit U.S. Citizenship is required for all positions with a government clearance and certain other restricted positions.
09/25/2026
Full time
RELOCATION ASSISTANCE: Relocation assistance may be available CLEARANCE REQUIRED FOR START: Yes CLEARANCE TYPE: Secret TRAVEL: Yes, 10% of the Time Description At Northrop Grumman, our employees have incredible opportunities to work on revolutionary systems that impact people's lives around the world today, and for generations to come. Our pioneering and inventive spirit has enabled us to be at the forefront of many technological advancements in our nation's history - from the first flight across the Atlantic Ocean, to stealth bombers, to landing on the moon. We look for people who have bold new ideas, courage and a pioneering spirit to join forces to invent the future, and have fun along the way. Our culture thrives on intellectual curiosity, cognitive diversity and bringing your whole self to work - and we have an insatiable drive to do what others think is impossible. Our employees are not only part of history, they're making history. Northrop Grumman Space Systems is seeking a Staff Software Signal Processing Engineer to join our team supporting a radar program in Colorado Springs, Colorado. In this staff-level role, you will provide technical leadership for the development of advanced radar digital signal processing software in support of identification and tracking of objects in geosynchronous orbit. You will act as a senior technical authority for radar DSP software-guiding architecture, design decisions, and implementation approaches across the team-while remaining hands-on with code and technical problem solving. Job responsibilities will include, but are not limited to, the following: Serve as a technical leader and subject-matter expert for radar digital signal processing software, providing guidance on system architecture, algorithm selection, implementation strategies, and performance optimization. Design, develop, document, test, and debug complex applications software and software systems that contain logical and mathematical solutions, with particular emphasis on high performance radar DSP pipelines. Lead multidisciplinary research and collaborate closely with systems, hardware, and equipment designers to plan, design, develop, and utilize electronic data processing systems for mission and product software. Work with stakeholders to determine computer user and mission needs; analyze system capabilities to resolve problems on program intent, output requirements, input data acquisition, programming techniques, and controls; prepare operating instructions and technical documentation. Guide the design and development of compilers, utility programs, and/or supporting infrastructure where needed to enable scalable signal processing workflows. Provide mentorship and technical direction to other engineers, including code reviews, design reviews, and coaching on best practices for performance, reliability, and maintainability. Demonstrated experience developing and optimizing radar signal processing algorithms and pipelines (DSP), including use of common engineering toolchains such as DSP System Toolbox / Signal Processing Toolbox / Phased Array System toolsets. Demonstrated experience with GPGPU acceleration (e.g., CUDA) for performance critical signal processing and/or analytics workloads, including profiling, optimization, and balancing CPU/GPU resources. Experience using AI enabled engineering tooling (e.g., an AI coding assistant) to accelerate development tasks such as code explanation, refactoring, and test generation-while maintaining clear engineer accountability for reviewed code and technical decisions. Basic Qualifications: Must have an active U.S. Government Secret security clearance at time of application, current and within scope. Demonstrated proficiency with C++, Java, Python, and CUDA, including application to performance critical workloads. Experience with Linux development environments. Experience with RF signal processing and radar DSP concepts. Experience with GPU optimization for real time or near real time processing. Experience with NVIDIA MatX or similar high performance numerical libraries. Experience leveraging AI tooling for software development workflows. Experience with high-side development and processes, including secure environments and appropriate handling of classified data. Preferred Qualifications: Able to explain how to process a radar signal to generate a Range-Doppler Map. Describe the process of using a sparse array to accomplish beamforming. Discuss the concept of loss budgets and system tradeoffs. Explain the application of calibration for a distributed system. C++ expertise: Strong knowledge of real time processing, current C++ standards, and modern C++ techniques and patterns for high performance, maintainable code. CUDA expertise: Deep understanding of the benefits and constraints of running signal processing code on GPGPUs; familiarity with cuFFT and/or cuSignal; experience with MatX preferred but not required. Functional familiarity with Matlab and Python for signal processing, prototyping, and data analysis. Self starter and technical leader: Able to take minimal direction, quickly learn complex systems, and independently drive solutions. Approaches every story with the broader system in mind-identifying defects, optimizations, and architectural implications along the way. Comfortable influencing designs across teams and defending technical decisions to stakeholders. Experience writing and implementing signal processing algorithms is expected; at the staff level, candidates should additionally demonstrate the ability to own system level DSP architectures, make trade studies, and lead others in implementing robust, scalable solutions. Primary Level Salary Range: $161,000.00 - $241,400.00 The above salary range represents a general guideline; however, Northrop Grumman considers a number of factors when determining base salary offers such as the scope and responsibilities of the position and the candidate's experience, education, skills and current market conditions. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay. Annual bonuses are designed to reward individual contributions as well as allow employees to share in company results. Employees in Vice President or Director positions may be eligible for Long Term Incentives. In addition, Northrop Grumman provides a variety of benefits including health insurance coverage, life and disability insurance, savings plan, Company paid holidays and paid time off (PTO) for vacation and/or personal business. The application period for the job is estimated to be 20 days from the job posting date. However, this timeline may be shortened or extended depending on business needs and the availability of qualified candidates. Northrop Grumman is an Equal Opportunity Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class. For our complete EEO and pay transparency statement, please visit U.S. Citizenship is required for all positions with a government clearance and certain other restricted positions.
Quick Position Facts! Location: San Diego, CA at our G2 Ops office and customer site. Work Setting: Primarily at the Customer Site. Salary Range: $140,000 - $185,000plus comprehensive benefits package. Years of Industry Experience: 7+ years of relevant experience. Security Clearance Requirement: Must be able to obtain and maintain Active DoD Secret Clearance. About the Role. G2 Ops supports mission-critical defense and technology initiatives by solving complex operational and engineering challenges. We are seeking a motivated, collaborative Senior Systems Engineer to support the PMW 770 Link 16 program and its mission-critical systems engineering activities. In this role, you'll contribute to technical program support, Model-Based Systems Engineering, engineering change management, technical reviews, acquisition documentation, and systems test and evaluation activities while working alongside a team that values innovation, technical excellence, and continuous growth. What You'll Do Provide technical systems engineering support to the Assistant Program Manager for the PMW 770 Link 16 program. Develop technical briefs and deliver presentations to senior government and program stakeholders. Contribute to, review, and validate Model-Based Systems Engineering diagrams and engineering artifacts. Manage and coordinate Engineering Change Requests through applicable review and approval processes. Conduct and support naval systems engineering technical reviews. Draft and maintain systems engineering acquisition documentation and related technical products. Coordinate systems test and evaluation activities with engineering, program, and customer stakeholders. Identify technical issues, communicate engineering impacts, and support informed program decisions. What You'll Bring Required Qualifications Bachelor's degree in an Engineering discipline. 7+ years of relevant professional experience, including experience with enterprise architecture for IT infrastructure management. Experience applying systems engineering principles to complex technical programs or environments. Experience developing technical documentation, engineering briefs, and decision-support materials. Ability to evaluate and contribute to Model-Based Systems Engineering diagrams and technical artifacts. Strong communication and collaboration skills. Ability to communicate technical information effectively to senior stakeholders. Ability to work effectively in a team-oriented and customer-facing environment. Ability to obtain and maintain required security clearance. Preferred Qualifications Experience supporting DoD, Department of the Navy, NAVWAR, or similar government programs. Familiarity with Link 16, tactical data links, or related naval communications systems. Experience with Model-Based Systems Engineering methodologies, tools, or modeling environments. Experience managing Engineering Change Requests or similar configuration and change-management processes. Experience supporting systems engineering technical reviews, acquisition documentation, or systems test and evaluation activities. Experience working in cross-functional or customer-facing environments. Additional Considerations This is a full-time position supporting mission-focused customer environments Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $140,000-$185,000 depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
09/25/2026
Full time
Quick Position Facts! Location: San Diego, CA at our G2 Ops office and customer site. Work Setting: Primarily at the Customer Site. Salary Range: $140,000 - $185,000plus comprehensive benefits package. Years of Industry Experience: 7+ years of relevant experience. Security Clearance Requirement: Must be able to obtain and maintain Active DoD Secret Clearance. About the Role. G2 Ops supports mission-critical defense and technology initiatives by solving complex operational and engineering challenges. We are seeking a motivated, collaborative Senior Systems Engineer to support the PMW 770 Link 16 program and its mission-critical systems engineering activities. In this role, you'll contribute to technical program support, Model-Based Systems Engineering, engineering change management, technical reviews, acquisition documentation, and systems test and evaluation activities while working alongside a team that values innovation, technical excellence, and continuous growth. What You'll Do Provide technical systems engineering support to the Assistant Program Manager for the PMW 770 Link 16 program. Develop technical briefs and deliver presentations to senior government and program stakeholders. Contribute to, review, and validate Model-Based Systems Engineering diagrams and engineering artifacts. Manage and coordinate Engineering Change Requests through applicable review and approval processes. Conduct and support naval systems engineering technical reviews. Draft and maintain systems engineering acquisition documentation and related technical products. Coordinate systems test and evaluation activities with engineering, program, and customer stakeholders. Identify technical issues, communicate engineering impacts, and support informed program decisions. What You'll Bring Required Qualifications Bachelor's degree in an Engineering discipline. 7+ years of relevant professional experience, including experience with enterprise architecture for IT infrastructure management. Experience applying systems engineering principles to complex technical programs or environments. Experience developing technical documentation, engineering briefs, and decision-support materials. Ability to evaluate and contribute to Model-Based Systems Engineering diagrams and technical artifacts. Strong communication and collaboration skills. Ability to communicate technical information effectively to senior stakeholders. Ability to work effectively in a team-oriented and customer-facing environment. Ability to obtain and maintain required security clearance. Preferred Qualifications Experience supporting DoD, Department of the Navy, NAVWAR, or similar government programs. Familiarity with Link 16, tactical data links, or related naval communications systems. Experience with Model-Based Systems Engineering methodologies, tools, or modeling environments. Experience managing Engineering Change Requests or similar configuration and change-management processes. Experience supporting systems engineering technical reviews, acquisition documentation, or systems test and evaluation activities. Experience working in cross-functional or customer-facing environments. Additional Considerations This is a full-time position supporting mission-focused customer environments Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $140,000-$185,000 depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
Quick Position Facts! Location: San Diego, CA. Work Setting: Two openings are available. One position requires regular onsite support at the customer site; the second offers a more flexible/hybrid work arrangement based on program requirements. Salary Range: Mid-Level: $90,000 - $140,000 Senior-Level: $140,000 - $200,000 Final compensation will depend on qualifications, experience, level, and program requirements. Years of Industry Experience: Mid-Level: 3-6 years of relevant experience Senior-Level: 7+ years of relevant experience Security Clearance Requirement: Must be able to obtain and maintain Active DoD Secret Clearance. About the Role. G2 Ops supports mission-critical defense and technology initiatives by solving complex operational and engineering challenges. We are seeking motivated, collaborative Model-Based Systems Engineers (MBSE) to support Digital Engineering initiatives for DoD customers in San Diego. We are hiring for two MBSE openings at different experience levels and with different work arrangements. One role requires regular customer-site support, while the other offers greater flexibility based on program requirements. Depending on your experience and program fit, you may be considered for a mid-level or senior-level opportunity. In this role, you'll apply MBSE methodologies to develop and mature system models, architectures, standards, and engineering products. You may support model development, Digital Engineering implementation, model governance, or Navy programs transitioning from traditional systems engineering to model-based approaches. What You'll Do Develop and maintain system models and architectures using Cameo/MagicDraw, SysML, and related MBSE methodologies. Collaborate with Government stakeholders, Program Executive Offices, and multidisciplinary engineering teams to resolve technical issues and advance MBSE adoption. Improve model consistency, traceability, interoperability, information reuse, and collaboration across engineering teams. Support model governance, configuration management, cybersecurity, systems engineering, and acquisition-related documentation and analysis. Develop and maintain MBSE standards, implementation guidance, modeling conventions, technical documentation, and configuration-managed engineering products. Support Navy programs transitioning from traditional systems engineering to Digital Engineering and model-based approaches. Translate technical requirements, existing engineering documentation, and stakeholder needs into structured models, data requirements, workflows, and technical products. What You'll Bring Required Qualifications Bachelor's degree in Engineering or a related technical discipline. Demonstrated professional experience applying Model-Based Systems Engineering (MBSE) methodologies. Experience developing or maintaining system models using Cameo Systems Modeler/MagicDraw or comparable SysML-based tools. Working knowledge of systems engineering principles, lifecycle processes, and model-based engineering methodologies. Ability to translate engineering requirements and stakeholder needs into structured models, architectures, workflows, and technical documentation. Strong written and verbal communication skills and the ability to collaborate with multidisciplinary engineering and Government teams. Ability to obtain and maintain Secret clearance. Preferred Qualifications Experience supporting NAVWAR, PEOs, Naval acquisition, or other DoD programs. Experience with SysML, UML, executable architectures, or model governance. Experience developing MBSE standards, guides, training, or model catalogs. Familiarity with cybersecurity, configuration management, Digital Thread, or PLM concepts. Experience working in cross-functional or customer-facing environments. Additional Considerations This is a full-time position supporting mission-focused customer environments Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position are: Mid-Level: $90,000 - $140,000 Senior-Level: $140,000 - $200,000 Final compensation will depend on qualifications, experience, level, and program requirements. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
09/25/2026
Full time
Quick Position Facts! Location: San Diego, CA. Work Setting: Two openings are available. One position requires regular onsite support at the customer site; the second offers a more flexible/hybrid work arrangement based on program requirements. Salary Range: Mid-Level: $90,000 - $140,000 Senior-Level: $140,000 - $200,000 Final compensation will depend on qualifications, experience, level, and program requirements. Years of Industry Experience: Mid-Level: 3-6 years of relevant experience Senior-Level: 7+ years of relevant experience Security Clearance Requirement: Must be able to obtain and maintain Active DoD Secret Clearance. About the Role. G2 Ops supports mission-critical defense and technology initiatives by solving complex operational and engineering challenges. We are seeking motivated, collaborative Model-Based Systems Engineers (MBSE) to support Digital Engineering initiatives for DoD customers in San Diego. We are hiring for two MBSE openings at different experience levels and with different work arrangements. One role requires regular customer-site support, while the other offers greater flexibility based on program requirements. Depending on your experience and program fit, you may be considered for a mid-level or senior-level opportunity. In this role, you'll apply MBSE methodologies to develop and mature system models, architectures, standards, and engineering products. You may support model development, Digital Engineering implementation, model governance, or Navy programs transitioning from traditional systems engineering to model-based approaches. What You'll Do Develop and maintain system models and architectures using Cameo/MagicDraw, SysML, and related MBSE methodologies. Collaborate with Government stakeholders, Program Executive Offices, and multidisciplinary engineering teams to resolve technical issues and advance MBSE adoption. Improve model consistency, traceability, interoperability, information reuse, and collaboration across engineering teams. Support model governance, configuration management, cybersecurity, systems engineering, and acquisition-related documentation and analysis. Develop and maintain MBSE standards, implementation guidance, modeling conventions, technical documentation, and configuration-managed engineering products. Support Navy programs transitioning from traditional systems engineering to Digital Engineering and model-based approaches. Translate technical requirements, existing engineering documentation, and stakeholder needs into structured models, data requirements, workflows, and technical products. What You'll Bring Required Qualifications Bachelor's degree in Engineering or a related technical discipline. Demonstrated professional experience applying Model-Based Systems Engineering (MBSE) methodologies. Experience developing or maintaining system models using Cameo Systems Modeler/MagicDraw or comparable SysML-based tools. Working knowledge of systems engineering principles, lifecycle processes, and model-based engineering methodologies. Ability to translate engineering requirements and stakeholder needs into structured models, architectures, workflows, and technical documentation. Strong written and verbal communication skills and the ability to collaborate with multidisciplinary engineering and Government teams. Ability to obtain and maintain Secret clearance. Preferred Qualifications Experience supporting NAVWAR, PEOs, Naval acquisition, or other DoD programs. Experience with SysML, UML, executable architectures, or model governance. Experience developing MBSE standards, guides, training, or model catalogs. Familiarity with cybersecurity, configuration management, Digital Thread, or PLM concepts. Experience working in cross-functional or customer-facing environments. Additional Considerations This is a full-time position supporting mission-focused customer environments Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position are: Mid-Level: $90,000 - $140,000 Senior-Level: $140,000 - $200,000 Final compensation will depend on qualifications, experience, level, and program requirements. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
Quick Position Facts! Location: Virginia Beach, Virginia at our G2 Ops office. Work Setting: In-Office with Hybrid or Flexible Schedule Available Based on Program Requirements Salary Range: $140,000 - $165,000 plus comprehensive benefits package. Years of Industry Experience: 10 + years of relevant experience. Security Clearance Requirement: Active DoD Secret clearance required at start, with the ability to obtain and maintain a DoD Top Secret clearance after starting. About the Role. As a Principal Software Engineer , you'll shape the technical direction of software capabilities supporting the AEGIS portfolio and Department of Defense missions. You'll serve as the senior software engineering lead across a portfolio of cybersecurity and digital engineering applications, with primary responsibility for the continued development and evolution of the SOFIA Cyber Analytics Platform and technical oversight of MONARCH, SMART, and related capabilities. This is a hands-on technical leadership role for an experienced software engineer who enjoys solving complex problems, setting technical direction, and mentoring other engineers. You'll work closely with software, cybersecurity, and systems engineers, program leadership, and customers to turn complex mission needs into secure, scalable, and maintainable software solutions. What You'll Do Lead the technical direction of SOFIA, MONARCH, SMART, and related software capabilities supporting the AEGIS portfolio. Own software architecture and guide the design, development, integration, deployment, and continued evolution of mission-focused applications. Provide hands-on technical leadership to the SOFIA development team, guiding technical planning, development priorities, design decisions, code quality, and engineering execution. Establish and promote software engineering standards and best practices across front-end and back-end development, APIs, databases, analytics, cloud integration, and plugin development. Translate user and mission needs into software architectures, technical requirements, development priorities, and executable solutions. Guide development of capabilities supporting cyber analytics, vulnerability and reachability analysis, mission impact assessment, data visualization, digital engineering, and decision support. Review software designs and source code, troubleshoot complex technical challenges, and mentor engineers through architecture reviews, code reviews, and technical planning. Advance modern development practices, including Agile, DevOps/DevSecOps, CI/CD, automated testing, configuration management, and secure software development. Evaluate emerging technologies, including AI, advanced analytics, cloud services, and modern software frameworks, and apply them where they provide meaningful value. Collaborate across engineering disciplines and with program leadership and customer stakeholders to support technical demonstrations, customer engagements, technical planning, proposals, and R&D efforts. What You'll Bring Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred. 10+ years of professional software engineering experience, including designing, developing, enhancing, and sustaining software applications. Experience serving as a Lead Software Engineer, Principal Software Engineer, Software Architect, Technical Lead, or similar technical leadership role. Strong programming foundation with professional development experience in one or more modern programming languages. Experience with modern web applications, distributed systems, APIs, databases, cloud technologies, and data-driven software solutions. Strong understanding of software architecture and secure software development, with the ability to review designs and source code, troubleshoot complex issues, and provide technical direction to other engineers. Experience applying Agile and DevOps/DevSecOps practices, including CI/CD and modern software delivery approaches. Strong communication skills with the ability to lead technical discussions across multidisciplinary engineering teams, program leadership, and customer stakeholders. Experience supporting Department of Defense or Federal Government software development programs preferred. Experience with cybersecurity analytics, digital engineering, MBSE, system-of-systems environments, AI/advanced analytics, or mission-focused engineering applications preferred. Certified Scrum Developer, Agile Developer, comparable software development certification, and/or CompTIA Security+ preferred. Additional Considerations This is a full-time position supporting mission-focused customer environments. Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $140,000 - $165,000, depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
09/25/2026
Full time
Quick Position Facts! Location: Virginia Beach, Virginia at our G2 Ops office. Work Setting: In-Office with Hybrid or Flexible Schedule Available Based on Program Requirements Salary Range: $140,000 - $165,000 plus comprehensive benefits package. Years of Industry Experience: 10 + years of relevant experience. Security Clearance Requirement: Active DoD Secret clearance required at start, with the ability to obtain and maintain a DoD Top Secret clearance after starting. About the Role. As a Principal Software Engineer , you'll shape the technical direction of software capabilities supporting the AEGIS portfolio and Department of Defense missions. You'll serve as the senior software engineering lead across a portfolio of cybersecurity and digital engineering applications, with primary responsibility for the continued development and evolution of the SOFIA Cyber Analytics Platform and technical oversight of MONARCH, SMART, and related capabilities. This is a hands-on technical leadership role for an experienced software engineer who enjoys solving complex problems, setting technical direction, and mentoring other engineers. You'll work closely with software, cybersecurity, and systems engineers, program leadership, and customers to turn complex mission needs into secure, scalable, and maintainable software solutions. What You'll Do Lead the technical direction of SOFIA, MONARCH, SMART, and related software capabilities supporting the AEGIS portfolio. Own software architecture and guide the design, development, integration, deployment, and continued evolution of mission-focused applications. Provide hands-on technical leadership to the SOFIA development team, guiding technical planning, development priorities, design decisions, code quality, and engineering execution. Establish and promote software engineering standards and best practices across front-end and back-end development, APIs, databases, analytics, cloud integration, and plugin development. Translate user and mission needs into software architectures, technical requirements, development priorities, and executable solutions. Guide development of capabilities supporting cyber analytics, vulnerability and reachability analysis, mission impact assessment, data visualization, digital engineering, and decision support. Review software designs and source code, troubleshoot complex technical challenges, and mentor engineers through architecture reviews, code reviews, and technical planning. Advance modern development practices, including Agile, DevOps/DevSecOps, CI/CD, automated testing, configuration management, and secure software development. Evaluate emerging technologies, including AI, advanced analytics, cloud services, and modern software frameworks, and apply them where they provide meaningful value. Collaborate across engineering disciplines and with program leadership and customer stakeholders to support technical demonstrations, customer engagements, technical planning, proposals, and R&D efforts. What You'll Bring Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred. 10+ years of professional software engineering experience, including designing, developing, enhancing, and sustaining software applications. Experience serving as a Lead Software Engineer, Principal Software Engineer, Software Architect, Technical Lead, or similar technical leadership role. Strong programming foundation with professional development experience in one or more modern programming languages. Experience with modern web applications, distributed systems, APIs, databases, cloud technologies, and data-driven software solutions. Strong understanding of software architecture and secure software development, with the ability to review designs and source code, troubleshoot complex issues, and provide technical direction to other engineers. Experience applying Agile and DevOps/DevSecOps practices, including CI/CD and modern software delivery approaches. Strong communication skills with the ability to lead technical discussions across multidisciplinary engineering teams, program leadership, and customer stakeholders. Experience supporting Department of Defense or Federal Government software development programs preferred. Experience with cybersecurity analytics, digital engineering, MBSE, system-of-systems environments, AI/advanced analytics, or mission-focused engineering applications preferred. Certified Scrum Developer, Agile Developer, comparable software development certification, and/or CompTIA Security+ preferred. Additional Considerations This is a full-time position supporting mission-focused customer environments. Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $140,000 - $165,000, depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
Quick Position Facts! Location: Arlington, Virginia at our G2 Ops office. Work Setting: In-Office with Hybrid or Flexible Schedule Available Based on Program Requirements Salary Range: $145,000 - $170,000 plus comprehensive benefits package. Years of Industry Experience: 10 + years of relevant experience. Security Clearance Requirement: Active DoD Secret clearance required at start, with the ability to obtain and maintain a DoD Top Secret clearance after starting. About the Role. As a Principal Software Engineer , you'll shape the technical direction of software capabilities supporting the AEGIS portfolio and Department of Defense missions. You'll serve as the senior software engineering lead across a portfolio of cybersecurity and digital engineering applications, with primary responsibility for the continued development and evolution of the SOFIA Cyber Analytics Platform and technical oversight of MONARCH, SMART, and related capabilities. This is a hands-on technical leadership role for an experienced software engineer who enjoys solving complex problems, setting technical direction, and mentoring other engineers. You'll work closely with software, cybersecurity, and systems engineers, program leadership, and customers to turn complex mission needs into secure, scalable, and maintainable software solutions. What You'll Do Lead the technical direction of SOFIA, MONARCH, SMART, and related software capabilities supporting the AEGIS portfolio. Own software architecture and guide the design, development, integration, deployment, and continued evolution of mission-focused applications. Provide hands-on technical leadership to the SOFIA development team, guiding technical planning, development priorities, design decisions, code quality, and engineering execution. Establish and promote software engineering standards and best practices across front-end and back-end development, APIs, databases, analytics, cloud integration, and plugin development. Translate user and mission needs into software architectures, technical requirements, development priorities, and executable solutions. Guide development of capabilities supporting cyber analytics, vulnerability and reachability analysis, mission impact assessment, data visualization, digital engineering, and decision support. Review software designs and source code, troubleshoot complex technical challenges, and mentor engineers through architecture reviews, code reviews, and technical planning. Advance modern development practices, including Agile, DevOps/DevSecOps, CI/CD, automated testing, configuration management, and secure software development. Evaluate emerging technologies, including AI, advanced analytics, cloud services, and modern software frameworks, and apply them where they provide meaningful value. Collaborate across engineering disciplines and with program leadership and customer stakeholders to support technical demonstrations, customer engagements, technical planning, proposals, and R&D efforts. What You'll Bring Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred. 10+ years of professional software engineering experience, including designing, developing, enhancing, and sustaining software applications. Experience serving as a Lead Software Engineer, Principal Software Engineer, Software Architect, Technical Lead, or similar technical leadership role. Strong programming foundation with professional development experience in one or more modern programming languages. Experience with modern web applications, distributed systems, APIs, databases, cloud technologies, and data-driven software solutions. Strong understanding of software architecture and secure software development, with the ability to review designs and source code, troubleshoot complex issues, and provide technical direction to other engineers. Experience applying Agile and DevOps/DevSecOps practices, including CI/CD and modern software delivery approaches. Strong communication skills with the ability to lead technical discussions across multidisciplinary engineering teams, program leadership, and customer stakeholders. Experience supporting Department of Defense or Federal Government software development programs preferred. Experience with cybersecurity analytics, digital engineering, MBSE, system-of-systems environments, AI/advanced analytics, or mission-focused engineering applications preferred. Certified Scrum Developer, Agile Developer, comparable software development certification, and/or CompTIA Security+ preferred. Additional Considerations This is a full-time position supporting mission-focused customer environments. Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $145,000 - $170,000, depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
09/25/2026
Full time
Quick Position Facts! Location: Arlington, Virginia at our G2 Ops office. Work Setting: In-Office with Hybrid or Flexible Schedule Available Based on Program Requirements Salary Range: $145,000 - $170,000 plus comprehensive benefits package. Years of Industry Experience: 10 + years of relevant experience. Security Clearance Requirement: Active DoD Secret clearance required at start, with the ability to obtain and maintain a DoD Top Secret clearance after starting. About the Role. As a Principal Software Engineer , you'll shape the technical direction of software capabilities supporting the AEGIS portfolio and Department of Defense missions. You'll serve as the senior software engineering lead across a portfolio of cybersecurity and digital engineering applications, with primary responsibility for the continued development and evolution of the SOFIA Cyber Analytics Platform and technical oversight of MONARCH, SMART, and related capabilities. This is a hands-on technical leadership role for an experienced software engineer who enjoys solving complex problems, setting technical direction, and mentoring other engineers. You'll work closely with software, cybersecurity, and systems engineers, program leadership, and customers to turn complex mission needs into secure, scalable, and maintainable software solutions. What You'll Do Lead the technical direction of SOFIA, MONARCH, SMART, and related software capabilities supporting the AEGIS portfolio. Own software architecture and guide the design, development, integration, deployment, and continued evolution of mission-focused applications. Provide hands-on technical leadership to the SOFIA development team, guiding technical planning, development priorities, design decisions, code quality, and engineering execution. Establish and promote software engineering standards and best practices across front-end and back-end development, APIs, databases, analytics, cloud integration, and plugin development. Translate user and mission needs into software architectures, technical requirements, development priorities, and executable solutions. Guide development of capabilities supporting cyber analytics, vulnerability and reachability analysis, mission impact assessment, data visualization, digital engineering, and decision support. Review software designs and source code, troubleshoot complex technical challenges, and mentor engineers through architecture reviews, code reviews, and technical planning. Advance modern development practices, including Agile, DevOps/DevSecOps, CI/CD, automated testing, configuration management, and secure software development. Evaluate emerging technologies, including AI, advanced analytics, cloud services, and modern software frameworks, and apply them where they provide meaningful value. Collaborate across engineering disciplines and with program leadership and customer stakeholders to support technical demonstrations, customer engagements, technical planning, proposals, and R&D efforts. What You'll Bring Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred. 10+ years of professional software engineering experience, including designing, developing, enhancing, and sustaining software applications. Experience serving as a Lead Software Engineer, Principal Software Engineer, Software Architect, Technical Lead, or similar technical leadership role. Strong programming foundation with professional development experience in one or more modern programming languages. Experience with modern web applications, distributed systems, APIs, databases, cloud technologies, and data-driven software solutions. Strong understanding of software architecture and secure software development, with the ability to review designs and source code, troubleshoot complex issues, and provide technical direction to other engineers. Experience applying Agile and DevOps/DevSecOps practices, including CI/CD and modern software delivery approaches. Strong communication skills with the ability to lead technical discussions across multidisciplinary engineering teams, program leadership, and customer stakeholders. Experience supporting Department of Defense or Federal Government software development programs preferred. Experience with cybersecurity analytics, digital engineering, MBSE, system-of-systems environments, AI/advanced analytics, or mission-focused engineering applications preferred. Certified Scrum Developer, Agile Developer, comparable software development certification, and/or CompTIA Security+ preferred. Additional Considerations This is a full-time position supporting mission-focused customer environments. Outside employment or activities must not create conflicts of interest with company or customer responsibilities Why G2 Ops? What makes someone choose one company over another? Compensation, benefits, meaningful work, flexibility, growth opportunities, culture? At G2 Ops, we believe you shouldn't have to choose. We offer competitive pay and benefits, but what truly sets us apart is our collaborative culture. Team members support mission-focused projects alongside highly skilled technical professionals across engineering, cybersecurity, operations, and program teams. We encourage continuous learning, cross-training, and professional growth while empowering team members to contribute ideas that improve how we work and deliver value to our customers. At G2 Ops, your contributions matter to our customers' missions and to the continued success of our team. Compensation & Benefits. The annual salary range for this position is $145,000 - $170,000, depending on qualifications and experience. G2 Ops offers a competitive compensation and benefits package designed to support our employees both personally and professionally, including: 100% company-paid insurance for medical, dental, and vision for eligible employees and family members 100% company-paid insurance for life, short-term disability (STD), and long-term disability (LTD) for eligible employees 401(k) plan with discretionary employer matching 10 paid holidays Paid time off (PTO) Educational assistance In addition, we provide professional development opportunities, performance recognition programs, and resources that support long-term career growth and work-life balance. AI at G2 Ops. At G2 Ops, we don't just talk about AI - we actively use it to improve how we work. Our teams are integrating AI into engineering, cybersecurity, operations, and decision-making workflows. We continue to invest in secure AI tooling aligned with government security requirements while developing practical, mission-focused applications across the company. Whether your role is technical or operational, you'll have opportunities to explore how AI can enhance efficiency, innovation, and mission impact. Work Environment. Because we support classified and mission-critical DoD programs, many roles require onsite collaboration at G2 Ops offices and/or customer locations. Depending on program requirements, telework and flexible scheduling options may be available. We've built a collaborative environment where team members can learn, contribute, and grow while supporting meaningful customer missions. Ready to Apply? If you're excited about solving meaningful problems, working alongside talented teammates, and contributing to important mission-focused work, we'd love to hear from you. We look forward to learning more about you! About G2 Ops, Inc. (G2 Ops) G2 Ops leverages over a decade of experience integrating Systems, Cybersecurity, and Software Engineering techniques to provide solutions to a growing list of Government and private customers. We combine cutting edge tools with innovative engineering practices, data analytics, and risk algorithms that enhance visibility into complex infrastructures, optimizing resiliency in system design and operations. G2 Ops is a woman-owned small business led by an executive staff known for providing innovative solutions to solve our nation's most complex engineering challenges. G2 Ops has been named to the Inc. 5000 list of America's fastest growing companies each of the last 9 years () and has locations in Arlington, VA, Virginia Beach, VA, and San Diego, CA. G2 Ops, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law. G2 Ops, Inc. participates in the E-Verify program. Employment is contingent upon verification of identity and authorization to work in the United States. Applicants have rights under Federal Employment Laws: E-Verify Participation and Right to Work Notices:
Vast is developing next-generation space stations and space infrastructure using an incremental, hardware-rich, and low-cost approach. Vast is rapidly developing its multi-module Haven Station to ensure a continuous human presence in space for America and its allies, enabling advanced microgravity research and manufacturing, and unlocking a new space economy for government, corporate, and private customers. Haven Demo's 2025 success made Vast the only operational commercial space station company to fly and operate its own spacecraft. Next, Haven-1 is expected to become the world's first commercial space station when it launches in 2027, followed by additional Haven modules. Additionally, the company recently announced Vast Satellite , a high-power satellite product line leveraging its space station components and the heritage of Haven Demo. Headquartered in Long Beach, California , and with more than 1,000 employees and over a billion dollars in private capital, Vast has built the facilities required to manufacture and operate America's next space station. The company plans to develop future habitats and systems for the Moon and Mars, dedicated space stations for government partners, and other crewed systems that will unlock the expanding long-term space economy. Vast is looking for a Staff Software Engineer, AI Tooling reporting to the Senior Manager, Software Engineering to support the development of the systems that will be required for the design and build of artificial-gravity human-rated space stations. This will be a full-time , exempt position located in our Long Beach location. Responsibilities: Define the technical direction and long-term roadmap for Vast's internal AI platforms and tooling Architect and lead the development of full-stack AI applications that serve diverse use-cases across the company Establish engineering best practices, design patterns, and quality standards for AI systems development Build and maintain production-grade integrations connecting AI models with internal tools, data sources, and workflows Design and implement pipelines for AI response enforcement, content safety, and output formatting Drive the evaluation and adoption of emerging AI frameworks, models, and patterns to continuously improve internal tooling Collaborate cross-functionally with IT Infrastructure, security, and business stakeholders to integrate AI tooling with existing systems Mentor and guide other engineers contributing to AI initiatives, fostering a culture of technical excellence Ensure the reliability, observability, and scalability of production AI systems Rapidly iterate on AI tooling as the technology landscape and company needs evolve Minimum Qualifications: Bachelor's degree in computer science, math, an engineering discipline, or equivalent work experience 8+ years of software development or relevant industry experience Proven track record of leading technical initiatives from concept through production deployment Strong full-stack development proficiency with experience across both frontend and backend systems Deep understanding of AI/ML concepts, large language model architectures, prompt engineering, and modern AI tooling patterns Demonstrated experience implementing and maintaining production applications at scale Experience establishing engineering best practices and providing technical mentorship Strong cross-functional communication skills with the ability to translate technical concepts for diverse audiences Preferred Skills & Experience: Experience with Python and TypeScript in production environments Experience with Svelte or SvelteKit Experience with OpenWebUI or similar open-source AI interface platforms Experience building or integrating MCP (Model Context Protocol) servers and gateways Experience with LLM API integration and AI pipeline development Experience with containerization and deployment (Docker, Kubernetes) Familiarity with RAG (Retrieval-Augmented Generation) patterns and vector databases Experience building or leading AI/ML platforms at scale Prior experience as a tech lead or in a similar engineering leadership role Experience working in fast-paced, ambiguous environments with evolving requirements Additional Requirements: Ability to travel up to 10% of the time Willingness to work evenings and/or weekends to support critical mission milestones Ability to lift up to 25 lbs unassisted Specific certifications, as appropriate Pay Range: Senior Software Engineer, AI Tooling: $159,900 - $226,900 Staff Software Engineer, AI Tooling: $188,600 - $267,700 Pay Range: California $159,900-$267,700 USD We calibrate level and compensation to the selected candidate's experience and demonstrated capability. We welcome applicants across a range of experience levels to apply. COMPENSATION AND BENEFITS Base salary will vary depending on job-related knowledge, education, skills, experience, business needs, and market demand. Salary is just one component of our comprehensive compensation package. Full-time employees also receive company equity, as well as access to a full suite of compelling benefits and perks, including: medical, dental, and vision coverage for employees and dependents, generous paid time off; up to 20+ days of vacation for exempt staff and up to 10+ days of vacation for non-exempt staff with the ability to cash-out unused vacation annually, paid parental leave, short and long-term disability insurance, life insurance, access to a 401(k) retirement plan, ClassPass credits, personalized mental healthcare through Spring Health, and other discounts and perks. We also take pride in offering exceptional food perks, with snacks, drip coffee & onsite barista, cold drinks, and dinner meals remaining free of charge, and lunch subsidized as part of Vast's ongoing commitment to providing high-quality meals for employees. U.S. EXPORT CONTROL COMPLIANCE STATUS The person hired will have access to information and items subject to U.S. export controls, and therefore, must either be a "U.S. person" as defined by 22 C.F.R. 120.62 or otherwise eligible for deemed export licensing. This status includes U.S. citizens, U.S. nationals, lawful permanent residents (green card holders), and asylees and refugees with such status granted, not pending. EQUAL OPPORTUNITY Vast is an Equal Opportunity Employer; employment with Vast is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.
09/25/2026
Full time
Vast is developing next-generation space stations and space infrastructure using an incremental, hardware-rich, and low-cost approach. Vast is rapidly developing its multi-module Haven Station to ensure a continuous human presence in space for America and its allies, enabling advanced microgravity research and manufacturing, and unlocking a new space economy for government, corporate, and private customers. Haven Demo's 2025 success made Vast the only operational commercial space station company to fly and operate its own spacecraft. Next, Haven-1 is expected to become the world's first commercial space station when it launches in 2027, followed by additional Haven modules. Additionally, the company recently announced Vast Satellite , a high-power satellite product line leveraging its space station components and the heritage of Haven Demo. Headquartered in Long Beach, California , and with more than 1,000 employees and over a billion dollars in private capital, Vast has built the facilities required to manufacture and operate America's next space station. The company plans to develop future habitats and systems for the Moon and Mars, dedicated space stations for government partners, and other crewed systems that will unlock the expanding long-term space economy. Vast is looking for a Staff Software Engineer, AI Tooling reporting to the Senior Manager, Software Engineering to support the development of the systems that will be required for the design and build of artificial-gravity human-rated space stations. This will be a full-time , exempt position located in our Long Beach location. Responsibilities: Define the technical direction and long-term roadmap for Vast's internal AI platforms and tooling Architect and lead the development of full-stack AI applications that serve diverse use-cases across the company Establish engineering best practices, design patterns, and quality standards for AI systems development Build and maintain production-grade integrations connecting AI models with internal tools, data sources, and workflows Design and implement pipelines for AI response enforcement, content safety, and output formatting Drive the evaluation and adoption of emerging AI frameworks, models, and patterns to continuously improve internal tooling Collaborate cross-functionally with IT Infrastructure, security, and business stakeholders to integrate AI tooling with existing systems Mentor and guide other engineers contributing to AI initiatives, fostering a culture of technical excellence Ensure the reliability, observability, and scalability of production AI systems Rapidly iterate on AI tooling as the technology landscape and company needs evolve Minimum Qualifications: Bachelor's degree in computer science, math, an engineering discipline, or equivalent work experience 8+ years of software development or relevant industry experience Proven track record of leading technical initiatives from concept through production deployment Strong full-stack development proficiency with experience across both frontend and backend systems Deep understanding of AI/ML concepts, large language model architectures, prompt engineering, and modern AI tooling patterns Demonstrated experience implementing and maintaining production applications at scale Experience establishing engineering best practices and providing technical mentorship Strong cross-functional communication skills with the ability to translate technical concepts for diverse audiences Preferred Skills & Experience: Experience with Python and TypeScript in production environments Experience with Svelte or SvelteKit Experience with OpenWebUI or similar open-source AI interface platforms Experience building or integrating MCP (Model Context Protocol) servers and gateways Experience with LLM API integration and AI pipeline development Experience with containerization and deployment (Docker, Kubernetes) Familiarity with RAG (Retrieval-Augmented Generation) patterns and vector databases Experience building or leading AI/ML platforms at scale Prior experience as a tech lead or in a similar engineering leadership role Experience working in fast-paced, ambiguous environments with evolving requirements Additional Requirements: Ability to travel up to 10% of the time Willingness to work evenings and/or weekends to support critical mission milestones Ability to lift up to 25 lbs unassisted Specific certifications, as appropriate Pay Range: Senior Software Engineer, AI Tooling: $159,900 - $226,900 Staff Software Engineer, AI Tooling: $188,600 - $267,700 Pay Range: California $159,900-$267,700 USD We calibrate level and compensation to the selected candidate's experience and demonstrated capability. We welcome applicants across a range of experience levels to apply. COMPENSATION AND BENEFITS Base salary will vary depending on job-related knowledge, education, skills, experience, business needs, and market demand. Salary is just one component of our comprehensive compensation package. Full-time employees also receive company equity, as well as access to a full suite of compelling benefits and perks, including: medical, dental, and vision coverage for employees and dependents, generous paid time off; up to 20+ days of vacation for exempt staff and up to 10+ days of vacation for non-exempt staff with the ability to cash-out unused vacation annually, paid parental leave, short and long-term disability insurance, life insurance, access to a 401(k) retirement plan, ClassPass credits, personalized mental healthcare through Spring Health, and other discounts and perks. We also take pride in offering exceptional food perks, with snacks, drip coffee & onsite barista, cold drinks, and dinner meals remaining free of charge, and lunch subsidized as part of Vast's ongoing commitment to providing high-quality meals for employees. U.S. EXPORT CONTROL COMPLIANCE STATUS The person hired will have access to information and items subject to U.S. export controls, and therefore, must either be a "U.S. person" as defined by 22 C.F.R. 120.62 or otherwise eligible for deemed export licensing. This status includes U.S. citizens, U.S. nationals, lawful permanent residents (green card holders), and asylees and refugees with such status granted, not pending. EQUAL OPPORTUNITY Vast is an Equal Opportunity Employer; employment with Vast is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.
Company Overview: Over the past 15 years, eTel has delivered essential solutions for the federal government by securing and managing data, providing scalable identity access, modernizing legacy systems, and building high-performance platforms. By integrating new technologies and ensuring reliable operations we help agencies stay prepared for future challenges As a premier technology solutions and services company to the US federal government, eTel possesses longstanding relationships across the federal civilian marketplace. Other customers include the broader Treasury Department, Commerce Department, and State Department. eTel offers integrated CMMI Level 3 processes, tools, and techniques with innovative, cost-efficient, and secure solutions to address complex challenges. eTel also holds ISO 9001:2015, ISO/IEC 27001:2013, and ISO/IEC 20000-1:2018 certifications, and offers dedicated subject matter experts (SMEs) and thought leaders that possess a deep understanding of customers' environments and challenges. Work Location and On-Site/Telework Requirements: Hybrid - NIH, Bethesda, MD. On site for workshops and pilots (typically 1-2 days/week during the first 120 days, then as scheduled). Citizenship: U.S. Citizenship required Clearance: All staff must obtain NIH suitability and a PIV credential and be fluent in English. Anyone doing risk or vulnerability testing needs a current T2 (BI) or higher investigation. Salary Range: $130,000-$140,000 yearly salary Overview : You will co-author the cloud, network, and identity design patterns behind NIH's Future State ZTA under the NIH Governance, Risk & Compliance (GRC) Zero Trust Architecture (ZTA) Support Services task order for the NIH Office of the Chief Information Officer (OCIO). You will work in the Architecture Pod on Tasks 2, 3, and 4. This is a design and advisory role: you will produce reference architectures, patterns, and control mappings for NIH teams to implement, not operate NIH production systems. Deep hands-on engineering experience is still essential. Responsibilities : Co-author the cloud, on-premises, and hybrid reference architectures (Subtask 2.2) with the Senior ZT Architect, as reusable design patterns with NIST SP 800-53 Rev 5 control mappings. Design microsegmentation and workload-identity patterns (NIST SP 800-207A, CISA and NSA Zero Trust guidance), software-defined perimeter for HPC clusters, and isolated VLANs with brokered remote access for laboratory instruments. Document how centrally provided network, endpoint, and logging services are inherited, as input to the Centrally Provided Services Matrix. Define technical policy anchors for the network and device pillars (segmentation policy in the controller, compliance policy in endpoint management). Map controls to telemetry so continuous-monitoring evidence is collected rather than written by hand. Support Task 3 network use cases: flow baselining to generate and validate microsegmentation policy, and lateral-movement anomaly detection. Support the Task 4 gap assessment for the network, device, and cloud pillars, and the evidence expectations in the ZTA Overlay. Validate patterns with IC engineering teams (CIT, HPC, and clinical platform owners). Tools & Technology Environment : AWS and Azure (including GovCloud), cloud-native IAM and CSPM; SSE/ZTNA (e.g., Zscaler, Palo Alto); microsegmentation platforms (e.g., Illumio); Palo Alto/Fortinet/Cisco firewalls; Microsoft Defender, CrowdStrike, Forescout; Splunk/Microsoft Sentinel; Terraform/IaC; Git. Required Qualifications : Bachelor's degree plus 7+ years of network and/or cloud security engineering. Hands-on experience with identity-aware access, microsegmentation, and cloud security patterns in enterprise environments. Experience with firewalls, SSE/ZTNA, and native AWS or Azure security services. Security+ or a higher security certification. Ability to obtain an NIH suitability determination and PIV credential; fluent in English. Preferred Qualifications : A cloud security certification (AWS Security Specialty, AZ-500, or CCSP); PCNSE or CCNP Security. Experience in FedRAMP-authorized or GovCloud environments. Experience with research networks, HPC (Linux/Slurm), or operational technology/IoT device segmentation. Infrastructure-as-code (Terraform, CloudFormation) and experience writing security patterns as code. Current National Institutes of Health (NIH) or U.S. Department of Health and Human Services (HHS) experience is highly preferred. Commitment to Diversity - eTelligent Group provides equal employment opportunities (EEO) to all applicants without regard to race, color, religion, gender, sexual orientation, gender identity, nations origin, age, disability, genetic information, marital status, amnesty, status as a covered veteran, and any other characteristic provided in accordance with applicable, federal, state and local laws.
09/25/2026
Full time
Company Overview: Over the past 15 years, eTel has delivered essential solutions for the federal government by securing and managing data, providing scalable identity access, modernizing legacy systems, and building high-performance platforms. By integrating new technologies and ensuring reliable operations we help agencies stay prepared for future challenges As a premier technology solutions and services company to the US federal government, eTel possesses longstanding relationships across the federal civilian marketplace. Other customers include the broader Treasury Department, Commerce Department, and State Department. eTel offers integrated CMMI Level 3 processes, tools, and techniques with innovative, cost-efficient, and secure solutions to address complex challenges. eTel also holds ISO 9001:2015, ISO/IEC 27001:2013, and ISO/IEC 20000-1:2018 certifications, and offers dedicated subject matter experts (SMEs) and thought leaders that possess a deep understanding of customers' environments and challenges. Work Location and On-Site/Telework Requirements: Hybrid - NIH, Bethesda, MD. On site for workshops and pilots (typically 1-2 days/week during the first 120 days, then as scheduled). Citizenship: U.S. Citizenship required Clearance: All staff must obtain NIH suitability and a PIV credential and be fluent in English. Anyone doing risk or vulnerability testing needs a current T2 (BI) or higher investigation. Salary Range: $130,000-$140,000 yearly salary Overview : You will co-author the cloud, network, and identity design patterns behind NIH's Future State ZTA under the NIH Governance, Risk & Compliance (GRC) Zero Trust Architecture (ZTA) Support Services task order for the NIH Office of the Chief Information Officer (OCIO). You will work in the Architecture Pod on Tasks 2, 3, and 4. This is a design and advisory role: you will produce reference architectures, patterns, and control mappings for NIH teams to implement, not operate NIH production systems. Deep hands-on engineering experience is still essential. Responsibilities : Co-author the cloud, on-premises, and hybrid reference architectures (Subtask 2.2) with the Senior ZT Architect, as reusable design patterns with NIST SP 800-53 Rev 5 control mappings. Design microsegmentation and workload-identity patterns (NIST SP 800-207A, CISA and NSA Zero Trust guidance), software-defined perimeter for HPC clusters, and isolated VLANs with brokered remote access for laboratory instruments. Document how centrally provided network, endpoint, and logging services are inherited, as input to the Centrally Provided Services Matrix. Define technical policy anchors for the network and device pillars (segmentation policy in the controller, compliance policy in endpoint management). Map controls to telemetry so continuous-monitoring evidence is collected rather than written by hand. Support Task 3 network use cases: flow baselining to generate and validate microsegmentation policy, and lateral-movement anomaly detection. Support the Task 4 gap assessment for the network, device, and cloud pillars, and the evidence expectations in the ZTA Overlay. Validate patterns with IC engineering teams (CIT, HPC, and clinical platform owners). Tools & Technology Environment : AWS and Azure (including GovCloud), cloud-native IAM and CSPM; SSE/ZTNA (e.g., Zscaler, Palo Alto); microsegmentation platforms (e.g., Illumio); Palo Alto/Fortinet/Cisco firewalls; Microsoft Defender, CrowdStrike, Forescout; Splunk/Microsoft Sentinel; Terraform/IaC; Git. Required Qualifications : Bachelor's degree plus 7+ years of network and/or cloud security engineering. Hands-on experience with identity-aware access, microsegmentation, and cloud security patterns in enterprise environments. Experience with firewalls, SSE/ZTNA, and native AWS or Azure security services. Security+ or a higher security certification. Ability to obtain an NIH suitability determination and PIV credential; fluent in English. Preferred Qualifications : A cloud security certification (AWS Security Specialty, AZ-500, or CCSP); PCNSE or CCNP Security. Experience in FedRAMP-authorized or GovCloud environments. Experience with research networks, HPC (Linux/Slurm), or operational technology/IoT device segmentation. Infrastructure-as-code (Terraform, CloudFormation) and experience writing security patterns as code. Current National Institutes of Health (NIH) or U.S. Department of Health and Human Services (HHS) experience is highly preferred. Commitment to Diversity - eTelligent Group provides equal employment opportunities (EEO) to all applicants without regard to race, color, religion, gender, sexual orientation, gender identity, nations origin, age, disability, genetic information, marital status, amnesty, status as a covered veteran, and any other characteristic provided in accordance with applicable, federal, state and local laws.
Join Axon and be a Force for Good. At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with candor and care, seeking out diverse perspectives from our customers, communities and each other. Life at Axon is fast-paced, challenging and meaningful. Here, you'll take ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter. Your Impact This role is primarily for the DEMS Ingestion team which is responsible for the upload of evidence files both from 3rd party clients as well as Axon devices such as Body Worn Cameras. We deal with large throughput upload requests ( Terabytes) and have an increasing need for a strong grasp of multi-cloud architecture (AWS & Azure). The team largely is a backend focused team leveraging mainly GoLang and Scala. Cloud technologies include blob storage, CDNs, stateful APIs, multi-part file upload, end-to-end encryption, SQL, event driven infrastructure, and caching. We are also making increasing use of AI to accelerate our development process through improved quality guardrails. An ideal candidate would be a strong backend engineer with a mind for system design across multi-cloud infrastructure. They should be a strong communicator of complex technical details and able to consider a wide range of use-case scenarios and risks in their designs. They should also possess well developed leadership soft skills as this role is expected to lead other engineers and act as a force multiplier for the team as whole. They should have a strong technical foundation that others on the team can rely on and be approachable and willing to help their team grow. What You'll Do Location: This role is based out of our Seattle Office and follows a hybrid schedule. We rely on in-person collaboration and ask that team members work onsite Tuesdays through Fridays, with the flexibility to work remotely on Mondays, unless there is an approved workplace accommodation. We believe that connection fuels innovation, and our in-office culture is designed to foster meaningful teamwork, mentorship, and shared success. Direct Reports: Manager of Software Engineering Design and develop scalable, secure, high-performance software in this mission-critical space Lead technical projects from concept to launch, ensuring solution meets business and technical requirements Collaborate across teams with Product, Design, and Engineering to create solutions that delight our customers Provide technical leadership and mentorship for engineers across the group Help to define a technical vision both within our team and the greater organization Partner with Engineering Managers, Directors, and Staff Engineers to define a strategic technical vision and direction for the group Lead engineering architecture design reviews and provide feedback to other engineers through channels like PR reviews Write and review design documents, library proposals, and technical vision documents. Define and manage roadmaps for the technical backlog Continuously evaluate and improve engineering processes across the group to enhance efficiency and effectiveness What You Bring Bachelor's Degree in Computer Science, Engineering, or related field 10+ years of professional software development experience Experience designing and delivering highly-available, scalable cloud-based systems Full-stack development experience in languages such as Go, Scala, Java, TypeScript, JavaScript, C# or similar Experience working with SQL or NoSQL data stores Experience with Realtime streaming event log or messaging technologies, such as Kafka or ActiveMQ (Bonus) Experience with a front end framework such as React (Bonus) Experience with mobile development (Bonus) Experience with an automated testing framework like Selenium, Puppeteer, or Playwright or automated API testing Benefits that Benefit You Competitive salary and 401k with employer match Discretionary paid time off Paid parental leave for all Medical, Dental, Vision plans Fitness Programs Emotional & Mental Wellness support Learning & Development programs Employee Resource Groups (ERGs) And yes, we have snacks in our offices Benefits listed herein may vary depending on the nature of your employment and the location where you work. Axon is a total compensation company, meaning compensation is made up of base pay, bonus, and stock awards. The actual base pay is dependent upon many factors, such as: level, function, training, transferable skills, work experience, business needs, geographic market, and often a combination of all these factors. Our benefits offer an array of options to help support you physically, financially and emotionally through the big milestones and in your everyday life. To see more details on our benefits offerings please visit Base Pay Range $148,500-$237,600 USD Don't meet every single requirement? That's ok. At Axon, we Aim Far. We think big with a long-term view because we want to reinvent the world to be a safer, better place. We are also committed to building diverse teams that reflect the communities we serve. Studies have shown that women and people of color are less likely to apply to jobs unless they check every box in the job description. If you're excited about this role and our mission to Protect Life but your experience doesn't align perfectly with every qualification listed here, we encourage you to apply anyways. You may be just the right candidate for this or other roles. Important Notes The above job description is not intended as, nor should it be construed as, exhaustive of all duties, responsibilities, skills, efforts, or working conditions associated with this job. The job description may change or be supplemented at any time in accordance with business needs and conditions. Some roles may also require legal eligibility to work in a firearms environment. We collect personal information from applicants to evaluate candidates for employment. You may request access, deletion, or exercise other CCPA rights at or via our Axon Privacy Web Form. For more information, please see the Your California Privacy Rights section of our Applicant and Candidate Privacy Notice. Axon's mission is to Protect Life and is committed to the well-being and safety of its employees as well as Axon's impact on the environment. All Axon employees must be aware of and committed to the appropriate environmental, health, and safety regulations, policies, and procedures. Axon employees are empowered to report safety concerns as they arise and activities potentially impacting the environment. We are an equal opportunity employer that promotes justice, advances equity, values diversity and fosters inclusion. We're committed to hiring the best talent - regardless of race, creed, color, ancestry, religion, sex (including pregnancy), national origin, sexual orientation, age, citizenship status, marital status, disability, gender identity, genetic information, veteran status, or any other characteristic protected by applicable laws, regulations and ordinances - and empowering all of our employees so they can do their best work. If you have a disability or special need that requires assistance or accommodation during the application or the recruiting process, please email . Please note that this email address is for accommodation purposes only. Axon will not respond to inquiries for other purposes. Phishing alert: Axon will never ask you to pay for any part of the hiring process, including training, equipment, or background checks. We do not make job offers via text message, WhatsApp, or instant messaging platforms without a formal interview process. All legitimate job openings are listed on our official careers page at If you receive a suspicious offer or outreach from an email address that is not , or if you are asked for sensitive personal information (bank details, Social Security Number) prematurely, please ignore the message and report it to .
09/25/2026
Full time
Join Axon and be a Force for Good. At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with candor and care, seeking out diverse perspectives from our customers, communities and each other. Life at Axon is fast-paced, challenging and meaningful. Here, you'll take ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter. Your Impact This role is primarily for the DEMS Ingestion team which is responsible for the upload of evidence files both from 3rd party clients as well as Axon devices such as Body Worn Cameras. We deal with large throughput upload requests ( Terabytes) and have an increasing need for a strong grasp of multi-cloud architecture (AWS & Azure). The team largely is a backend focused team leveraging mainly GoLang and Scala. Cloud technologies include blob storage, CDNs, stateful APIs, multi-part file upload, end-to-end encryption, SQL, event driven infrastructure, and caching. We are also making increasing use of AI to accelerate our development process through improved quality guardrails. An ideal candidate would be a strong backend engineer with a mind for system design across multi-cloud infrastructure. They should be a strong communicator of complex technical details and able to consider a wide range of use-case scenarios and risks in their designs. They should also possess well developed leadership soft skills as this role is expected to lead other engineers and act as a force multiplier for the team as whole. They should have a strong technical foundation that others on the team can rely on and be approachable and willing to help their team grow. What You'll Do Location: This role is based out of our Seattle Office and follows a hybrid schedule. We rely on in-person collaboration and ask that team members work onsite Tuesdays through Fridays, with the flexibility to work remotely on Mondays, unless there is an approved workplace accommodation. We believe that connection fuels innovation, and our in-office culture is designed to foster meaningful teamwork, mentorship, and shared success. Direct Reports: Manager of Software Engineering Design and develop scalable, secure, high-performance software in this mission-critical space Lead technical projects from concept to launch, ensuring solution meets business and technical requirements Collaborate across teams with Product, Design, and Engineering to create solutions that delight our customers Provide technical leadership and mentorship for engineers across the group Help to define a technical vision both within our team and the greater organization Partner with Engineering Managers, Directors, and Staff Engineers to define a strategic technical vision and direction for the group Lead engineering architecture design reviews and provide feedback to other engineers through channels like PR reviews Write and review design documents, library proposals, and technical vision documents. Define and manage roadmaps for the technical backlog Continuously evaluate and improve engineering processes across the group to enhance efficiency and effectiveness What You Bring Bachelor's Degree in Computer Science, Engineering, or related field 10+ years of professional software development experience Experience designing and delivering highly-available, scalable cloud-based systems Full-stack development experience in languages such as Go, Scala, Java, TypeScript, JavaScript, C# or similar Experience working with SQL or NoSQL data stores Experience with Realtime streaming event log or messaging technologies, such as Kafka or ActiveMQ (Bonus) Experience with a front end framework such as React (Bonus) Experience with mobile development (Bonus) Experience with an automated testing framework like Selenium, Puppeteer, or Playwright or automated API testing Benefits that Benefit You Competitive salary and 401k with employer match Discretionary paid time off Paid parental leave for all Medical, Dental, Vision plans Fitness Programs Emotional & Mental Wellness support Learning & Development programs Employee Resource Groups (ERGs) And yes, we have snacks in our offices Benefits listed herein may vary depending on the nature of your employment and the location where you work. Axon is a total compensation company, meaning compensation is made up of base pay, bonus, and stock awards. The actual base pay is dependent upon many factors, such as: level, function, training, transferable skills, work experience, business needs, geographic market, and often a combination of all these factors. Our benefits offer an array of options to help support you physically, financially and emotionally through the big milestones and in your everyday life. To see more details on our benefits offerings please visit Base Pay Range $148,500-$237,600 USD Don't meet every single requirement? That's ok. At Axon, we Aim Far. We think big with a long-term view because we want to reinvent the world to be a safer, better place. We are also committed to building diverse teams that reflect the communities we serve. Studies have shown that women and people of color are less likely to apply to jobs unless they check every box in the job description. If you're excited about this role and our mission to Protect Life but your experience doesn't align perfectly with every qualification listed here, we encourage you to apply anyways. You may be just the right candidate for this or other roles. Important Notes The above job description is not intended as, nor should it be construed as, exhaustive of all duties, responsibilities, skills, efforts, or working conditions associated with this job. The job description may change or be supplemented at any time in accordance with business needs and conditions. Some roles may also require legal eligibility to work in a firearms environment. We collect personal information from applicants to evaluate candidates for employment. You may request access, deletion, or exercise other CCPA rights at or via our Axon Privacy Web Form. For more information, please see the Your California Privacy Rights section of our Applicant and Candidate Privacy Notice. Axon's mission is to Protect Life and is committed to the well-being and safety of its employees as well as Axon's impact on the environment. All Axon employees must be aware of and committed to the appropriate environmental, health, and safety regulations, policies, and procedures. Axon employees are empowered to report safety concerns as they arise and activities potentially impacting the environment. We are an equal opportunity employer that promotes justice, advances equity, values diversity and fosters inclusion. We're committed to hiring the best talent - regardless of race, creed, color, ancestry, religion, sex (including pregnancy), national origin, sexual orientation, age, citizenship status, marital status, disability, gender identity, genetic information, veteran status, or any other characteristic protected by applicable laws, regulations and ordinances - and empowering all of our employees so they can do their best work. If you have a disability or special need that requires assistance or accommodation during the application or the recruiting process, please email . Please note that this email address is for accommodation purposes only. Axon will not respond to inquiries for other purposes. Phishing alert: Axon will never ask you to pay for any part of the hiring process, including training, equipment, or background checks. We do not make job offers via text message, WhatsApp, or instant messaging platforms without a formal interview process. All legitimate job openings are listed on our official careers page at If you receive a suspicious offer or outreach from an email address that is not , or if you are asked for sensitive personal information (bank details, Social Security Number) prematurely, please ignore the message and report it to .
Senior Systems Engineer III - CICS Systems Programmer Overview We are seeking a Senior Systems Engineer III - CICS Systems Programmer to support the integrity, performance, configuration, documentation, and reliability of server operating systems, hardware, storage, and/or application software within an assigned area of responsibility. This role is well suited for someone with strong mainframe, systems programming, troubleshooting, documentation, audit, disaster recovery, and performance management experience. Primary Responsibilities Maintain and monitor usage, performance, and resources for assigned systems, hardware, storage, and/or application software. Troubleshoot day-to-day technical issues related to assigned systems and platforms. Develop, configure, customize, document, and modify system configurations. Support patching and version maintenance activities. Schedule testing and maintain documentation related to audit requirements and disaster recovery planning. Maintain technical documentation for assigned software, hardware management tools, and appliances. Participate in audit and disaster recovery practices and activities. Participate in root-cause problem analysis across technology and business areas as needed. Perform capacity planning and performance tuning activities for assigned areas of responsibility. Assist in developing entry-level personnel through mentoring, technical training, and related support. Required Qualifications Seven years of experience in one or more of the following areas: system administration on Windows, Unix, or another distributed platform; systems programming on a mainframe platform; storage management within a large SAN environment; software development; internal audit; or information security. In lieu of seven years of experience, a bachelor's degree and five years of experience in one or more of the same areas is acceptable. Experience demonstrating strong logical reasoning, problem-solving, and analytical skills. Ability to translate technical knowledge in a way that supports another person's understanding. Experience working with team members, team leads, end users, business partners, and management during day-to-day operations and project-based activities. Ability to help ensure software and hardware functionality, performance, and availability. Willingness to develop and train entry-level staff on operational and systems management tasks related to the area of expertise. Experience managing day-to-day operational workload and assigned projects within established deadlines and timeframes. Organizational skills to work with vendors and notify management and team members of pending patches, upgrades, and functionality enhancements. Experience maintaining and monitoring usage, performance, and resources while troubleshooting day-to-day issues. Strong written and verbal communication skills. Experience documenting initial system configuration, change management, audit, disaster recovery, and business continuity information while keeping required system documentation up to date. Preferred Qualifications Bachelor's degree in IT or a related field. Advanced-level certification in mainframe, Db2 database, network, UNIX/Linux, or storage designations in a systems engineering track. Experience leading mid-scale IT infrastructure projects. Experience working in an IBM z/OS mainframe environment. Knowledge of CICS architecture, internal interfaces, and resource management. CICS systems programming experience. SMP/E experience. CICSPlex configuration and management experience. Understanding of common mainframe utilities and programming languages such as COBOL, JCL, SQL, and REXX. Performance and capacity management experience using BMC tools. CICS application programming experience. Familiarity with other core mainframe products such as Db2 and MQ. BMC tools support experience, including installation, configuration, and maintenance. Automation development experience using languages and tools such as REXX, Python, YAML, and Ansible. Ideal Candidate Experienced in supporting complex systems environments with a focus on reliability, performance, and availability. Comfortable troubleshooting operational issues and participating in root-cause analysis. Strong documentation habits across configuration, change management, audit, disaster recovery, and business continuity activities. Able to work with both technical and non-technical stakeholders across day-to-day operations and project work. Interested in mentoring and helping entry-level staff grow technically. Determining compensation for this role (and others) at Vaco/Highspring depends upon a wide array of factors including but not limited to the individual's skill sets, experience and training, licensure and certifications, office location and other geographic considerations, as well as other business and organizational needs. With that said, as required by local law in geographies that require salary range disclosure, Vaco/Highspring notes the salary range for the role is noted in this job posting. The individual may also be eligible for discretionary bonuses, and can participate in medical, dental, and vision benefits as well as the company's 401(k) retirement plan. Additional disclaimer: Unless otherwise noted in the job description, the position Vaco/Highspring is filing for is occupied. Please note, however, that Vaco/Highspring is regularly asked to provide talent to other organizations. By submitting to this position, you are agreeing to be included in our talent pool for future hiring for similarly qualified positions. Submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. Further assessment of candidates beyond this initial phase within Vaco/Highspring will be otherwise assessed by recruiters and hiring managers. Vaco/Highspring does not have knowledge of the tools used by its clients in making final hiring decisions and cannot opine on their use of AI products. EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law . click apply for full job details
09/24/2026
Full time
Senior Systems Engineer III - CICS Systems Programmer Overview We are seeking a Senior Systems Engineer III - CICS Systems Programmer to support the integrity, performance, configuration, documentation, and reliability of server operating systems, hardware, storage, and/or application software within an assigned area of responsibility. This role is well suited for someone with strong mainframe, systems programming, troubleshooting, documentation, audit, disaster recovery, and performance management experience. Primary Responsibilities Maintain and monitor usage, performance, and resources for assigned systems, hardware, storage, and/or application software. Troubleshoot day-to-day technical issues related to assigned systems and platforms. Develop, configure, customize, document, and modify system configurations. Support patching and version maintenance activities. Schedule testing and maintain documentation related to audit requirements and disaster recovery planning. Maintain technical documentation for assigned software, hardware management tools, and appliances. Participate in audit and disaster recovery practices and activities. Participate in root-cause problem analysis across technology and business areas as needed. Perform capacity planning and performance tuning activities for assigned areas of responsibility. Assist in developing entry-level personnel through mentoring, technical training, and related support. Required Qualifications Seven years of experience in one or more of the following areas: system administration on Windows, Unix, or another distributed platform; systems programming on a mainframe platform; storage management within a large SAN environment; software development; internal audit; or information security. In lieu of seven years of experience, a bachelor's degree and five years of experience in one or more of the same areas is acceptable. Experience demonstrating strong logical reasoning, problem-solving, and analytical skills. Ability to translate technical knowledge in a way that supports another person's understanding. Experience working with team members, team leads, end users, business partners, and management during day-to-day operations and project-based activities. Ability to help ensure software and hardware functionality, performance, and availability. Willingness to develop and train entry-level staff on operational and systems management tasks related to the area of expertise. Experience managing day-to-day operational workload and assigned projects within established deadlines and timeframes. Organizational skills to work with vendors and notify management and team members of pending patches, upgrades, and functionality enhancements. Experience maintaining and monitoring usage, performance, and resources while troubleshooting day-to-day issues. Strong written and verbal communication skills. Experience documenting initial system configuration, change management, audit, disaster recovery, and business continuity information while keeping required system documentation up to date. Preferred Qualifications Bachelor's degree in IT or a related field. Advanced-level certification in mainframe, Db2 database, network, UNIX/Linux, or storage designations in a systems engineering track. Experience leading mid-scale IT infrastructure projects. Experience working in an IBM z/OS mainframe environment. Knowledge of CICS architecture, internal interfaces, and resource management. CICS systems programming experience. SMP/E experience. CICSPlex configuration and management experience. Understanding of common mainframe utilities and programming languages such as COBOL, JCL, SQL, and REXX. Performance and capacity management experience using BMC tools. CICS application programming experience. Familiarity with other core mainframe products such as Db2 and MQ. BMC tools support experience, including installation, configuration, and maintenance. Automation development experience using languages and tools such as REXX, Python, YAML, and Ansible. Ideal Candidate Experienced in supporting complex systems environments with a focus on reliability, performance, and availability. Comfortable troubleshooting operational issues and participating in root-cause analysis. Strong documentation habits across configuration, change management, audit, disaster recovery, and business continuity activities. Able to work with both technical and non-technical stakeholders across day-to-day operations and project work. Interested in mentoring and helping entry-level staff grow technically. Determining compensation for this role (and others) at Vaco/Highspring depends upon a wide array of factors including but not limited to the individual's skill sets, experience and training, licensure and certifications, office location and other geographic considerations, as well as other business and organizational needs. With that said, as required by local law in geographies that require salary range disclosure, Vaco/Highspring notes the salary range for the role is noted in this job posting. The individual may also be eligible for discretionary bonuses, and can participate in medical, dental, and vision benefits as well as the company's 401(k) retirement plan. Additional disclaimer: Unless otherwise noted in the job description, the position Vaco/Highspring is filing for is occupied. Please note, however, that Vaco/Highspring is regularly asked to provide talent to other organizations. By submitting to this position, you are agreeing to be included in our talent pool for future hiring for similarly qualified positions. Submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. Further assessment of candidates beyond this initial phase within Vaco/Highspring will be otherwise assessed by recruiters and hiring managers. Vaco/Highspring does not have knowledge of the tools used by its clients in making final hiring decisions and cannot opine on their use of AI products. EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law . click apply for full job details
About Us Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. At Visa, you'll have the opportunity to create impact at scale - tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world. Join Visa and do work that matters - to you, to your community, and to the world. Progress starts with you. Job Description Shape the Future of Enterprise AI at Scale We are looking for a seasoned Software Architect (Sr. Consultant title at Visa) to join our Corporate Generative AI Technologies team. In this role, you will help architect, build, and scale enterprise-grade Generative AI and agentic applications. This is a senior, hands-on engineering role for someone who brings strong system design judgment, full-stack product engineering depth, and the ability to translate complex business workflows into scalable, reliable AI-enabled automation solutions. You will work on platforms and applications that use AI-native application patterns, including large language models, agentic workflows, retrieval-augmented generation, tool orchestration, API integrations, ETL pipelines, and data systems to transform business processes into intelligent, reliable, and scalable workflows. You will help solve hard engineering problems such as decomposing complex workflows into reusable agents, designing secure human-in-the-loop systems, and building production-grade GenAI applications that can be monitored, evaluated, governed, and continuously improved. As a Senior Consultant (Staff Software Architect), you will operate with a high degree of autonomy, ownership, and technical judgment while collaborating closely with your manager and the broader engineering team to align on technical direction. You will be responsible for building high-quality software while also influencing system design decisions, architectural direction, engineering standards, and best practices across the team. Key Responsibilities Design, build, and scale enterprise-grade GenAI and agentic applications, with a strong focus on maintainable, secure, scalable, reliable, and production-ready architecture. Own architecture and delivery of major GenAI subsystems; lead design reviews; mentor I4/I5 engineers; define reusable patterns and production standards. Apply strong system design judgment to build full-stack, production-grade applications with robust API design, workflow orchestration, secure data flows, observability, and operational readiness. Build modern frontend experiences using React and established frontend patterns, including component-based architecture, state management, reusable UI components, accessibility, performance optimization, and seamless integration with backend APIs and AI-enabled services. Design and implement scalable backend services using Python, Node.js, and/or Java, including secure APIs, asynchronous processing, background jobs, authentication, authorization, logging, error handling, and system resiliency. Work with databases like PostgreSQL, Redis, vector databases, and related technologies, including schema design, indexing strategies, query optimization, transaction management, caching patterns, migrations, and data access patterns. Implement backend capabilities for AI-enabled and agentic workflow automation, including intent routing, agent orchestration, tool execution, API integrations, data retrieval, multi-step execution, workflow state management, human approval flows, guardrails, auditability, and enterprise system integration. Develop AI-native capabilities using OpenAI, Anthropic, and related LLM APIs/SDKs, including prompt orchestration, tool/function calling, structured outputs, streaming responses, model routing, and evaluation patterns. Apply deep knowledge of modern LLM capabilities to make informed engineering decisions around model selection, context management, latency, cost, reliability, output quality, safety, and user experience. Design and implement retrieval-augmented generation solutions, including ingestion pipelines, ETL workflows, embeddings, vector database integration, retrieval strategies, relevance ranking, grounding, and retrieval quality evaluation. Build and deploy cloud-native applications using containers, DevOps practices, CI/CD pipelines, automated testing, monitoring, and operational automation. Work closely with engineering teammates and cross-functional partners to align with team priorities, translate ambiguous requirements into proof-of-concepts, then evolve them into production-quality solutions through shared ownership and hands-on collaboration. Implement observability and operational excellence for GenAI applications, including end-to-end tracing, workflow telemetry, model evaluation, and monitoring to ensure secure, reliable, and production-ready AI systems. Technical Skills: Languages & Frameworks: Python, FastAPI, LangGraph Cloud & Infrastructure: AWS, Azure, Docker, Kubernetes, ECS DevOps: Git, CI/CD pipelines Anthropic / OpenAI SDKs, MCP, A2A Databases & Storage: Pinecone, Redis, PostgreSQL Monitoring & Governance: Prometheus, Grafana, audit logging, access control Frontend: ReactJS, HTML, CSS, JavaScript Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager. Qualifications Basic Qualifications: 8+ years of relevant work experience with a Bachelor's Degree or at least 5 years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 2 years of work experience with a PhD, OR 11+ years of relevant work experience. At least 8 years of experience designing, building, and operating complex distributed software systems in production environments. Strong background in software architecture, distributed systems, API design, cloud-native platforms, and data-intensive applications. Experience with cloud platforms, containerized environments, CI/CD systems, and observability tooling. Hands-on experience with React and backend development using Python, Node.js, Java, or similar languages. Strong communication skills with the ability to translate complex technical concepts to business stakeholders Preferred Qualifications: 9 or more years of relevant work experience with a Bachelor Degree or 7 or more relevant years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 3 or more years of experience with a PhD Experience building AI-enabled, data-driven, workflow automation, search, conversational, or machine learning-powered applications is highly valued. Direct experience with Generative AI technologies, LLMs, RAG architectures, or agentic systems is preferred but not required for candidates with exceptional software architecture and distributed systems experience. Experience driving technical strategy and influencing architectural direction across organizations. Expertise in system design tradeoffs involving scalability, security, reliability, performance, cost, and developer productivity. Deep understanding of modern LLM ecosystems, agent frameworks, retrieval architectures, and enterprise AI deployment patterns. Familiarity with modern AI architectures including agentic workflows, memory systems, tool orchestration, MCP, and A2A frameworks Proven ability to lead cross-functional AI initiatives, including PoC development, stakeholder alignment, and enterprise rollout. Information for US Applicants For roles located in the US, the estimated salary range for this position is $162,500.00 to $ 260,400.00 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position may be eligible for bonus and equity.Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program. Work Hours Varies upon the needs of the department. Travel Requirements This position requires travel 5-10% of the time. Mental/Physical Requirements This position will be performed in an office setting. The position will require the incumbent to sit and stand at a desk, communicate in person and by telephone, frequently operate standard office equipment, such as telephones and computers. Visa is an EEO Employer Qualified applicants will receive consideration for employment without regard to race, color religion, sex, national origin, sexual orientation, gender identity, disability or protect veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with the EEOC guidelines and applicable local law.
09/24/2026
Full time
About Us Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. At Visa, you'll have the opportunity to create impact at scale - tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world. Join Visa and do work that matters - to you, to your community, and to the world. Progress starts with you. Job Description Shape the Future of Enterprise AI at Scale We are looking for a seasoned Software Architect (Sr. Consultant title at Visa) to join our Corporate Generative AI Technologies team. In this role, you will help architect, build, and scale enterprise-grade Generative AI and agentic applications. This is a senior, hands-on engineering role for someone who brings strong system design judgment, full-stack product engineering depth, and the ability to translate complex business workflows into scalable, reliable AI-enabled automation solutions. You will work on platforms and applications that use AI-native application patterns, including large language models, agentic workflows, retrieval-augmented generation, tool orchestration, API integrations, ETL pipelines, and data systems to transform business processes into intelligent, reliable, and scalable workflows. You will help solve hard engineering problems such as decomposing complex workflows into reusable agents, designing secure human-in-the-loop systems, and building production-grade GenAI applications that can be monitored, evaluated, governed, and continuously improved. As a Senior Consultant (Staff Software Architect), you will operate with a high degree of autonomy, ownership, and technical judgment while collaborating closely with your manager and the broader engineering team to align on technical direction. You will be responsible for building high-quality software while also influencing system design decisions, architectural direction, engineering standards, and best practices across the team. Key Responsibilities Design, build, and scale enterprise-grade GenAI and agentic applications, with a strong focus on maintainable, secure, scalable, reliable, and production-ready architecture. Own architecture and delivery of major GenAI subsystems; lead design reviews; mentor I4/I5 engineers; define reusable patterns and production standards. Apply strong system design judgment to build full-stack, production-grade applications with robust API design, workflow orchestration, secure data flows, observability, and operational readiness. Build modern frontend experiences using React and established frontend patterns, including component-based architecture, state management, reusable UI components, accessibility, performance optimization, and seamless integration with backend APIs and AI-enabled services. Design and implement scalable backend services using Python, Node.js, and/or Java, including secure APIs, asynchronous processing, background jobs, authentication, authorization, logging, error handling, and system resiliency. Work with databases like PostgreSQL, Redis, vector databases, and related technologies, including schema design, indexing strategies, query optimization, transaction management, caching patterns, migrations, and data access patterns. Implement backend capabilities for AI-enabled and agentic workflow automation, including intent routing, agent orchestration, tool execution, API integrations, data retrieval, multi-step execution, workflow state management, human approval flows, guardrails, auditability, and enterprise system integration. Develop AI-native capabilities using OpenAI, Anthropic, and related LLM APIs/SDKs, including prompt orchestration, tool/function calling, structured outputs, streaming responses, model routing, and evaluation patterns. Apply deep knowledge of modern LLM capabilities to make informed engineering decisions around model selection, context management, latency, cost, reliability, output quality, safety, and user experience. Design and implement retrieval-augmented generation solutions, including ingestion pipelines, ETL workflows, embeddings, vector database integration, retrieval strategies, relevance ranking, grounding, and retrieval quality evaluation. Build and deploy cloud-native applications using containers, DevOps practices, CI/CD pipelines, automated testing, monitoring, and operational automation. Work closely with engineering teammates and cross-functional partners to align with team priorities, translate ambiguous requirements into proof-of-concepts, then evolve them into production-quality solutions through shared ownership and hands-on collaboration. Implement observability and operational excellence for GenAI applications, including end-to-end tracing, workflow telemetry, model evaluation, and monitoring to ensure secure, reliable, and production-ready AI systems. Technical Skills: Languages & Frameworks: Python, FastAPI, LangGraph Cloud & Infrastructure: AWS, Azure, Docker, Kubernetes, ECS DevOps: Git, CI/CD pipelines Anthropic / OpenAI SDKs, MCP, A2A Databases & Storage: Pinecone, Redis, PostgreSQL Monitoring & Governance: Prometheus, Grafana, audit logging, access control Frontend: ReactJS, HTML, CSS, JavaScript Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager. Qualifications Basic Qualifications: 8+ years of relevant work experience with a Bachelor's Degree or at least 5 years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 2 years of work experience with a PhD, OR 11+ years of relevant work experience. At least 8 years of experience designing, building, and operating complex distributed software systems in production environments. Strong background in software architecture, distributed systems, API design, cloud-native platforms, and data-intensive applications. Experience with cloud platforms, containerized environments, CI/CD systems, and observability tooling. Hands-on experience with React and backend development using Python, Node.js, Java, or similar languages. Strong communication skills with the ability to translate complex technical concepts to business stakeholders Preferred Qualifications: 9 or more years of relevant work experience with a Bachelor Degree or 7 or more relevant years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 3 or more years of experience with a PhD Experience building AI-enabled, data-driven, workflow automation, search, conversational, or machine learning-powered applications is highly valued. Direct experience with Generative AI technologies, LLMs, RAG architectures, or agentic systems is preferred but not required for candidates with exceptional software architecture and distributed systems experience. Experience driving technical strategy and influencing architectural direction across organizations. Expertise in system design tradeoffs involving scalability, security, reliability, performance, cost, and developer productivity. Deep understanding of modern LLM ecosystems, agent frameworks, retrieval architectures, and enterprise AI deployment patterns. Familiarity with modern AI architectures including agentic workflows, memory systems, tool orchestration, MCP, and A2A frameworks Proven ability to lead cross-functional AI initiatives, including PoC development, stakeholder alignment, and enterprise rollout. Information for US Applicants For roles located in the US, the estimated salary range for this position is $162,500.00 to $ 260,400.00 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position may be eligible for bonus and equity.Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program. Work Hours Varies upon the needs of the department. Travel Requirements This position requires travel 5-10% of the time. Mental/Physical Requirements This position will be performed in an office setting. The position will require the incumbent to sit and stand at a desk, communicate in person and by telephone, frequently operate standard office equipment, such as telephones and computers. Visa is an EEO Employer Qualified applicants will receive consideration for employment without regard to race, color religion, sex, national origin, sexual orientation, gender identity, disability or protect veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with the EEOC guidelines and applicable local law.
At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology. We are seeking a highly skilled and experienced AI Systems Engineer to join our team. This is a hands-on, senior individual contributor role that will be pivotal in leading the development, operations, and support of our entire AI infrastructure. You will be responsible for the entire lifecycle of our AI systems, from architecting and building high-performance GPU clusters to deploying and optimizing our most advanced AI models and agentic services. Responsibilities AI Infrastructure Architecture & Strategy: Lead the design and implementation of our next-generation AI infrastructure to support our Agentic AI initiatives. You will define the technical strategy for our on-premise GPU clusters, storage solutions, and networking to ensure optimal performance, scalability, and reliability for all our AI workloads. Cloud AI Service Integration: Support and secure the use of public cloud AI services, including Azure OpenAI services and Google Cloud Platform (GCP) services like Gemini. This includes managing secure access, monitoring usage, and tracking billing to ensure cost-effectiveness. You will also have hands-on experience supporting compute, GPUs, and AI services on both GCP and Azure. Hands-on GPU Cluster Management: Take a leadership role in the configuration, installation, and optimization of GPU server clusters. This includes advanced troubleshooting of hardware and software, performance tuning, and implementing best practices for cluster utilization and resource management. You will be an expert in administering job schedulers like LSF in a production environment, including integration with Docker for containerized job submission. Full-Stack AI Tech Stack Development & Operations: Architect and deploy a robust and scalable AI tech stack. You will be responsible for the end-to-end operational lifecycle, including setting up and managing deep learning frameworks (PyTorch, TensorFlow), containerization with Docker and Kubernetes, and implementing CI/CD pipelines for AI model development. Advanced LLM Deployment & Optimization: Lead the deployment, serving, and optimization of Large Language Models (LLMs). You will be an expert in techniques such as model quantization, distillation, and using high-performance serving frameworks (e.g., vLLM, TGI, TensorRT-LLM) to maximize inference throughput and minimize latency. Agentic AI Workflow & Service Engineering: Architect and build production-grade Agentic AI workflows and services. You will be responsible for the technical design and implementation of systems that integrate LLMs with external tools, APIs, and databases, and will mentor other engineers on building robust and scalable AI agent applications. Automation & Monitoring: Develop and maintain automation scripts using languages like Python, Bash, or Perl to streamline system maintenance, deployment, and reporting. Implement and manage monitoring solutions for system health, job statuses, GPU utilization, and container performance to proactively identify and resolve issues. AI Systems Support & Mentorship: Act as the final escalation point for the most complex technical issues related to our AI infrastructure. You will also serve as a technical leader and mentor to other engineers, providing guidance on best practices in AI systems engineering, performance tuning, and operational excellence. Security and Compliance: Develop and implement security best practices for our AI systems and data, ensuring compliance with relevant regulations and protecting our intellectual property. Required Skills and Qualifications 10+ years of experience in a senior technical role, with at least 5 years focused on building and operating high-performance computing or AI infrastructure. Proven track record as a Principal or Senior Staff Engineer. Expert-level knowledge of NVIDIA GPU architecture and technologies like CUDA and cuDNN. Extensive experience with multi-GPU and multi-node training and inference. Proven experience with public cloud AI services, specifically managing access, usage, and billing for Azure OpenAI and Google Cloud Platform (GCP) services. Extensive hands-on experience with Docker: image management, container orchestration, and troubleshooting. Proficiency in scripting languages such as Python, Bash, or Perl. Deep expertise in Linux system administration (RHEL preferred), including networking, storage, and performance tuning. Familiarity with user authentication and integration using systems like LDAP or Active Directory. Strong problem-solving and communication skills with the ability to work in a multi-platform, cross-functional, and geographically distributed team. Preferred/Bonus Skills Understanding of AI job profiling and tuning (memory, GPU, I/O). Experience administering LSF clusters in a production or research environment. Familiarity with other job schedulers like Slurm is a plus. Experience with LSF Docker integration and job submission using container images. Experience with macOS/AppleSilicon system admin tasks and troubleshooting. The annual salary range for California is $136,500 to $253,500. You may also be eligible to receive incentive compensation: bonus, equity, and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the salary range is a guideline and compensation may vary based on factors such as qualifications, skill level, competencies and work location. Our benefits programs include: paid vacation and paid holidays, 401(k) plan with employer match, employee stock purchase plan, a variety of medical, dental and vision plan options, and more. We're doing work that matters. Help us solve what others can't.
09/24/2026
Full time
At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology. We are seeking a highly skilled and experienced AI Systems Engineer to join our team. This is a hands-on, senior individual contributor role that will be pivotal in leading the development, operations, and support of our entire AI infrastructure. You will be responsible for the entire lifecycle of our AI systems, from architecting and building high-performance GPU clusters to deploying and optimizing our most advanced AI models and agentic services. Responsibilities AI Infrastructure Architecture & Strategy: Lead the design and implementation of our next-generation AI infrastructure to support our Agentic AI initiatives. You will define the technical strategy for our on-premise GPU clusters, storage solutions, and networking to ensure optimal performance, scalability, and reliability for all our AI workloads. Cloud AI Service Integration: Support and secure the use of public cloud AI services, including Azure OpenAI services and Google Cloud Platform (GCP) services like Gemini. This includes managing secure access, monitoring usage, and tracking billing to ensure cost-effectiveness. You will also have hands-on experience supporting compute, GPUs, and AI services on both GCP and Azure. Hands-on GPU Cluster Management: Take a leadership role in the configuration, installation, and optimization of GPU server clusters. This includes advanced troubleshooting of hardware and software, performance tuning, and implementing best practices for cluster utilization and resource management. You will be an expert in administering job schedulers like LSF in a production environment, including integration with Docker for containerized job submission. Full-Stack AI Tech Stack Development & Operations: Architect and deploy a robust and scalable AI tech stack. You will be responsible for the end-to-end operational lifecycle, including setting up and managing deep learning frameworks (PyTorch, TensorFlow), containerization with Docker and Kubernetes, and implementing CI/CD pipelines for AI model development. Advanced LLM Deployment & Optimization: Lead the deployment, serving, and optimization of Large Language Models (LLMs). You will be an expert in techniques such as model quantization, distillation, and using high-performance serving frameworks (e.g., vLLM, TGI, TensorRT-LLM) to maximize inference throughput and minimize latency. Agentic AI Workflow & Service Engineering: Architect and build production-grade Agentic AI workflows and services. You will be responsible for the technical design and implementation of systems that integrate LLMs with external tools, APIs, and databases, and will mentor other engineers on building robust and scalable AI agent applications. Automation & Monitoring: Develop and maintain automation scripts using languages like Python, Bash, or Perl to streamline system maintenance, deployment, and reporting. Implement and manage monitoring solutions for system health, job statuses, GPU utilization, and container performance to proactively identify and resolve issues. AI Systems Support & Mentorship: Act as the final escalation point for the most complex technical issues related to our AI infrastructure. You will also serve as a technical leader and mentor to other engineers, providing guidance on best practices in AI systems engineering, performance tuning, and operational excellence. Security and Compliance: Develop and implement security best practices for our AI systems and data, ensuring compliance with relevant regulations and protecting our intellectual property. Required Skills and Qualifications 10+ years of experience in a senior technical role, with at least 5 years focused on building and operating high-performance computing or AI infrastructure. Proven track record as a Principal or Senior Staff Engineer. Expert-level knowledge of NVIDIA GPU architecture and technologies like CUDA and cuDNN. Extensive experience with multi-GPU and multi-node training and inference. Proven experience with public cloud AI services, specifically managing access, usage, and billing for Azure OpenAI and Google Cloud Platform (GCP) services. Extensive hands-on experience with Docker: image management, container orchestration, and troubleshooting. Proficiency in scripting languages such as Python, Bash, or Perl. Deep expertise in Linux system administration (RHEL preferred), including networking, storage, and performance tuning. Familiarity with user authentication and integration using systems like LDAP or Active Directory. Strong problem-solving and communication skills with the ability to work in a multi-platform, cross-functional, and geographically distributed team. Preferred/Bonus Skills Understanding of AI job profiling and tuning (memory, GPU, I/O). Experience administering LSF clusters in a production or research environment. Familiarity with other job schedulers like Slurm is a plus. Experience with LSF Docker integration and job submission using container images. Experience with macOS/AppleSilicon system admin tasks and troubleshooting. The annual salary range for California is $136,500 to $253,500. You may also be eligible to receive incentive compensation: bonus, equity, and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the salary range is a guideline and compensation may vary based on factors such as qualifications, skill level, competencies and work location. Our benefits programs include: paid vacation and paid holidays, 401(k) plan with employer match, employee stock purchase plan, a variety of medical, dental and vision plan options, and more. We're doing work that matters. Help us solve what others can't.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.