it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

95 jobs found

Email me jobs like this
Refine Search
Current Search
sr product security engineer
Senior Cloud Engineer
Gridware San Francisco, California
Job Description Job Description About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety. Gridware's advanced Active Grid Response platform uses high-precision sensors to detect potential issues early, enabling proactive maintenance and fault mitigation. This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently. The company is backed by climate-tech and Silicon Valley investors. For more information, please visit . Role Description We're scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires. The platform you'll build and operate ingests millions of events per day from devices in the field, powers customer-facing dashboards and alerting, and supports the data science work that turns raw signals into grid intelligence. You will own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end, partnering with our Cloud Security team to keep the platform safe and compliant, and with backend, firmware, and data teams to keep them shipping fast. As an early member of the DevOps team, you'll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come. Responsibilities Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS. Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices. Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability. Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines. Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments. Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations. Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience. Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements. Required Skills 5+ years of experience in DevOps, SRE, or Platform Engineering operating production AWS environments Deep expertise with Kubernetes (EKS preferred), GitOps workflows (Argo CD/Flux), and Infrastructure as Code (Terraform) Strong experience building and maintaining CI/CD pipelines, ideally with GitHub Actions Hands-on experience operating distributed systems and cloud-native platforms (e.g., Kafka/MSK) Solid understanding of networking, DNS, TLS, identity/access management, and cloud security best practices Experience with observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar Strong Linux, scripting, and troubleshooting skills with the ability to debug complex production issues end-to-end Bonus Skills Experience operating Apollo Router / GraphQL federation gateways in production. Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production. Familiarity with Databricks or ML Ops pipelines for data and model deployment. Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks. Experience with Tailscale or other zero-trust networking tools. Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity. Experience in high-growth startup environments where you must wear many hats. This describes the ideal candidate; many of us have picked up this expertise along the way. Even if you meet only part of this list, we encourage you to apply! Gridware Technologies Inc. is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to any characteristic protected by applicable federal, state, or local law. Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees. Commuter allowance Company-paid training
08/05/2026
Full time
Job Description Job Description About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety. Gridware's advanced Active Grid Response platform uses high-precision sensors to detect potential issues early, enabling proactive maintenance and fault mitigation. This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently. The company is backed by climate-tech and Silicon Valley investors. For more information, please visit . Role Description We're scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires. The platform you'll build and operate ingests millions of events per day from devices in the field, powers customer-facing dashboards and alerting, and supports the data science work that turns raw signals into grid intelligence. You will own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end, partnering with our Cloud Security team to keep the platform safe and compliant, and with backend, firmware, and data teams to keep them shipping fast. As an early member of the DevOps team, you'll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come. Responsibilities Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS. Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices. Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability. Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines. Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments. Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations. Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience. Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements. Required Skills 5+ years of experience in DevOps, SRE, or Platform Engineering operating production AWS environments Deep expertise with Kubernetes (EKS preferred), GitOps workflows (Argo CD/Flux), and Infrastructure as Code (Terraform) Strong experience building and maintaining CI/CD pipelines, ideally with GitHub Actions Hands-on experience operating distributed systems and cloud-native platforms (e.g., Kafka/MSK) Solid understanding of networking, DNS, TLS, identity/access management, and cloud security best practices Experience with observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar Strong Linux, scripting, and troubleshooting skills with the ability to debug complex production issues end-to-end Bonus Skills Experience operating Apollo Router / GraphQL federation gateways in production. Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production. Familiarity with Databricks or ML Ops pipelines for data and model deployment. Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks. Experience with Tailscale or other zero-trust networking tools. Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity. Experience in high-growth startup environments where you must wear many hats. This describes the ideal candidate; many of us have picked up this expertise along the way. Even if you meet only part of this list, we encourage you to apply! Gridware Technologies Inc. is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to any characteristic protected by applicable federal, state, or local law. Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees. Commuter allowance Company-paid training
Sr. Finance Reporting Engineer
Align Technology San Jose, California
Job Description Job Description Description This position is ideal for senior-level professionals to join the Finance Systems & Reporting team as a Sr. Finance Reporting Engineer. The Finance Engineer will be reporting to the Director, Global Finance Systems & Reporting based in San Jose, California (US) and would strategically engage with key members of finance function defining and delivering on the enterprise roadmap projects as well as own and help resolve day to day tactical problems and challenges. The Sr. Finance Reporting Engineer will be responsible for designing, developing, and maintaining enterprise financial reporting and analytics solutions. This role focuses on delivering accurate, scalable, and automated reporting capabilities by leveraging SAP technologies, cloud-based data platforms, and business intelligence tools to support strategic decision-making across the organization. Role expectations Lead the financial reporting life cycle end-to-end that includes conducting requirement gathering sessions with business, design the reporting solution, development, deployment, maintenance and training of the reporting solution to broader regional audience. Oversee the design and development for the planning system, OneStream, and be prepared to use other planning and reporting tools used at Align, including ERP, Consolidation Planning, and Business Intelligence software packages. Support financial close activities through data reconciliation, validation, and coordination of data loads for month-end, quarter-end, and year-end reporting. Responsible for designing, developing, and maintaining scalable data pipelines and integration workflows to ensure accurate and efficient data processing across systems. Monitor data flows, task chains, and job performance; implementing data quality controls and validation processes; troubleshooting and resolving data load issues; and optimizing data transformation and loading performance. Develop and troubleshoot a variety of Onestream components that include transformation rules, security models, custom dashboards, reports, business rules, and member formulas, with a focus on enhancing system functionality and efficiency. Lead the collaboration and development efforts for various financial Power BI dashboards required for finance, that will make analytics easier and reporting simplified for business at the same time maintaining data accuracy standards as required for financial reporting. Acting as a liaison between business and IT groups. Work with IT to help them understand business requirements and translate it in technical terminologies for them. Help business understand the technical solutions deployed and train them on how they can effectively use it for their needs. Identifying the gaps with existing solutions in place, finding out solutions to resolve them and work with IT/business to get them implemented. Lead the integration testing and user acceptance testing with business and IT collaboratively. Key Personnel responsible for financial close reporting of critical reports Key person responsible for ensuring financial data integrity and completeness for all finance data that will be needed for financial close reporting. Comply with data security and access control standards, maintains thorough process documentation, and provides production support, including participation in on-call rotations with the offshore team when required. Assist with special projects and ad-hoc requests, as necessary. What we're looking for Requires Bachelor's degree in CS, Information Technology, or a related field 12+ years of experience in Data loading and Monitoring systems to support the Finance organization. Strong understanding of financial reporting, FP&A, close processes, and management reporting. Hands-on experience with SAP Datasphere for data modeling, data warehousing, integration, and governance across SAP and non-SAP landscapes. Strong understanding of SAP RTR business processes, hands on experience working with FI-GL, COPA, AR/AP, Should be a self-starter who is able to work with minimal direction and exercises considerable latitude in determining objectives and approaches to assignments. Hands-on experience working with Onestream planning system or similar planning platforms such as Anaplan, Planful. Hands-on experience with SQL and object-oriented (VB.Net, C#) coding experience is preferred. The candidate will serve as a liaison between the finance user group, corporate report development team, and IT. Should be a team player and possess good interpersonal and communication skills, reflecting an ability to be patient and outgoing with people. Should be highly motivated, result focused, and act with a high sense of urgency. Should possess excellent planning and prioritization skills with the ability to multitask and maintain by adapting to change. Complementary skills Advanced Microsoft Outlook, Word, Excel and PowerPoint skills. Must have the ability to independently create spreadsheets and perform quantitative analysis. Prior experience working with Azure Datalake, S/4 HANA, SAP-ECC or SAP BW is a plus. Pay Transparency If provided, base salary or wage rate ranges are the range in which Align reasonably expects to set a candidate's pay for the posted position. Actual placement depends on the individual skills and experience level of a candidate plus the total compensation and equity across team members. For other locations outside of the primary location, the base salary range will be adjusted geographically. For Field Sales roles, the salary listed is the base pay only and does not include the applicable incentive compensation plan. A cost of living adjustment may be added to base pay for higher cost areas in the U.S. Our internship hourly rates are a standard pay determined based on the position and your location, year in school, degree, and experience. General Description of All Benefits We are pleased to provide a general description of the benefits Align offers to full-time employees in this position. Family Benefits. Align offers employees and their eligible dependents medical (with a Health Savings Account option for some plan offerings), dental, and vision in accordance with those plans. Align also offers to employees: Discounts on Invisalign and Vivera to employees and their eligible dependents after 90 days of employment Back-up Child/Elder Care and access to a caregiving concierge Family Forming Benefits - Available to Employees, and their spouse or domestic partner, covered under one of Align's health plans Breast Milk Delivery and Lactation Support Services Employee Assistance Program Hinge Health Virtual Physical Therapy - Available to all employees and eligible dependents (age 18+) enrolled in an Align medical Plan Employee benefits. Align offers its employees: Short-term and long-term disability insurance in accordance with those plans. Basic Life Insurance and Accidental Death and Dismemberment. Voluntary Supplemental Life Insurance for Employee, Spouse/Domestic Partner, and Child(ren) are available for purchase in accordance with those plans. Flexible Spending Accounts- Employees may be eligible to participate in a health care account (including a limited health FSA if enrolled in a HDHP), dependent care account, and a pre-tax commuter benefit plan. 401k plan (with a discretionary Company match of 50% up to 6% of eligible earnings up to a maximum match of 3%.). Employer match vests after two years - 25% year one and 100% at year two. Align offers traditional, Roth, and after-tax options. Employee Stock Purchase Program (Employees must work 20 hours or more and be employed on purchase date to be eligible). Paid vacation of up to 17 days during the first full year of employment (currently accrued at the rate of 5.24 hours each pay-period), which carries over to a maximum cap of 30 days. Annual paid vacation time accrual increases based on tenure. Both exempt and non-exempt employees who work 32 hours or more per week receive prorated vacation accrual based on their regularly scheduled work hours and tenure. Sick time is accrued throughout the year at the rate of one hour for every thirty worked. Employees can carry over unused sick leave each year, up to a maximum balance of 80 hours. 11 Company-designated paid holidays throughout the year. If employed for at least 12 consecutive months, Align will grant up to 6 weeks of paid Parental Leave. If employed for less than 12 consecutive months, Align will grant up to 4 weeks of paid Parental Leave. All parental leave must be completed within one year of the birth or placement of the child. Parental leave is in addition to any state and/or local parental leave benefits. Three days of paid bereavement leave. In some cases, due to travel the amount of paid leave may be extended to 5 paid days off. To the extent applicable state or local law offers more generous benefits, Align complies with any such law. Non-exempt employees will receive full pay for up to 10 days of jury duty. Exempt employees will receive their full salary during any week they serve and perform any work. Other insurance such as legal, critical illness, voluntary accident, long-term care, auto, home and pet insurance are available for purchase. To the extent applicable state or local law offers more generous benefits, Align complies with any such law.
08/05/2026
Full time
Job Description Job Description Description This position is ideal for senior-level professionals to join the Finance Systems & Reporting team as a Sr. Finance Reporting Engineer. The Finance Engineer will be reporting to the Director, Global Finance Systems & Reporting based in San Jose, California (US) and would strategically engage with key members of finance function defining and delivering on the enterprise roadmap projects as well as own and help resolve day to day tactical problems and challenges. The Sr. Finance Reporting Engineer will be responsible for designing, developing, and maintaining enterprise financial reporting and analytics solutions. This role focuses on delivering accurate, scalable, and automated reporting capabilities by leveraging SAP technologies, cloud-based data platforms, and business intelligence tools to support strategic decision-making across the organization. Role expectations Lead the financial reporting life cycle end-to-end that includes conducting requirement gathering sessions with business, design the reporting solution, development, deployment, maintenance and training of the reporting solution to broader regional audience. Oversee the design and development for the planning system, OneStream, and be prepared to use other planning and reporting tools used at Align, including ERP, Consolidation Planning, and Business Intelligence software packages. Support financial close activities through data reconciliation, validation, and coordination of data loads for month-end, quarter-end, and year-end reporting. Responsible for designing, developing, and maintaining scalable data pipelines and integration workflows to ensure accurate and efficient data processing across systems. Monitor data flows, task chains, and job performance; implementing data quality controls and validation processes; troubleshooting and resolving data load issues; and optimizing data transformation and loading performance. Develop and troubleshoot a variety of Onestream components that include transformation rules, security models, custom dashboards, reports, business rules, and member formulas, with a focus on enhancing system functionality and efficiency. Lead the collaboration and development efforts for various financial Power BI dashboards required for finance, that will make analytics easier and reporting simplified for business at the same time maintaining data accuracy standards as required for financial reporting. Acting as a liaison between business and IT groups. Work with IT to help them understand business requirements and translate it in technical terminologies for them. Help business understand the technical solutions deployed and train them on how they can effectively use it for their needs. Identifying the gaps with existing solutions in place, finding out solutions to resolve them and work with IT/business to get them implemented. Lead the integration testing and user acceptance testing with business and IT collaboratively. Key Personnel responsible for financial close reporting of critical reports Key person responsible for ensuring financial data integrity and completeness for all finance data that will be needed for financial close reporting. Comply with data security and access control standards, maintains thorough process documentation, and provides production support, including participation in on-call rotations with the offshore team when required. Assist with special projects and ad-hoc requests, as necessary. What we're looking for Requires Bachelor's degree in CS, Information Technology, or a related field 12+ years of experience in Data loading and Monitoring systems to support the Finance organization. Strong understanding of financial reporting, FP&A, close processes, and management reporting. Hands-on experience with SAP Datasphere for data modeling, data warehousing, integration, and governance across SAP and non-SAP landscapes. Strong understanding of SAP RTR business processes, hands on experience working with FI-GL, COPA, AR/AP, Should be a self-starter who is able to work with minimal direction and exercises considerable latitude in determining objectives and approaches to assignments. Hands-on experience working with Onestream planning system or similar planning platforms such as Anaplan, Planful. Hands-on experience with SQL and object-oriented (VB.Net, C#) coding experience is preferred. The candidate will serve as a liaison between the finance user group, corporate report development team, and IT. Should be a team player and possess good interpersonal and communication skills, reflecting an ability to be patient and outgoing with people. Should be highly motivated, result focused, and act with a high sense of urgency. Should possess excellent planning and prioritization skills with the ability to multitask and maintain by adapting to change. Complementary skills Advanced Microsoft Outlook, Word, Excel and PowerPoint skills. Must have the ability to independently create spreadsheets and perform quantitative analysis. Prior experience working with Azure Datalake, S/4 HANA, SAP-ECC or SAP BW is a plus. Pay Transparency If provided, base salary or wage rate ranges are the range in which Align reasonably expects to set a candidate's pay for the posted position. Actual placement depends on the individual skills and experience level of a candidate plus the total compensation and equity across team members. For other locations outside of the primary location, the base salary range will be adjusted geographically. For Field Sales roles, the salary listed is the base pay only and does not include the applicable incentive compensation plan. A cost of living adjustment may be added to base pay for higher cost areas in the U.S. Our internship hourly rates are a standard pay determined based on the position and your location, year in school, degree, and experience. General Description of All Benefits We are pleased to provide a general description of the benefits Align offers to full-time employees in this position. Family Benefits. Align offers employees and their eligible dependents medical (with a Health Savings Account option for some plan offerings), dental, and vision in accordance with those plans. Align also offers to employees: Discounts on Invisalign and Vivera to employees and their eligible dependents after 90 days of employment Back-up Child/Elder Care and access to a caregiving concierge Family Forming Benefits - Available to Employees, and their spouse or domestic partner, covered under one of Align's health plans Breast Milk Delivery and Lactation Support Services Employee Assistance Program Hinge Health Virtual Physical Therapy - Available to all employees and eligible dependents (age 18+) enrolled in an Align medical Plan Employee benefits. Align offers its employees: Short-term and long-term disability insurance in accordance with those plans. Basic Life Insurance and Accidental Death and Dismemberment. Voluntary Supplemental Life Insurance for Employee, Spouse/Domestic Partner, and Child(ren) are available for purchase in accordance with those plans. Flexible Spending Accounts- Employees may be eligible to participate in a health care account (including a limited health FSA if enrolled in a HDHP), dependent care account, and a pre-tax commuter benefit plan. 401k plan (with a discretionary Company match of 50% up to 6% of eligible earnings up to a maximum match of 3%.). Employer match vests after two years - 25% year one and 100% at year two. Align offers traditional, Roth, and after-tax options. Employee Stock Purchase Program (Employees must work 20 hours or more and be employed on purchase date to be eligible). Paid vacation of up to 17 days during the first full year of employment (currently accrued at the rate of 5.24 hours each pay-period), which carries over to a maximum cap of 30 days. Annual paid vacation time accrual increases based on tenure. Both exempt and non-exempt employees who work 32 hours or more per week receive prorated vacation accrual based on their regularly scheduled work hours and tenure. Sick time is accrued throughout the year at the rate of one hour for every thirty worked. Employees can carry over unused sick leave each year, up to a maximum balance of 80 hours. 11 Company-designated paid holidays throughout the year. If employed for at least 12 consecutive months, Align will grant up to 6 weeks of paid Parental Leave. If employed for less than 12 consecutive months, Align will grant up to 4 weeks of paid Parental Leave. All parental leave must be completed within one year of the birth or placement of the child. Parental leave is in addition to any state and/or local parental leave benefits. Three days of paid bereavement leave. In some cases, due to travel the amount of paid leave may be extended to 5 paid days off. To the extent applicable state or local law offers more generous benefits, Align complies with any such law. Non-exempt employees will receive full pay for up to 10 days of jury duty. Exempt employees will receive their full salary during any week they serve and perform any work. Other insurance such as legal, critical illness, voluntary accident, long-term care, auto, home and pet insurance are available for purchase. To the extent applicable state or local law offers more generous benefits, Align complies with any such law.
Site Reliability Engineer (SRE) / DevOps Engineer
E-Space Saratoga, California
Job Description Job Description Ready to make connectivity from space universally accessible, secure and actionable? Then you've come to the right place! E-Space is bridging Earth and space to enable hyper-scaled deployments of Internet of Things (IoT) solutions and services. We are building a highly-advanced low Earth orbit (LEO) space system that will fundamentally change the design, economics, manufacturing and service delivery associated with traditional satellite and terrestrial IoT systems. We're intentional, we're unapologetically curious and we're 100% committed to innovate space-based communications and deliver actionable intelligence that will expand global economies, protect space and our planet and enhance our overall quality of life. We are seeking a DevOps Engineer who is eager to have an immediate impact in establishing and building out our development, testing, and release platforms, taking us from square one to a sophisticated DevOps environment. You will be working with key engineering stakeholders across the company to ensure we are following test-and-release best practices, establishing best-in-class technology and tool stacks, and keeping our environments safe and secure from outside threats. What you will be doing: Design, deploy, and maintain highly-scalable, highly-available software systems in AWS Architect and manage containerized applications on Amazon EKS with focus on reliability and performance Build and maintain Infrastructure as Code using Terraform for AWS cloud resources Develop and optimize CI/CD pipelines for automated testing, deployment, and rollback capabilities Implement comprehensive monitoring, alerting, and observability solutions using CloudWatch, Prometheus, and Grafana Ensure system reliability through SLI/SLO definition, error budgets, and incident response procedures Collaborate directly with engineering teams to optimize application deployment and operations Manage deployments and scaling strategies to support mission-critical operations Automate and enforce cloud security, governance, and compliance controls Participate in on-call rotation and lead incident response for production level systems What you bring to this role: 5+ years of experience in SRE, DevOps, or Platform Engineering roles Proven experience designing and operating mission-critical, highly-available systems within AWS Advanced proficiency in Infrastructure as Code using Terraform (OpenTofu) Deep experience with Kubernetes, EKS, Helm, and container orchestration Strong CI/CD pipeline development and management experience (Bitbucket preferred) Proficiency in Python and Bash scripting for automation Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack) Knowledge of capacity planning and performance optimization Experience with database operations and scaling (RDS, Aurora, or similar) Extra bonus points for the following: AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA), or equivalent expertise Experience with incident management and post-mortem processes Experience with GitOps workflows and tools (ArgoCD, Flux) Knowledge of service mesh technologies (Istio, Linkerd) Experience with chaos engineering and disaster recovery planning Experience with Zero Trust Networking (ZTNA) or VPN solutions Background in aerospace, defense, or other mission-critical industries Strong intellectual curiosity and commitment to continuous learning Exceptional attention to detail and an ownership mentality The estimated range is meant to reflect an anticipated salary range for the position in question, which is based on market data and other factors, all of which are subject to change. Individual pay is based on location, skills and expertise, depth of relevant experience, and other relevant factors. For questions about this, please speak to the recruiter if you decide to apply for the role and are selected for an interview . This is a full time, exempt position, based out of our Saratoga office. The target base pay for this position is $100,000 - $170,000 annually. The total compensation packaged will be determined by various factors such as your relevant job-related knowledge, skills, and experience. We are redefining how satellites are designed, manufactured and used-so we're looking for candidates with passion, deep knowledge and direct experience on LEO satellite component development, design and in-orbit activities. If that's your experience - then we'll be immediately wow-ed. E-Space is not currently able to provide employment sponsorship for candidates who do not hold work authorization for the location of this role. Why E-Space is right for you: As a member of our team, you will play a crucial role in driving our success. Our team members have a strong sense of dedication and responsibility; this includes a strong commitment to our mission to create an entirely new suite of global capabilities to improve lives, business efficiencies and build a smarter planet. This means that there will be times when extra hours, including nights and weekends, may be needed to meet critical deadlines and mission goals. In return, we offer a dynamic work environment with opportunities for professional growth and development and the chance to make a meaningful impact in a high-growth industry. We want you to make the most of your journey at E-Space. That's why we support and invest in the physical, emotional and financial well-being of our team members and their families. Some of what you can expect when working at E-Space: • An opportunity to really make a difference • Sustainability at our core • Fair and honest workplace • Innovative thinking is encouraged • Competitive salaries • Continuous learning and development • Health and wellness care options • Financial solutions for the future • Optional legal services (US only) • Paid holidays • Paid time off We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
08/05/2026
Full time
Job Description Job Description Ready to make connectivity from space universally accessible, secure and actionable? Then you've come to the right place! E-Space is bridging Earth and space to enable hyper-scaled deployments of Internet of Things (IoT) solutions and services. We are building a highly-advanced low Earth orbit (LEO) space system that will fundamentally change the design, economics, manufacturing and service delivery associated with traditional satellite and terrestrial IoT systems. We're intentional, we're unapologetically curious and we're 100% committed to innovate space-based communications and deliver actionable intelligence that will expand global economies, protect space and our planet and enhance our overall quality of life. We are seeking a DevOps Engineer who is eager to have an immediate impact in establishing and building out our development, testing, and release platforms, taking us from square one to a sophisticated DevOps environment. You will be working with key engineering stakeholders across the company to ensure we are following test-and-release best practices, establishing best-in-class technology and tool stacks, and keeping our environments safe and secure from outside threats. What you will be doing: Design, deploy, and maintain highly-scalable, highly-available software systems in AWS Architect and manage containerized applications on Amazon EKS with focus on reliability and performance Build and maintain Infrastructure as Code using Terraform for AWS cloud resources Develop and optimize CI/CD pipelines for automated testing, deployment, and rollback capabilities Implement comprehensive monitoring, alerting, and observability solutions using CloudWatch, Prometheus, and Grafana Ensure system reliability through SLI/SLO definition, error budgets, and incident response procedures Collaborate directly with engineering teams to optimize application deployment and operations Manage deployments and scaling strategies to support mission-critical operations Automate and enforce cloud security, governance, and compliance controls Participate in on-call rotation and lead incident response for production level systems What you bring to this role: 5+ years of experience in SRE, DevOps, or Platform Engineering roles Proven experience designing and operating mission-critical, highly-available systems within AWS Advanced proficiency in Infrastructure as Code using Terraform (OpenTofu) Deep experience with Kubernetes, EKS, Helm, and container orchestration Strong CI/CD pipeline development and management experience (Bitbucket preferred) Proficiency in Python and Bash scripting for automation Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack) Knowledge of capacity planning and performance optimization Experience with database operations and scaling (RDS, Aurora, or similar) Extra bonus points for the following: AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA), or equivalent expertise Experience with incident management and post-mortem processes Experience with GitOps workflows and tools (ArgoCD, Flux) Knowledge of service mesh technologies (Istio, Linkerd) Experience with chaos engineering and disaster recovery planning Experience with Zero Trust Networking (ZTNA) or VPN solutions Background in aerospace, defense, or other mission-critical industries Strong intellectual curiosity and commitment to continuous learning Exceptional attention to detail and an ownership mentality The estimated range is meant to reflect an anticipated salary range for the position in question, which is based on market data and other factors, all of which are subject to change. Individual pay is based on location, skills and expertise, depth of relevant experience, and other relevant factors. For questions about this, please speak to the recruiter if you decide to apply for the role and are selected for an interview . This is a full time, exempt position, based out of our Saratoga office. The target base pay for this position is $100,000 - $170,000 annually. The total compensation packaged will be determined by various factors such as your relevant job-related knowledge, skills, and experience. We are redefining how satellites are designed, manufactured and used-so we're looking for candidates with passion, deep knowledge and direct experience on LEO satellite component development, design and in-orbit activities. If that's your experience - then we'll be immediately wow-ed. E-Space is not currently able to provide employment sponsorship for candidates who do not hold work authorization for the location of this role. Why E-Space is right for you: As a member of our team, you will play a crucial role in driving our success. Our team members have a strong sense of dedication and responsibility; this includes a strong commitment to our mission to create an entirely new suite of global capabilities to improve lives, business efficiencies and build a smarter planet. This means that there will be times when extra hours, including nights and weekends, may be needed to meet critical deadlines and mission goals. In return, we offer a dynamic work environment with opportunities for professional growth and development and the chance to make a meaningful impact in a high-growth industry. We want you to make the most of your journey at E-Space. That's why we support and invest in the physical, emotional and financial well-being of our team members and their families. Some of what you can expect when working at E-Space: • An opportunity to really make a difference • Sustainability at our core • Fair and honest workplace • Innovative thinking is encouraged • Competitive salaries • Continuous learning and development • Health and wellness care options • Financial solutions for the future • Optional legal services (US only) • Paid holidays • Paid time off We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Senior Site Reliability Engineer- Palo Alto, the US
Kody Palo Alto, California
Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment. Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Competitive packages aligned with California market standards Lead a dynamic and innovative team in a very rapidly growing company Collaborative, inclusive environment where your contributions are recognized and valued
08/05/2026
Full time
Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment. Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Competitive packages aligned with California market standards Lead a dynamic and innovative team in a very rapidly growing company Collaborative, inclusive environment where your contributions are recognized and valued
Senior Site Reliability Engineer- Sunnyvale, CA, the US
Kody Sunnyvale, California
Job Description Job Description About the Role Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Requirements Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Lead a dynamic and innovative team in a very rapidly growing company. Competitive package. Collaborative, inclusive environment where your contributions are recognized and valued.
08/05/2026
Full time
Job Description Job Description About the Role Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Requirements Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Lead a dynamic and innovative team in a very rapidly growing company. Competitive package. Collaborative, inclusive environment where your contributions are recognized and valued.
Site Reliability Engineer
VantageScore San Francisco, California
Job Description Job Description About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure, APIs, and software supply chain. You will drive patch management programs, harden our Cloud infrastructure, and maintain our code repositories to ensure all systems remain compliant, secure, and scalable. This role is ideal for an engineer who thrives at the intersection of operations and security, is passionate about automation, and takes pride in keeping complex environments clean, auditable, and resilient. Key Responsibilities Own and execute end-to-end patch management across AWS compute resources (EC2, ECS, Lambda runtimes, EKS nodes), third-party dependencies, and OS-level packages. Monitor, triage, and remediate vulnerabilities identified by security scanning tools (e.g., AWS Inspector, Dependabot, Security Hub, or equivalent), prioritizing by CVSS severity and business impact. Maintain and enforce branch protection rules, secret scanning policies, and dependency update workflows across all code repositories. Design and implement automated pipelines for continuous compliance checking, security testing (SAST/DAST/SCA), and infrastructure drift detection. Collaborate with IT & Info-Sec SMEs on AWS IAM roles and policies, VPC configurations, Security Groups, CloudTrail, Config, and GuardDuty to ensure least-privilege access and auditability. Collaborate with development teams to embed security controls into CI/CD pipelines (GitHub Actions, CodePipeline, or equivalent) without impeding developer velocity. Support the reliability and availability of production APIs - including uptime monitoring, incident response, runbook creation, and post-incident reviews. Partner with Legal and Data Governance SMEs on API access procedures and monitoring. Define and track SLOs/SLAs for internal and external APIs; implement alerting and dashboards using observability tooling (e.g., CloudWatch, Datadog, Grafana). Lead periodic infrastructure and dependency audits; produce clear reports on patch compliance status and open risk items for engineering and security leadership. Maintain thorough documentation of patching schedules, runbooks, access policies, and environment configurations. Participate in on-call rotation and contribute to a culture of continuous improvement. Required Qualifications Bachelor's Degree in Computer Science, Information Systems, or a related field (or equivalent practical experience). 5+ years of professional experience in a Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. Demonstrated expertise managing AWS environments - including EC2, Lambda, ECS/EKS, S3, RDS, IAM, VPC, CloudTrail, Config, and GuardDuty. Experience with various cloud environments: AWS, Azure, GPC Strong experience with GitHub administration: branch protection, Actions workflows, secret scanning, Dependabot, and code owners. Hands-on experience with patch management and vulnerability remediation at scale, including OS-level patching (Amazon Linux, Ubuntu) and dependency lifecycle management. Proficiency with infrastructure-as-code tools (Terraform, CloudFormation, or AWS CDK). Experience integrating security tooling (SAST, DAST, SCA, container scanning) into CI/CD pipelines. Solid understanding of API reliability patterns: health checks, rate limiting, circuit breakers, and observability. Familiarity with compliance frameworks relevant to cloud environments (SOC 2, CIS Benchmarks, NIST CSF). Strong scripting skills in Python, Bash, or similar for automation and tooling. Excellent communication skills and ability to translate technical risk for non-technical stakeholders. Build observation (logging, metrics, alerting) systems to make sure system works well, and develop response plans. Preferred Qualifications AWS certifications (e.g., AWS Certified Security - Specialty, AWS Certified DevOps Engineer - Professional). Experience with container security and Kubernetes (EKS) hardening. Familiarity with CSPM tools (e.g., Wiz, Prisma Cloud, AWS Security Hub) for continuous cloud posture management. Experience managing API gateways (AWS API Gateway, Kong, or similar) including security policy enforcement. Exposure to secrets management solutions (AWS Secrets Manager, HashiCorp Vault). Knowledge of SBOM (Software Bill of Materials) generation and management. Experience with incident response playbooks and tabletop exercises. Familiarity with Agile/Scrum methodologies and cross-functional engineering teams. Compensation The anticipated base salary range for this position is $150,000 annually, plus eligibility for a 15% annual performance bonus. Actual compensation will be determined based on several factors, including skills, experience, education, certifications, and geographic location. In addition to base salary and bonus eligibility, we offer a competitive benefits package, including medical, dental, vision, 401(k), paid time off, and other employee benefits.
08/05/2026
Full time
Job Description Job Description About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure, APIs, and software supply chain. You will drive patch management programs, harden our Cloud infrastructure, and maintain our code repositories to ensure all systems remain compliant, secure, and scalable. This role is ideal for an engineer who thrives at the intersection of operations and security, is passionate about automation, and takes pride in keeping complex environments clean, auditable, and resilient. Key Responsibilities Own and execute end-to-end patch management across AWS compute resources (EC2, ECS, Lambda runtimes, EKS nodes), third-party dependencies, and OS-level packages. Monitor, triage, and remediate vulnerabilities identified by security scanning tools (e.g., AWS Inspector, Dependabot, Security Hub, or equivalent), prioritizing by CVSS severity and business impact. Maintain and enforce branch protection rules, secret scanning policies, and dependency update workflows across all code repositories. Design and implement automated pipelines for continuous compliance checking, security testing (SAST/DAST/SCA), and infrastructure drift detection. Collaborate with IT & Info-Sec SMEs on AWS IAM roles and policies, VPC configurations, Security Groups, CloudTrail, Config, and GuardDuty to ensure least-privilege access and auditability. Collaborate with development teams to embed security controls into CI/CD pipelines (GitHub Actions, CodePipeline, or equivalent) without impeding developer velocity. Support the reliability and availability of production APIs - including uptime monitoring, incident response, runbook creation, and post-incident reviews. Partner with Legal and Data Governance SMEs on API access procedures and monitoring. Define and track SLOs/SLAs for internal and external APIs; implement alerting and dashboards using observability tooling (e.g., CloudWatch, Datadog, Grafana). Lead periodic infrastructure and dependency audits; produce clear reports on patch compliance status and open risk items for engineering and security leadership. Maintain thorough documentation of patching schedules, runbooks, access policies, and environment configurations. Participate in on-call rotation and contribute to a culture of continuous improvement. Required Qualifications Bachelor's Degree in Computer Science, Information Systems, or a related field (or equivalent practical experience). 5+ years of professional experience in a Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. Demonstrated expertise managing AWS environments - including EC2, Lambda, ECS/EKS, S3, RDS, IAM, VPC, CloudTrail, Config, and GuardDuty. Experience with various cloud environments: AWS, Azure, GPC Strong experience with GitHub administration: branch protection, Actions workflows, secret scanning, Dependabot, and code owners. Hands-on experience with patch management and vulnerability remediation at scale, including OS-level patching (Amazon Linux, Ubuntu) and dependency lifecycle management. Proficiency with infrastructure-as-code tools (Terraform, CloudFormation, or AWS CDK). Experience integrating security tooling (SAST, DAST, SCA, container scanning) into CI/CD pipelines. Solid understanding of API reliability patterns: health checks, rate limiting, circuit breakers, and observability. Familiarity with compliance frameworks relevant to cloud environments (SOC 2, CIS Benchmarks, NIST CSF). Strong scripting skills in Python, Bash, or similar for automation and tooling. Excellent communication skills and ability to translate technical risk for non-technical stakeholders. Build observation (logging, metrics, alerting) systems to make sure system works well, and develop response plans. Preferred Qualifications AWS certifications (e.g., AWS Certified Security - Specialty, AWS Certified DevOps Engineer - Professional). Experience with container security and Kubernetes (EKS) hardening. Familiarity with CSPM tools (e.g., Wiz, Prisma Cloud, AWS Security Hub) for continuous cloud posture management. Experience managing API gateways (AWS API Gateway, Kong, or similar) including security policy enforcement. Exposure to secrets management solutions (AWS Secrets Manager, HashiCorp Vault). Knowledge of SBOM (Software Bill of Materials) generation and management. Experience with incident response playbooks and tabletop exercises. Familiarity with Agile/Scrum methodologies and cross-functional engineering teams. Compensation The anticipated base salary range for this position is $150,000 annually, plus eligibility for a 15% annual performance bonus. Actual compensation will be determined based on several factors, including skills, experience, education, certifications, and geographic location. In addition to base salary and bonus eligibility, we offer a competitive benefits package, including medical, dental, vision, 401(k), paid time off, and other employee benefits.
Site Reliability Engineer, Client Platform
Hadrian Automation Bodega Bay, California
Job Description Job Description Hadrian - Manufacturing the Future Hadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x cheaper. By combining advanced software, robotics, and full-stack manufacturing, we are reinventing how America produces its most critical parts. We're accelerating our mission with the launch of Factory 3 in Mesa, Arizona, a 290,000-square-foot facility creating 350 new jobs. We are expanding rapidly to support thousands of future hires, launching Hadrian Maritime to expand into naval production, and introducing a Factory-as-a-Service model that delivers complete systems instead of individual parts. Hadrian is backed by leading investors including T. Rowe Price, Lux Capital, Founders Fund, and Andreessen Horowitz, our fast-growing team is united around reindustrializing American manufacturing for the 21st century and beyond. The Role: What You'll Do Focus on building scalable , automated solutions that ensure seamless deployments, security configurations, and efficient operational workflows for our end-users. Own , administer , and optimize MDM platforms (Fleet DM, Intune, Workspace ONE) to enforce configuration and drive self-healing by writing OS-level scripts and lightweight tools that resolve recurring user-impacting issues (disk pressure, certificate expiry, drift, broken agents) at the source instead of via tickets. Proactive Response . We want to gather telemetry data and analytics to develop an understanding of device lifecycles. We want to prevent end-user disruption by understanding when and how to act. Partner with Security, IT, and Infrastructure to translate compliance requirements (CMMC) into enforceable, code-managed baselines. Build dashboards and alerts that measure end-user experience as an SLO, not a helpdesk metric. What We're Looking For Ownership. You treat the fleet as a product, take incidents personally, and close the loop with automation rather than a runbook. Strong scripting in Python, Bash, and PowerShell. Scalability. Hands-on experience with Infrastructure as Code (IaC) and configuration management: Ansible and Terraform (or equivalents like Chef, Salt, Puppet, Pulumi). Device Management. Working knowledge of at least one major MDM (Fleet DM, Intune, Jamf, Workspace ONE) and its API surface. Security Remediation. Practical experience with patch management , vulnerability remediation , and endpoint hardening on both macOS and Windows. What Will Set You Apart Strong computer science fundamentals. You can reason about systems from the operating system up, demonstrating durable and sustainable solutions. Experience building self-healing or auto-remediation platforms (remote actions, osquery + response, custom agents). Exposure to OT (Operational Technology) systems. Comfort operating in an SRE culture: SLOs, error budgets, and blameless postmortems applied to the end-user experience. Compensation For this role, the target salary range is $164,000 - $270,000 (actual range may vary based on experience). This is the lowest to highest salary we reasonably and in good faith believe we would pay for this role at the time of this posting. We may ultimately pay more or less than the posted range, and the range may be modified in the future. An employee's pay position within the salary range will be based on several factors, including, but not limited to, relevant education, qualifications, certifications, experience, skills, geographic location, performance, and business or organizational needs. Benefits for Full-time Employees Medical, dental, vision, and life insurance plans for employees 401k Relocation support may be provided for certain situations, based on business need. Flexible vacation policy Equity ITAR Requirements To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR) you must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. Hadrian Is An Equal Opportunity Employer It is the Company's policy to provide equal employment opportunity for all applicants and employees. The Company does not unlawfully discriminate on the basis of race inclusive of traits historically associated with race (including, but not limited to, hair texture and protective hairstyles, such as braids, locks and twists), color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender identity, gender expression, transgender status, national origin (including, in California, possession of a drivers license), ancestry, citizenship, age, physical or mental disability, height or weight, medical condition, family care status, military or veteran status, marital status, domestic partner status, sexual orientation, genetic information, exercise of reproductive rights, any other basis protected by local, state, or federal laws, or any combination of the above characteristics. When necessary, the Company also makes reasonable accommodations for disabled candidates and employees, including for candidates or employees who are disabled by pregnancy, childbirth, or related medical conditions.
08/05/2026
Full time
Job Description Job Description Hadrian - Manufacturing the Future Hadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x cheaper. By combining advanced software, robotics, and full-stack manufacturing, we are reinventing how America produces its most critical parts. We're accelerating our mission with the launch of Factory 3 in Mesa, Arizona, a 290,000-square-foot facility creating 350 new jobs. We are expanding rapidly to support thousands of future hires, launching Hadrian Maritime to expand into naval production, and introducing a Factory-as-a-Service model that delivers complete systems instead of individual parts. Hadrian is backed by leading investors including T. Rowe Price, Lux Capital, Founders Fund, and Andreessen Horowitz, our fast-growing team is united around reindustrializing American manufacturing for the 21st century and beyond. The Role: What You'll Do Focus on building scalable , automated solutions that ensure seamless deployments, security configurations, and efficient operational workflows for our end-users. Own , administer , and optimize MDM platforms (Fleet DM, Intune, Workspace ONE) to enforce configuration and drive self-healing by writing OS-level scripts and lightweight tools that resolve recurring user-impacting issues (disk pressure, certificate expiry, drift, broken agents) at the source instead of via tickets. Proactive Response . We want to gather telemetry data and analytics to develop an understanding of device lifecycles. We want to prevent end-user disruption by understanding when and how to act. Partner with Security, IT, and Infrastructure to translate compliance requirements (CMMC) into enforceable, code-managed baselines. Build dashboards and alerts that measure end-user experience as an SLO, not a helpdesk metric. What We're Looking For Ownership. You treat the fleet as a product, take incidents personally, and close the loop with automation rather than a runbook. Strong scripting in Python, Bash, and PowerShell. Scalability. Hands-on experience with Infrastructure as Code (IaC) and configuration management: Ansible and Terraform (or equivalents like Chef, Salt, Puppet, Pulumi). Device Management. Working knowledge of at least one major MDM (Fleet DM, Intune, Jamf, Workspace ONE) and its API surface. Security Remediation. Practical experience with patch management , vulnerability remediation , and endpoint hardening on both macOS and Windows. What Will Set You Apart Strong computer science fundamentals. You can reason about systems from the operating system up, demonstrating durable and sustainable solutions. Experience building self-healing or auto-remediation platforms (remote actions, osquery + response, custom agents). Exposure to OT (Operational Technology) systems. Comfort operating in an SRE culture: SLOs, error budgets, and blameless postmortems applied to the end-user experience. Compensation For this role, the target salary range is $164,000 - $270,000 (actual range may vary based on experience). This is the lowest to highest salary we reasonably and in good faith believe we would pay for this role at the time of this posting. We may ultimately pay more or less than the posted range, and the range may be modified in the future. An employee's pay position within the salary range will be based on several factors, including, but not limited to, relevant education, qualifications, certifications, experience, skills, geographic location, performance, and business or organizational needs. Benefits for Full-time Employees Medical, dental, vision, and life insurance plans for employees 401k Relocation support may be provided for certain situations, based on business need. Flexible vacation policy Equity ITAR Requirements To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR) you must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. Hadrian Is An Equal Opportunity Employer It is the Company's policy to provide equal employment opportunity for all applicants and employees. The Company does not unlawfully discriminate on the basis of race inclusive of traits historically associated with race (including, but not limited to, hair texture and protective hairstyles, such as braids, locks and twists), color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender identity, gender expression, transgender status, national origin (including, in California, possession of a drivers license), ancestry, citizenship, age, physical or mental disability, height or weight, medical condition, family care status, military or veteran status, marital status, domestic partner status, sexual orientation, genetic information, exercise of reproductive rights, any other basis protected by local, state, or federal laws, or any combination of the above characteristics. When necessary, the Company also makes reasonable accommodations for disabled candidates and employees, including for candidates or employees who are disabled by pregnancy, childbirth, or related medical conditions.
Sr. Site Reliability Engineer
PayNearMe, Inc. Alviso, California
Job Description Job Description Company Description At PayNearMe, we're on a mission to make paying and getting paid as simple as possible. We build innovative technology that transforms the way businesses and their customers experience payments. Our industry-leading platform, PayXM , is the first of its kind-designed to manage the entire payment experience from start to finish. Every click, swipe or tap is seamless, fast and secure, helping non-commerce businesses boost customer satisfaction, accelerate payments, and reduce costs. Our single platform handles it all: cards, ACH, digital wallets such as PayPal, Venmo, Cash App Pay, Apple Pay and Google Pay, and even cash at more than 62,000 retail locations nationwide. Today, thousands of businesses across consumer lending, iGaming and online sports betting, property management, and tolling trust PayNearMe to deliver a payment experience that drives real results. In September 2025, we raised a $50 million Series E funding round to accelerate our growth. We're a team of 300+ employees across 41 states, headquartered in Silicon Valley with satellite offices in Dallas, TX and Holmdel, NJ. Join us and be part of a team that's shaping the future of payments-one experience at a time. As our Site Reliability Engineer, you will design, build, and maintain the systems and infrastructure that power our applications, ensuring their reliability, scalability, and performance. You will bring a software engineering approach to operations, automating processes, and continuously improving the infrastructure and tools to support our business needs. Responsibilities Infrastructure Management: Design, implement, and maintain scalable and resilient infrastructure using Terraform for infrastructure as code, ensuring high availability and performance Kubernetes and Containers: Deploy, manage, and optimize Kubernetes clusters and containerized applications using Docker. Implement best practices for container orchestration and management Systems and Application Monitoring/Observability: Develop and maintain comprehensive monitoring and observability solutions using Datadog. Ensure detailed visibility into system performance and application health SLOs and SLA Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to ensure reliable and consistent service delivery Incident Response and Troubleshooting: Respond to incidents, perform root cause analysis, and implement solutions to prevent recurrence. Participate in post-incident reviews and contribute to blameless postmortems Reliability and Production Environment Management: Ensure the reliability and stability of our production environments. Continuously assess and improve system reliability, identifying and addressing potential points of failure Automation and Scripting: Develop automation scripts and tools to reduce manual intervention and improve system reliability using Python, Bash, or Go. Implement and improve CI/CD pipelines CI/CD Pipeline Management: Enhance and maintain continuous integration and continuous deployment pipelines using GitLab CI. Ensure seamless and reliable deployment processes Capacity Planning and Scaling: Assist in capacity planning and ensure that systems are scalable to meet future demands. Implement auto-scaling strategies where applicable Security and Compliance: Implement security best practices and ensure compliance with industry standards. Regularly review and update security policies and procedures Collaboration and Support: Work closely with development teams to ensure reliability and scalability of new features and services. Provide technical support and guidance on infrastructure-related issues Software Engineering for Operations: Develop and maintain internal tools and services that enhance the efficiency and reliability of our operations On-Call Rotation: Participate in an on-call rotation to address production issues and collaborate in incident response efforts Qualifications +3 years of experience in SRE, DevOps, or a related role Cloud Platform Experience: Proficient with cloud platforms such as AWS, GCP, or Azure Experience with EC2, RDS, VPCs, and security groups is essential. Kubernetes and Containers: Strong experience with Kubernetes and Docker, including deployment, scaling, and management of containerized applications Infrastructure as Code: Expert in using Terraform for infrastructure as code. Proficient with configuration management tools such as Ansible, Puppet, or Chef Monitoring and Observability: Extensive experience with monitoring and observability tools like Datadog, Prometheus, Grafana, ELK stack, or Splunk. Skilled in setting up detailed monitoring and logging systems SLOs and SLA Management: Proven ability to define, monitor, and maintain SLOs and SLAs to ensure reliable service delivery Scripting and Automation: Strong skills in scripting languages like Python, Bash, or Go. Experience automating repetitive tasks and processes CI/CD Practices: Familiarity with GitLab CI or similar tool for continuous integration and deployment. Experience in setting up and managing pipelines Production Environments: Experience supporting production environments running Go or Ruby/Rails applications Tool Development: Ability to write and update tools to support infrastructure and application management, demonstrating the principle that "SRE is what happens when you ask a software engineer to design an operations team DevOps Best Practices: Deep understanding of DevOps principles, practices, and tools to drive continuous improvement in the software development lifecycle Soft Skills: Strong organizational skills, attention to detail, and the ability to work collaboratively in a team environment. Excellent documentation skills to ensure accurate and detailed records Problem-Solving Ability: Excellent analytical and problem-solving skills to diagnose and resolve complex system issues quickly and effectively The annual base salary range for this role represents PayNearMe's good-faith estimate of the base salary it reasonably expects to offer for this position at the time of hire. Actual compensation may vary based on factors including the candidate's experience, qualifications, skills, and work location. PayNearMe may offer compensation outside of this range in certain circumstances. This position will remain posted until filled. Annual Salary Range $180,000-$200,000 USD Why Join Us?: Competitive salary and benefits with growth-company options grant Fast- paced and professional work culture Stock options with standard startup vesting - 1 year cliff; 4 years total $50 monthly communication expense stipend to go towards your phone/internet bill $250 stipend to enhance your WFH setup Reimbursement for peripheral equipment: monitor (up to $400), keyboard and mouse (up to $200) Premium medical benefits including vision and dental (100% coverage for employees) Company-sponsored life and disability insurance Paid parental bonding leave Paid sick leave, jury duty, bereavement 401k plan Flexible Time Off (our team members typically take off 3-4 weeks per year) Volunteer Time Off 13 scheduled holidays PayNearMe strives to create a workplace where all employees thrive. Our core values represent who we are today and we take pride in the way we work with each other as well as with our stakeholders. We're in this together to do the right thing. We deliver real results we are proud of while remaining respectful, transparent, and flexible. PayNearMe is an equal opportunity employer. We are diligently and thoughtfully working towards cultivating a diverse workforce which in turn, enhances our products and services for the communities we serve. Applicants who represent all backgrounds are strongly encouraged to apply. CALIFORNIA CONSUMER PRIVACY ACT: APPLICANT NOTICE Effective Date: January 1, 2020 Last Reviewed on: December 23, 2019 PayNearMe, Inc. (the "Company") is providing you with this Notice ("Notice") to inform you about: the categories of Personal Information that the Company collects and maintains about applicants; and the purposes for which the Company uses that Personal Information. For purposes of this Notice, "Personal Information" means information that identifies, relates to, describes, is capable of being associated with, or could reasonably be linked, directly or indirectly with, a natural person that the Company may collect in connection with screening applicants for job openings at the Company. Identifiers and Professional or Employment-Related Information. The Company collects identifiers and professional or employment-related information, which may include some or all the following: real name, nickname or alias, postal address, telephone number, e-mail address, membership in professional organizations, professional certifications, language skills, and current and past employment history. The Company collects this Personal Information to evaluate previous job performance and consider applicants for positions, to develop a talent pool and plan for succession . click apply for full job details
08/05/2026
Full time
Job Description Job Description Company Description At PayNearMe, we're on a mission to make paying and getting paid as simple as possible. We build innovative technology that transforms the way businesses and their customers experience payments. Our industry-leading platform, PayXM , is the first of its kind-designed to manage the entire payment experience from start to finish. Every click, swipe or tap is seamless, fast and secure, helping non-commerce businesses boost customer satisfaction, accelerate payments, and reduce costs. Our single platform handles it all: cards, ACH, digital wallets such as PayPal, Venmo, Cash App Pay, Apple Pay and Google Pay, and even cash at more than 62,000 retail locations nationwide. Today, thousands of businesses across consumer lending, iGaming and online sports betting, property management, and tolling trust PayNearMe to deliver a payment experience that drives real results. In September 2025, we raised a $50 million Series E funding round to accelerate our growth. We're a team of 300+ employees across 41 states, headquartered in Silicon Valley with satellite offices in Dallas, TX and Holmdel, NJ. Join us and be part of a team that's shaping the future of payments-one experience at a time. As our Site Reliability Engineer, you will design, build, and maintain the systems and infrastructure that power our applications, ensuring their reliability, scalability, and performance. You will bring a software engineering approach to operations, automating processes, and continuously improving the infrastructure and tools to support our business needs. Responsibilities Infrastructure Management: Design, implement, and maintain scalable and resilient infrastructure using Terraform for infrastructure as code, ensuring high availability and performance Kubernetes and Containers: Deploy, manage, and optimize Kubernetes clusters and containerized applications using Docker. Implement best practices for container orchestration and management Systems and Application Monitoring/Observability: Develop and maintain comprehensive monitoring and observability solutions using Datadog. Ensure detailed visibility into system performance and application health SLOs and SLA Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to ensure reliable and consistent service delivery Incident Response and Troubleshooting: Respond to incidents, perform root cause analysis, and implement solutions to prevent recurrence. Participate in post-incident reviews and contribute to blameless postmortems Reliability and Production Environment Management: Ensure the reliability and stability of our production environments. Continuously assess and improve system reliability, identifying and addressing potential points of failure Automation and Scripting: Develop automation scripts and tools to reduce manual intervention and improve system reliability using Python, Bash, or Go. Implement and improve CI/CD pipelines CI/CD Pipeline Management: Enhance and maintain continuous integration and continuous deployment pipelines using GitLab CI. Ensure seamless and reliable deployment processes Capacity Planning and Scaling: Assist in capacity planning and ensure that systems are scalable to meet future demands. Implement auto-scaling strategies where applicable Security and Compliance: Implement security best practices and ensure compliance with industry standards. Regularly review and update security policies and procedures Collaboration and Support: Work closely with development teams to ensure reliability and scalability of new features and services. Provide technical support and guidance on infrastructure-related issues Software Engineering for Operations: Develop and maintain internal tools and services that enhance the efficiency and reliability of our operations On-Call Rotation: Participate in an on-call rotation to address production issues and collaborate in incident response efforts Qualifications +3 years of experience in SRE, DevOps, or a related role Cloud Platform Experience: Proficient with cloud platforms such as AWS, GCP, or Azure Experience with EC2, RDS, VPCs, and security groups is essential. Kubernetes and Containers: Strong experience with Kubernetes and Docker, including deployment, scaling, and management of containerized applications Infrastructure as Code: Expert in using Terraform for infrastructure as code. Proficient with configuration management tools such as Ansible, Puppet, or Chef Monitoring and Observability: Extensive experience with monitoring and observability tools like Datadog, Prometheus, Grafana, ELK stack, or Splunk. Skilled in setting up detailed monitoring and logging systems SLOs and SLA Management: Proven ability to define, monitor, and maintain SLOs and SLAs to ensure reliable service delivery Scripting and Automation: Strong skills in scripting languages like Python, Bash, or Go. Experience automating repetitive tasks and processes CI/CD Practices: Familiarity with GitLab CI or similar tool for continuous integration and deployment. Experience in setting up and managing pipelines Production Environments: Experience supporting production environments running Go or Ruby/Rails applications Tool Development: Ability to write and update tools to support infrastructure and application management, demonstrating the principle that "SRE is what happens when you ask a software engineer to design an operations team DevOps Best Practices: Deep understanding of DevOps principles, practices, and tools to drive continuous improvement in the software development lifecycle Soft Skills: Strong organizational skills, attention to detail, and the ability to work collaboratively in a team environment. Excellent documentation skills to ensure accurate and detailed records Problem-Solving Ability: Excellent analytical and problem-solving skills to diagnose and resolve complex system issues quickly and effectively The annual base salary range for this role represents PayNearMe's good-faith estimate of the base salary it reasonably expects to offer for this position at the time of hire. Actual compensation may vary based on factors including the candidate's experience, qualifications, skills, and work location. PayNearMe may offer compensation outside of this range in certain circumstances. This position will remain posted until filled. Annual Salary Range $180,000-$200,000 USD Why Join Us?: Competitive salary and benefits with growth-company options grant Fast- paced and professional work culture Stock options with standard startup vesting - 1 year cliff; 4 years total $50 monthly communication expense stipend to go towards your phone/internet bill $250 stipend to enhance your WFH setup Reimbursement for peripheral equipment: monitor (up to $400), keyboard and mouse (up to $200) Premium medical benefits including vision and dental (100% coverage for employees) Company-sponsored life and disability insurance Paid parental bonding leave Paid sick leave, jury duty, bereavement 401k plan Flexible Time Off (our team members typically take off 3-4 weeks per year) Volunteer Time Off 13 scheduled holidays PayNearMe strives to create a workplace where all employees thrive. Our core values represent who we are today and we take pride in the way we work with each other as well as with our stakeholders. We're in this together to do the right thing. We deliver real results we are proud of while remaining respectful, transparent, and flexible. PayNearMe is an equal opportunity employer. We are diligently and thoughtfully working towards cultivating a diverse workforce which in turn, enhances our products and services for the communities we serve. Applicants who represent all backgrounds are strongly encouraged to apply. CALIFORNIA CONSUMER PRIVACY ACT: APPLICANT NOTICE Effective Date: January 1, 2020 Last Reviewed on: December 23, 2019 PayNearMe, Inc. (the "Company") is providing you with this Notice ("Notice") to inform you about: the categories of Personal Information that the Company collects and maintains about applicants; and the purposes for which the Company uses that Personal Information. For purposes of this Notice, "Personal Information" means information that identifies, relates to, describes, is capable of being associated with, or could reasonably be linked, directly or indirectly with, a natural person that the Company may collect in connection with screening applicants for job openings at the Company. Identifiers and Professional or Employment-Related Information. The Company collects identifiers and professional or employment-related information, which may include some or all the following: real name, nickname or alias, postal address, telephone number, e-mail address, membership in professional organizations, professional certifications, language skills, and current and past employment history. The Company collects this Personal Information to evaluate previous job performance and consider applicants for positions, to develop a talent pool and plan for succession . click apply for full job details
Senior Site Reliability Engineer- San Francisco, CA, the US
Kody San Francisco, California
Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Competitive packages aligned with California market standards Lead a dynamic and innovative team in a very rapidly growing company Collaborative, inclusive environment where your contributions are recognized and valued
08/05/2026
Full time
Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Responsibilities Participate in a follow-the-sun production on-call rotation as a primary incident responder. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure. Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes. Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews. Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR. Requirements 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems. Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization. Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements. Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence. Leadership & Operational Excellence Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact. Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure. Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements. Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil. Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events. Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability. Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design. Benefits Competitive packages aligned with California market standards Lead a dynamic and innovative team in a very rapidly growing company Collaborative, inclusive environment where your contributions are recognized and valued
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible)
Capital One Mc Lean, Virginia
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. In this role, you will: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. You will contribute to the north star platform architecture, continuously publishing and refining living diagrams and canonical APIs that cover agent orchestration, RAG pipelines, prompt libraries and multi-tenant policy enforcement. A major emphasis is around standardizing and automating agentic workflows : you will evaluate agentic frameworks such LangGraph, AutoGen, Semantic Kernal, CrewAI and LlamaIndex and then harden / blend patterns that best meet enterprise SLAs do that 90% of new apps adopt them. Developer experience is another cornerstone. You will contribute to crafting an end to end GenAI SDK, CLI and starter kits that let AI engineers spin up secure, observable agentic workflows in under minutes, shrinking prototyping to production timelines by 30%. Trust and safety remain paramount; you will help bring together a vision of central guardrail services - prompt firewalls, content-filter hooks, red team harnesses and audit APIs - consumed by every application to ensure zero Sev4 incidents. You will collaborate with cross organization architects to drive end to end performance by optimizing orchestration - level batching, retrieval caching, heuristic tuning to achieve reductions in per token spend. You will accelerate innovation by incubating proof of concepts and driving RFCs such as hierarchical agent memory, multimodal guardrails, multimodal RAG. You'll own central Helm charts, operators and CRDs that auto scale agents to hit tenant SLAs Finally you will coach and evangelize - hosting architecture office hours, mentoring Staff, Principal and Senior engineers, authoring technical design documents and blogs and representing Capital One at Tier1 AI conferences - to amplify platform vision across internal and external communities. The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good. Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production. You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven. You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enable you to see and exploit optimization opportunities that others miss. You are a resilient trail blazer who can forge new paths to achieve business goals when the route is unknown. Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, Engineering, or AI plus at least 10 years of experience developing AI and ML algorithms or technologies, or Master's degree plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications: 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) 2+ years of experience supporting Agentic Frameworks (LangChain, CrewAI, Semantic Kernel (Microsoft), or AutoGen) 2+ years of experience with LLMOps (Google Cloud Vertex AI, Amazon SageMaker, Azure Machine Learning) 8+ years of experience designing mission-critical machine learning platforms 2+ years of experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the VP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang Master's degree in Computer Science, Computer Engineering, or relevant technical field Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Experience leading GenAI or LLM-Powered application architectures in production Deep understanding of Responsible AI, data privacy and multi-tenant security patterns K8s mastery (multi-region clusters, service mesh) Experience staying abreast of the latest AI research and AI systems and applying novel techniques in production Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation . click apply for full job details
08/05/2026
Full time
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. In this role, you will: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. You will contribute to the north star platform architecture, continuously publishing and refining living diagrams and canonical APIs that cover agent orchestration, RAG pipelines, prompt libraries and multi-tenant policy enforcement. A major emphasis is around standardizing and automating agentic workflows : you will evaluate agentic frameworks such LangGraph, AutoGen, Semantic Kernal, CrewAI and LlamaIndex and then harden / blend patterns that best meet enterprise SLAs do that 90% of new apps adopt them. Developer experience is another cornerstone. You will contribute to crafting an end to end GenAI SDK, CLI and starter kits that let AI engineers spin up secure, observable agentic workflows in under minutes, shrinking prototyping to production timelines by 30%. Trust and safety remain paramount; you will help bring together a vision of central guardrail services - prompt firewalls, content-filter hooks, red team harnesses and audit APIs - consumed by every application to ensure zero Sev4 incidents. You will collaborate with cross organization architects to drive end to end performance by optimizing orchestration - level batching, retrieval caching, heuristic tuning to achieve reductions in per token spend. You will accelerate innovation by incubating proof of concepts and driving RFCs such as hierarchical agent memory, multimodal guardrails, multimodal RAG. You'll own central Helm charts, operators and CRDs that auto scale agents to hit tenant SLAs Finally you will coach and evangelize - hosting architecture office hours, mentoring Staff, Principal and Senior engineers, authoring technical design documents and blogs and representing Capital One at Tier1 AI conferences - to amplify platform vision across internal and external communities. The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good. Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production. You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven. You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enable you to see and exploit optimization opportunities that others miss. You are a resilient trail blazer who can forge new paths to achieve business goals when the route is unknown. Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, Engineering, or AI plus at least 10 years of experience developing AI and ML algorithms or technologies, or Master's degree plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications: 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) 2+ years of experience supporting Agentic Frameworks (LangChain, CrewAI, Semantic Kernel (Microsoft), or AutoGen) 2+ years of experience with LLMOps (Google Cloud Vertex AI, Amazon SageMaker, Azure Machine Learning) 8+ years of experience designing mission-critical machine learning platforms 2+ years of experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the VP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang Master's degree in Computer Science, Computer Engineering, or relevant technical field Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Experience leading GenAI or LLM-Powered application architectures in production Deep understanding of Responsible AI, data privacy and multi-tenant security patterns K8s mastery (multi-region clusters, service mesh) Experience staying abreast of the latest AI research and AI systems and applying novel techniques in production Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation . click apply for full job details
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible)
Capital One New York, New York
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. In this role, you will: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. You will contribute to the north star platform architecture, continuously publishing and refining living diagrams and canonical APIs that cover agent orchestration, RAG pipelines, prompt libraries and multi-tenant policy enforcement. A major emphasis is around standardizing and automating agentic workflows : you will evaluate agentic frameworks such LangGraph, AutoGen, Semantic Kernal, CrewAI and LlamaIndex and then harden / blend patterns that best meet enterprise SLAs do that 90% of new apps adopt them. Developer experience is another cornerstone. You will contribute to crafting an end to end GenAI SDK, CLI and starter kits that let AI engineers spin up secure, observable agentic workflows in under minutes, shrinking prototyping to production timelines by 30%. Trust and safety remain paramount; you will help bring together a vision of central guardrail services - prompt firewalls, content-filter hooks, red team harnesses and audit APIs - consumed by every application to ensure zero Sev4 incidents. You will collaborate with cross organization architects to drive end to end performance by optimizing orchestration - level batching, retrieval caching, heuristic tuning to achieve reductions in per token spend. You will accelerate innovation by incubating proof of concepts and driving RFCs such as hierarchical agent memory, multimodal guardrails, multimodal RAG. You'll own central Helm charts, operators and CRDs that auto scale agents to hit tenant SLAs Finally you will coach and evangelize - hosting architecture office hours, mentoring Staff, Principal and Senior engineers, authoring technical design documents and blogs and representing Capital One at Tier1 AI conferences - to amplify platform vision across internal and external communities. The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good. Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production. You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven. You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enable you to see and exploit optimization opportunities that others miss. You are a resilient trail blazer who can forge new paths to achieve business goals when the route is unknown. Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, Engineering, or AI plus at least 10 years of experience developing AI and ML algorithms or technologies, or Master's degree plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications: 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) 2+ years of experience supporting Agentic Frameworks (LangChain, CrewAI, Semantic Kernel (Microsoft), or AutoGen) 2+ years of experience with LLMOps (Google Cloud Vertex AI, Amazon SageMaker, Azure Machine Learning) 8+ years of experience designing mission-critical machine learning platforms 2+ years of experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the VP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang Master's degree in Computer Science, Computer Engineering, or relevant technical field Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Experience leading GenAI or LLM-Powered application architectures in production Deep understanding of Responsible AI, data privacy and multi-tenant security patterns K8s mastery (multi-region clusters, service mesh) Experience staying abreast of the latest AI research and AI systems and applying novel techniques in production Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation . click apply for full job details
08/05/2026
Full time
Senior. Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. In this role, you will: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. You will contribute to the north star platform architecture, continuously publishing and refining living diagrams and canonical APIs that cover agent orchestration, RAG pipelines, prompt libraries and multi-tenant policy enforcement. A major emphasis is around standardizing and automating agentic workflows : you will evaluate agentic frameworks such LangGraph, AutoGen, Semantic Kernal, CrewAI and LlamaIndex and then harden / blend patterns that best meet enterprise SLAs do that 90% of new apps adopt them. Developer experience is another cornerstone. You will contribute to crafting an end to end GenAI SDK, CLI and starter kits that let AI engineers spin up secure, observable agentic workflows in under minutes, shrinking prototyping to production timelines by 30%. Trust and safety remain paramount; you will help bring together a vision of central guardrail services - prompt firewalls, content-filter hooks, red team harnesses and audit APIs - consumed by every application to ensure zero Sev4 incidents. You will collaborate with cross organization architects to drive end to end performance by optimizing orchestration - level batching, retrieval caching, heuristic tuning to achieve reductions in per token spend. You will accelerate innovation by incubating proof of concepts and driving RFCs such as hierarchical agent memory, multimodal guardrails, multimodal RAG. You'll own central Helm charts, operators and CRDs that auto scale agents to hit tenant SLAs Finally you will coach and evangelize - hosting architecture office hours, mentoring Staff, Principal and Senior engineers, authoring technical design documents and blogs and representing Capital One at Tier1 AI conferences - to amplify platform vision across internal and external communities. The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good. Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production. You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven. You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enable you to see and exploit optimization opportunities that others miss. You are a resilient trail blazer who can forge new paths to achieve business goals when the route is unknown. Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, Engineering, or AI plus at least 10 years of experience developing AI and ML algorithms or technologies, or Master's degree plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications: 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) 2+ years of experience supporting Agentic Frameworks (LangChain, CrewAI, Semantic Kernel (Microsoft), or AutoGen) 2+ years of experience with LLMOps (Google Cloud Vertex AI, Amazon SageMaker, Azure Machine Learning) 8+ years of experience designing mission-critical machine learning platforms 2+ years of experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the VP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang Master's degree in Computer Science, Computer Engineering, or relevant technical field Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Experience leading GenAI or LLM-Powered application architectures in production Deep understanding of Responsible AI, data privacy and multi-tenant security patterns K8s mastery (multi-region clusters, service mesh) Experience staying abreast of the latest AI research and AI systems and applying novel techniques in production Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation . click apply for full job details
Senior Platform Engineer
OnHires San Francisco, California
Job Description Job Description About Us: We are building a robust, scalable trading platform to serve high-traffic, latency-sensitive applications. Our infrastructure leverages state-of-the-art technologies to support real-time trading while providing unparalleled reliability and performance. Join us to shape the future of our platform and engineering culture. Job Summary: We are looking for a Senior DevOps & Platform Engineer to lead the design, implementation, and management of our AWS-centric infrastructure. You will play a pivotal role in maximizing the velocity of our product engineering team, ensuring platform scalability, reliability, and security. This is a high-impact role, combining elements of DevOps, Platform Engineering, and Site Reliability Engineering (SRE). You will champion best practices, shape the engineering culture, and ensure our platform is robust, efficient, and ready for the future. Key Responsibilities: Platform Engineering Infrastructure Design: Architect and implement scalable infrastructure to support the deployment and management of our trading platform. Developer Tooling: Build and maintain internal tools to streamline developer workflows, including advanced CI/CD pipelines. Infrastructure as Code (IaC): Champion IaC practices using Terraform, CloudFormation, or Pulumi. Core Services Management: Manage and optimize platform-critical services such as: NATS Cluster RabbitMQ AWS RDS PostgreSQL Redis Cluster DevOps Automation and CI/CD: Automate and optimize deployment processes to ensure seamless continuous integration and delivery. Container Orchestration: Manage and scale containerized workloads using Kubernetes and Docker. Cloud Optimization: Monitor and optimize cloud resource usage for performance and cost efficiency. Site Reliability Engineering (SRE) Reliability Metrics: Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Monitoring & Observability: Implement observability tools and dashboards (e.g., Prometheus, Datadog, Grafana) for real-time system monitoring. Incident Management: Lead incident response efforts, conduct root cause analysis, and implement actionable postmortem reviews. Infrastructure Management AWS Expertise: Architect and manage cloud-based systems to handle high-traffic, latency-sensitive applications. Disaster Recovery: Implement robust disaster recovery and business continuity strategies, including backups and multi-region failover. Security Practices: Collaborate with security teams to enforce best practices for IAM, encryption, and compliance. Collaboration & Leadership Cross-Team Collaboration: Partner with software engineers to design infrastructure solutions tailored to their application needs. Culture Building: Help shape the engineering culture, promoting a philosophy of security, velocity, and reliability. Mentorship: Mentor junior engineers and document best practices to drive knowledge sharing and operational excellence. Long-Term Tech Evolution Backend Transition: Contribute to evolving our backend microservices (currently NodeJS, with some Python and C#) towards Go and Rust. Third-Party Integration: Evaluate and integrate critical third-party software and infrastructure, such as payment gateways and mobility stacks. Your Impact: Simplify infrastructure concerns for product teams to accelerate builds, deployments, and scaling. Advocate for modern practices like Zero Trust Networking and continuously improve platform architecture. Balance the demands of product velocity with a well-managed, secure, and scalable platform. Required Skills & Experience: Technical Expertise Cloud Experience: 5-8+ years of hands-on experience with cloud platforms, particularly AWS, including services like EC2, RDS, S3, Lambda, and VPC. Containerization: Proficiency with Docker and Kubernetes (EKS) or ECS. Infrastructure as Code (IaC): Strong experience with Terraform, CloudFormation, or Pulumi. Programming Skills: Proficiency in at least one programming language (e.g., Python, Go, TypeScript/JavaScript, Ruby, Java). DevOps & SRE CI/CD Pipelines: Expertise in building and maintaining CI/CD workflows using tools like GitLab CI, Jenkins, or GitHub Actions. Monitoring Tools: Experience with observability platforms (e.g., Prometheus, Datadog, Grafana). Incident Management: Proven ability to handle incident response, root cause analysis, and postmortem reviews. Soft Skills Problem-Solving: Ability to research, design, and deliver solutions to complex infrastructure challenges. Collaboration: Experience working directly with product engineers to improve workflows incrementally. Leadership: Ownership mindset with the ability to mentor team members and advocate for best practices. Preferred Skills (Nice-to-Have): Familiarity with backend languages like Go or Rust. AWS certifications (e.g., Solutions Architect, DevOps Engineer). Experience with networking concepts (e.g., load balancers, DNS, VPNs) and traffic optimization. Knowledge of emerging CNCF technologies and CI/CD trends. What We Offer: Competitive salary with future equity options Opportunities to work with cutting-edge technologies and evolve our platform. Flexible working hours and a remote-friendly environment. Professional growth through certifications, conferences, and internal training. Collaborative culture focused on innovation and operational excellence.
08/05/2026
Full time
Job Description Job Description About Us: We are building a robust, scalable trading platform to serve high-traffic, latency-sensitive applications. Our infrastructure leverages state-of-the-art technologies to support real-time trading while providing unparalleled reliability and performance. Join us to shape the future of our platform and engineering culture. Job Summary: We are looking for a Senior DevOps & Platform Engineer to lead the design, implementation, and management of our AWS-centric infrastructure. You will play a pivotal role in maximizing the velocity of our product engineering team, ensuring platform scalability, reliability, and security. This is a high-impact role, combining elements of DevOps, Platform Engineering, and Site Reliability Engineering (SRE). You will champion best practices, shape the engineering culture, and ensure our platform is robust, efficient, and ready for the future. Key Responsibilities: Platform Engineering Infrastructure Design: Architect and implement scalable infrastructure to support the deployment and management of our trading platform. Developer Tooling: Build and maintain internal tools to streamline developer workflows, including advanced CI/CD pipelines. Infrastructure as Code (IaC): Champion IaC practices using Terraform, CloudFormation, or Pulumi. Core Services Management: Manage and optimize platform-critical services such as: NATS Cluster RabbitMQ AWS RDS PostgreSQL Redis Cluster DevOps Automation and CI/CD: Automate and optimize deployment processes to ensure seamless continuous integration and delivery. Container Orchestration: Manage and scale containerized workloads using Kubernetes and Docker. Cloud Optimization: Monitor and optimize cloud resource usage for performance and cost efficiency. Site Reliability Engineering (SRE) Reliability Metrics: Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Monitoring & Observability: Implement observability tools and dashboards (e.g., Prometheus, Datadog, Grafana) for real-time system monitoring. Incident Management: Lead incident response efforts, conduct root cause analysis, and implement actionable postmortem reviews. Infrastructure Management AWS Expertise: Architect and manage cloud-based systems to handle high-traffic, latency-sensitive applications. Disaster Recovery: Implement robust disaster recovery and business continuity strategies, including backups and multi-region failover. Security Practices: Collaborate with security teams to enforce best practices for IAM, encryption, and compliance. Collaboration & Leadership Cross-Team Collaboration: Partner with software engineers to design infrastructure solutions tailored to their application needs. Culture Building: Help shape the engineering culture, promoting a philosophy of security, velocity, and reliability. Mentorship: Mentor junior engineers and document best practices to drive knowledge sharing and operational excellence. Long-Term Tech Evolution Backend Transition: Contribute to evolving our backend microservices (currently NodeJS, with some Python and C#) towards Go and Rust. Third-Party Integration: Evaluate and integrate critical third-party software and infrastructure, such as payment gateways and mobility stacks. Your Impact: Simplify infrastructure concerns for product teams to accelerate builds, deployments, and scaling. Advocate for modern practices like Zero Trust Networking and continuously improve platform architecture. Balance the demands of product velocity with a well-managed, secure, and scalable platform. Required Skills & Experience: Technical Expertise Cloud Experience: 5-8+ years of hands-on experience with cloud platforms, particularly AWS, including services like EC2, RDS, S3, Lambda, and VPC. Containerization: Proficiency with Docker and Kubernetes (EKS) or ECS. Infrastructure as Code (IaC): Strong experience with Terraform, CloudFormation, or Pulumi. Programming Skills: Proficiency in at least one programming language (e.g., Python, Go, TypeScript/JavaScript, Ruby, Java). DevOps & SRE CI/CD Pipelines: Expertise in building and maintaining CI/CD workflows using tools like GitLab CI, Jenkins, or GitHub Actions. Monitoring Tools: Experience with observability platforms (e.g., Prometheus, Datadog, Grafana). Incident Management: Proven ability to handle incident response, root cause analysis, and postmortem reviews. Soft Skills Problem-Solving: Ability to research, design, and deliver solutions to complex infrastructure challenges. Collaboration: Experience working directly with product engineers to improve workflows incrementally. Leadership: Ownership mindset with the ability to mentor team members and advocate for best practices. Preferred Skills (Nice-to-Have): Familiarity with backend languages like Go or Rust. AWS certifications (e.g., Solutions Architect, DevOps Engineer). Experience with networking concepts (e.g., load balancers, DNS, VPNs) and traffic optimization. Knowledge of emerging CNCF technologies and CI/CD trends. What We Offer: Competitive salary with future equity options Opportunities to work with cutting-edge technologies and evolve our platform. Flexible working hours and a remote-friendly environment. Professional growth through certifications, conferences, and internal training. Collaborative culture focused on innovation and operational excellence.
Principal Site Reliability Engineer, Google Cloud
Saviynt Milpitas, California
Job Description Job Description Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit . Why This Role Matters Saviynt's platform is mission-critical for our customers. As we scale globally, reliability, availability, and performance are not optional-they are core product features. As a Principal Engineer, you will define and drive the reliability strategy for our SaaS platform. This is a high-impact, hands-on engineering role with broad influence across infrastructure, platform, and application teams. You will shape how Saviynt designs, operates, and measures reliability at scale. This role is ideal for engineers who want to work on hard reliability problems, influence architecture across teams, and leave a lasting mark on a growing SaaS platform. What You Will Do In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on • You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment • Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers • Develop robust, internal-facing tools and automation for infrastructure provisioning and management primarily using Go (Golang) • Architect and optimize foundational solutions within Cloud environments (AWS, Azure, etc.), focusing on creating reusable patterns and modules for other teams • Design and implement shared Event-Driven Architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub that product teams can easily utilize • Develop and maintain robust CI/CD pipelines (e.g., GitLab CI and ArgoCD) as a service, providing standardized and automated deployment workflows for various development teams • Design and build resilient Distributed Systems components that serve as building blocks for other applications, focusing on reliability, fault tolerance, and performance • Manage and optimize our shared infrastructure across Multi-Region Cloud Environments, ensuring that platform services are globally available and performant for all consumers • Establish and enhance centralized Observability and Monitoring platforms and tools that provide self-service insights for consuming teams • Define and implement clear, well-documented RESTful API designs for the infrastructure services you build, ensuring ease of integration for internal clients • Implement and manage Service Mesh (e.g., Envoy, Istio) capabilities, providing traffic management, security, and policy enforcement as a shared platform for services • Design, implement, and optimize highly available Relational Database services or shared data platforms for broad organizational use • Collaborate closely with product development teams to understand their infrastructure needs and pain points, providing technical guidance and support • Participate in on-call rotations to support the critical shared infrastructure you build What Are We Looking For • 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers • Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures) • Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation • Extensive hands-on experience with at least one major Cloud Provider (GCP is a must); multi-cloud experience is a strong plus, especially in building abstractions over them. • Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services • Solid understanding and practical experience with CI/CD pipeline tools (especially GitLab CI) and experience establishing automated delivery processes for other teams • Demonstrable experience designing and operating Distributed Systems, with an understanding of patterns for creating reliable, shared components • Familiarity with Multi-Region Cloud Environments and strategies for building globally distributed and highly available platform • Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure • Strong experience with RESTful API design principles and building well-documented, consumable APIs • Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context • Hands-on experience with Relational Databases (e.g., MySQL, PostgresSQL), ideally in managing them as a service • Excellent communication skills and the ability to clearly articulate complex technical concepts to both technical and non-technical audiences • A strong customer-centric mindset, treating internal development teams as your primary customers • Advanced Professional GCP Certification is required. • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience required If required for this role, you will: - Complete security & privacy literacy and awareness training during onboarding and annually thereafter - Review (initially and annually thereafter), understand, and adhere to Information Security/Privacy Policies and Procedures such as (but not limited to): Incident Response Policy/Procedures > Business Continuity/Disaster Recovery Policy/Procedures > Mobile Device Policy > Account Management Policy > Access Control Policy > Personnel Security Policy > Privacy Policy Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us! Saviynt is an equal opportunity employer and we welcome everyone to our team. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
08/05/2026
Full time
Job Description Job Description Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit . Why This Role Matters Saviynt's platform is mission-critical for our customers. As we scale globally, reliability, availability, and performance are not optional-they are core product features. As a Principal Engineer, you will define and drive the reliability strategy for our SaaS platform. This is a high-impact, hands-on engineering role with broad influence across infrastructure, platform, and application teams. You will shape how Saviynt designs, operates, and measures reliability at scale. This role is ideal for engineers who want to work on hard reliability problems, influence architecture across teams, and leave a lasting mark on a growing SaaS platform. What You Will Do In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on • You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment • Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers • Develop robust, internal-facing tools and automation for infrastructure provisioning and management primarily using Go (Golang) • Architect and optimize foundational solutions within Cloud environments (AWS, Azure, etc.), focusing on creating reusable patterns and modules for other teams • Design and implement shared Event-Driven Architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub that product teams can easily utilize • Develop and maintain robust CI/CD pipelines (e.g., GitLab CI and ArgoCD) as a service, providing standardized and automated deployment workflows for various development teams • Design and build resilient Distributed Systems components that serve as building blocks for other applications, focusing on reliability, fault tolerance, and performance • Manage and optimize our shared infrastructure across Multi-Region Cloud Environments, ensuring that platform services are globally available and performant for all consumers • Establish and enhance centralized Observability and Monitoring platforms and tools that provide self-service insights for consuming teams • Define and implement clear, well-documented RESTful API designs for the infrastructure services you build, ensuring ease of integration for internal clients • Implement and manage Service Mesh (e.g., Envoy, Istio) capabilities, providing traffic management, security, and policy enforcement as a shared platform for services • Design, implement, and optimize highly available Relational Database services or shared data platforms for broad organizational use • Collaborate closely with product development teams to understand their infrastructure needs and pain points, providing technical guidance and support • Participate in on-call rotations to support the critical shared infrastructure you build What Are We Looking For • 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers • Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures) • Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation • Extensive hands-on experience with at least one major Cloud Provider (GCP is a must); multi-cloud experience is a strong plus, especially in building abstractions over them. • Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services • Solid understanding and practical experience with CI/CD pipeline tools (especially GitLab CI) and experience establishing automated delivery processes for other teams • Demonstrable experience designing and operating Distributed Systems, with an understanding of patterns for creating reliable, shared components • Familiarity with Multi-Region Cloud Environments and strategies for building globally distributed and highly available platform • Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure • Strong experience with RESTful API design principles and building well-documented, consumable APIs • Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context • Hands-on experience with Relational Databases (e.g., MySQL, PostgresSQL), ideally in managing them as a service • Excellent communication skills and the ability to clearly articulate complex technical concepts to both technical and non-technical audiences • A strong customer-centric mindset, treating internal development teams as your primary customers • Advanced Professional GCP Certification is required. • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience required If required for this role, you will: - Complete security & privacy literacy and awareness training during onboarding and annually thereafter - Review (initially and annually thereafter), understand, and adhere to Information Security/Privacy Policies and Procedures such as (but not limited to): Incident Response Policy/Procedures > Business Continuity/Disaster Recovery Policy/Procedures > Mobile Device Policy > Account Management Policy > Access Control Policy > Personnel Security Policy > Privacy Policy Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us! Saviynt is an equal opportunity employer and we welcome everyone to our team. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Senior Site Reliability Engineer - Compute Platforms
Five9 San Ramon, California
Job Description Job Description Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-first culture and enables us to innovate, grow, and thrive while enjoying the journey together. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer - Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment. This role is responsible for the architecture, design, and standardization of enterprise compute and hypervisor environments spanning bare metal infrastructure, operating systems, hypervisors, private cloud orchestration, and Kubernetes using Infrastructure-as-Code and GitOps practices. This is a deeply technical role requiring expert-level understanding of compute hardware management, Kubernetes, OpenStack, hypervisors and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multi-tenant-isolated services that serve high-throughput, mission-critical applications. Key Responsibilities Lead the architecture and design of enterprise compute and hypervisor platform solutions across hardware, OS, virtualization, cloud orchestration, and container orchestration layers Define standards and automation frameworks for bare metal provisioning and lifecycle management Design and implement Bare Metal as a Service (BMaaS) capabilities for scalable infrastructure consumption Architect and design Kubernetes platforms on bare metal with QoS and Affinity (ArgoCD) Architect and validate automated deployments of operating systems and hypervisors including Ubuntu and Harvester Design and maintain PXE-based provisioning environments leveraging Redfish APIs for large-scale server deployments Develop Infrastructure-as-Code using Ansible, Terraform, Helm and Git, with Python/Bash automation. Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback. Design automated workflows for server build, firmware lifecycle management, patching, and hardware validation Evaluate and standardize enterprise hardware platforms to meet performance, scalability, and reliability requirements Produce detailed high-level and low-level design documentation , build guides, and operational handoff materials Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems Partner with operations, network, storage, and platform teams to ensure designs are supportable and production-ready Participate in on-call escalation support for complex platform-related issues Collaborate globally on change management , documentation, and operational best practices Minimum Qualifications 6 + years of experience in infrastructure engineering, platform engineering, or DevOps with a strong focus on Compute system design Proven experience designing and automating bare metal compute environments at scale Strong hands-on experience with PXE boot, network-based OS provisioning, and automated server imaging Experience implementing or supporting Bare Metal as a Service (BMaaS) platforms Practical experience using Redfish APIs for hardware provisioning, power management, and remote lifecycle operations Deep expertise with Ubuntu Linux in enterprise environments Strong Hands-on experience with KVM hypervisors (Suse Harvester, OpenStack). Experience designing and deploying production-grade Kubernetes clusters Strong background with enterprise compute hardware platforms , including Cisco UCS, Dell PowerEdge, Supermicro systems & HPE Proficiency with Infrastructure as Code tools (e.g., Terraform, Ansible, or similar) Experience building or supporting CI/CD pipelines for infrastructure and platform automation Strong scripting skills in Python, Bash, or similar languages Demonstrated ability to produce clear, structured technical design documentation Excellent written and verbal communication skills Bachelor's degree in computer science or equivalent professional experience Preferred Qualifications OpenStack, Ubuntu KVM administration. BareMetal as a Service (PXE, Redfish). Kubernetes on BareMetal CIS/NIST security and infrastructure lifecycle management. ITIL Foundation/advanced certifications in support of ITSM standard methodology. Background in telco, edge cloud, or large enterprise environments. Ubuntu Certifications, CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS) Master's degree in computer science, IT, Engineering, or a related field preferred; equivalent experience and relevant industry certifications will also be considered What You'll Get A collaborative team that's deeply invested in infrastructure excellence. Complex technical challenges that require creative, scalable solutions. The opportunity to shape a next-generation private cloud platform-built reliability Access to the latest tools, frameworks, and upstream project developments Skills and Attributes: Analytical Thinking & Problem Solving: Demonstrated ability to translate complex, cross-domain requirements into scalable and resilient cloud infrastructure and automation solutions Collaboration & Teamwork: Strong interpersonal and communication skills with a proven track record of effective collaboration across multidisciplinary teams, including developers, operations, security, and product stakeholders Mentorship & Leadership: Passionate about knowledge-sharing and mentorship, with experience guiding junior engineers and fostering a team culture of continuous learning, innovation, and technical excellence in cloud engineering and DevOps practices Work Location: This role is fully remote for candidates who reside outside the 30 mile radius of one of our offices. For candidates who reside within a 30 mile radius of one of our offices, this role is Hybrid and would require 3 days a week (T, W, TH) in office. As part of our continued commitment to diversity, equity, and inclusion, Five9 supports pay transparency during the entire recruitment process. Actual compensation packages are based on several factors that are unique to each candidate including, but not limited to: skill set, depth of experience, certifications, and specific work location. The range displayed reflects the minimum and maximum target for new hire salaries for the job across the United States. Your recruiter can share more about the specific compensation package during your hiring process. Additionally, the total compensation package for this position may also include an annual performance bonus, stock, and/or other applicable incentive compensation plans. Our total reward package also includes: Health, dental, and vision coverage, beginning on the first day of employment. Five9 covers 100% of the employee portion of the health, dental and vision coverage and shares a high portion of the dependent cost. We also offer Short & Long-Term Disability, Basic Life Insurance, and a 401k saving plan with employer matching. Access to an innovative mental health support platform that offers personalized care and resources in areas such as: therapy, coaching and self-guided mindfulness exercises for all covered employees and their covered dependents. Generous employee stock purchase plan. Paid Time Off, Company paid holidays, paid volunteer hours and 12 weeks paid parental leave. All compensation and benefits are subject to the requirements and restrictions set forth in the applicable plan documents and any written agreements between the parties. The US base salary range for this role is below. $82,300-$228,800 USD Five9 embraces diversity and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better we are. Five9 is an equal opportunity employer. View our privacy policy, including our privacy notice to California residents here: -pt/legal. Note: Five9 will never request that an applicant send money as a prerequisite for commencing employment with Five9.
08/05/2026
Full time
Job Description Job Description Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-first culture and enables us to innovate, grow, and thrive while enjoying the journey together. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer - Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment. This role is responsible for the architecture, design, and standardization of enterprise compute and hypervisor environments spanning bare metal infrastructure, operating systems, hypervisors, private cloud orchestration, and Kubernetes using Infrastructure-as-Code and GitOps practices. This is a deeply technical role requiring expert-level understanding of compute hardware management, Kubernetes, OpenStack, hypervisors and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multi-tenant-isolated services that serve high-throughput, mission-critical applications. Key Responsibilities Lead the architecture and design of enterprise compute and hypervisor platform solutions across hardware, OS, virtualization, cloud orchestration, and container orchestration layers Define standards and automation frameworks for bare metal provisioning and lifecycle management Design and implement Bare Metal as a Service (BMaaS) capabilities for scalable infrastructure consumption Architect and design Kubernetes platforms on bare metal with QoS and Affinity (ArgoCD) Architect and validate automated deployments of operating systems and hypervisors including Ubuntu and Harvester Design and maintain PXE-based provisioning environments leveraging Redfish APIs for large-scale server deployments Develop Infrastructure-as-Code using Ansible, Terraform, Helm and Git, with Python/Bash automation. Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback. Design automated workflows for server build, firmware lifecycle management, patching, and hardware validation Evaluate and standardize enterprise hardware platforms to meet performance, scalability, and reliability requirements Produce detailed high-level and low-level design documentation , build guides, and operational handoff materials Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems Partner with operations, network, storage, and platform teams to ensure designs are supportable and production-ready Participate in on-call escalation support for complex platform-related issues Collaborate globally on change management , documentation, and operational best practices Minimum Qualifications 6 + years of experience in infrastructure engineering, platform engineering, or DevOps with a strong focus on Compute system design Proven experience designing and automating bare metal compute environments at scale Strong hands-on experience with PXE boot, network-based OS provisioning, and automated server imaging Experience implementing or supporting Bare Metal as a Service (BMaaS) platforms Practical experience using Redfish APIs for hardware provisioning, power management, and remote lifecycle operations Deep expertise with Ubuntu Linux in enterprise environments Strong Hands-on experience with KVM hypervisors (Suse Harvester, OpenStack). Experience designing and deploying production-grade Kubernetes clusters Strong background with enterprise compute hardware platforms , including Cisco UCS, Dell PowerEdge, Supermicro systems & HPE Proficiency with Infrastructure as Code tools (e.g., Terraform, Ansible, or similar) Experience building or supporting CI/CD pipelines for infrastructure and platform automation Strong scripting skills in Python, Bash, or similar languages Demonstrated ability to produce clear, structured technical design documentation Excellent written and verbal communication skills Bachelor's degree in computer science or equivalent professional experience Preferred Qualifications OpenStack, Ubuntu KVM administration. BareMetal as a Service (PXE, Redfish). Kubernetes on BareMetal CIS/NIST security and infrastructure lifecycle management. ITIL Foundation/advanced certifications in support of ITSM standard methodology. Background in telco, edge cloud, or large enterprise environments. Ubuntu Certifications, CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS) Master's degree in computer science, IT, Engineering, or a related field preferred; equivalent experience and relevant industry certifications will also be considered What You'll Get A collaborative team that's deeply invested in infrastructure excellence. Complex technical challenges that require creative, scalable solutions. The opportunity to shape a next-generation private cloud platform-built reliability Access to the latest tools, frameworks, and upstream project developments Skills and Attributes: Analytical Thinking & Problem Solving: Demonstrated ability to translate complex, cross-domain requirements into scalable and resilient cloud infrastructure and automation solutions Collaboration & Teamwork: Strong interpersonal and communication skills with a proven track record of effective collaboration across multidisciplinary teams, including developers, operations, security, and product stakeholders Mentorship & Leadership: Passionate about knowledge-sharing and mentorship, with experience guiding junior engineers and fostering a team culture of continuous learning, innovation, and technical excellence in cloud engineering and DevOps practices Work Location: This role is fully remote for candidates who reside outside the 30 mile radius of one of our offices. For candidates who reside within a 30 mile radius of one of our offices, this role is Hybrid and would require 3 days a week (T, W, TH) in office. As part of our continued commitment to diversity, equity, and inclusion, Five9 supports pay transparency during the entire recruitment process. Actual compensation packages are based on several factors that are unique to each candidate including, but not limited to: skill set, depth of experience, certifications, and specific work location. The range displayed reflects the minimum and maximum target for new hire salaries for the job across the United States. Your recruiter can share more about the specific compensation package during your hiring process. Additionally, the total compensation package for this position may also include an annual performance bonus, stock, and/or other applicable incentive compensation plans. Our total reward package also includes: Health, dental, and vision coverage, beginning on the first day of employment. Five9 covers 100% of the employee portion of the health, dental and vision coverage and shares a high portion of the dependent cost. We also offer Short & Long-Term Disability, Basic Life Insurance, and a 401k saving plan with employer matching. Access to an innovative mental health support platform that offers personalized care and resources in areas such as: therapy, coaching and self-guided mindfulness exercises for all covered employees and their covered dependents. Generous employee stock purchase plan. Paid Time Off, Company paid holidays, paid volunteer hours and 12 weeks paid parental leave. All compensation and benefits are subject to the requirements and restrictions set forth in the applicable plan documents and any written agreements between the parties. The US base salary range for this role is below. $82,300-$228,800 USD Five9 embraces diversity and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better we are. Five9 is an equal opportunity employer. View our privacy policy, including our privacy notice to California residents here: -pt/legal. Note: Five9 will never request that an applicant send money as a prerequisite for commencing employment with Five9.
Java SRE Engineer
Eitacies Inc Santa Clara, California
Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and Kubernetes platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill. Required Skill : AWS, AWS EKS, Kubernetes, DevOps / SRE, Java Key Responsibilities: Lead large-scale migrations of business-critical applications to AWS and Kubernetes (EKS) Design and operate production-grade AWS EKS platforms Implement GitOps-based deployment strategies using ArgoCD and Spinnaker Build and manage CI/CD pipelines and automated release strategies (blue/green, canary) Develop Python-based automation for infrastructure and operations Create and maintain Helm charts and deployment standards Troubleshoot and optimize Linux-based systems in production environments Support production systems including on-call, incident response, and RCA Collaborate with SRE and Security teams to ensure system reliability and scalability Drive architectural decisions and contribute to long-term platform strategy Mentor team members and improve engineering practices Qualifications: 10+ years of experience in Cloud / DevOps / SRE / Platform Engineering Strong hands-on experience with: AWS (EKS, EC2, VPC, IAM, ALB/NLB, CloudWatch, S3, RDS) Kubernetes, Linux systems, Python, ArgoCD (GitOps) Spinnaker, Helm Experience with Infrastructure as Code (Terraform or CloudFormation) Proven experience supporting production environments Experience leading or contributing to AWS migration projects Strong understanding of distributed systems and networking Preferred Qualifications: Experience with Akamai CDN and caching strategies Experience with Redis and Kafka Familiarity with observability tools (Prometheus, Grafana, Datadog, Splunk) Experience with service mesh (Istio, Linkerd) Knowledge of SRE practices (SLIs, SLOs, error budgets) Strong communication and documentation skills
08/05/2026
Full time
Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and Kubernetes platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill. Required Skill : AWS, AWS EKS, Kubernetes, DevOps / SRE, Java Key Responsibilities: Lead large-scale migrations of business-critical applications to AWS and Kubernetes (EKS) Design and operate production-grade AWS EKS platforms Implement GitOps-based deployment strategies using ArgoCD and Spinnaker Build and manage CI/CD pipelines and automated release strategies (blue/green, canary) Develop Python-based automation for infrastructure and operations Create and maintain Helm charts and deployment standards Troubleshoot and optimize Linux-based systems in production environments Support production systems including on-call, incident response, and RCA Collaborate with SRE and Security teams to ensure system reliability and scalability Drive architectural decisions and contribute to long-term platform strategy Mentor team members and improve engineering practices Qualifications: 10+ years of experience in Cloud / DevOps / SRE / Platform Engineering Strong hands-on experience with: AWS (EKS, EC2, VPC, IAM, ALB/NLB, CloudWatch, S3, RDS) Kubernetes, Linux systems, Python, ArgoCD (GitOps) Spinnaker, Helm Experience with Infrastructure as Code (Terraform or CloudFormation) Proven experience supporting production environments Experience leading or contributing to AWS migration projects Strong understanding of distributed systems and networking Preferred Qualifications: Experience with Akamai CDN and caching strategies Experience with Redis and Kafka Familiarity with observability tools (Prometheus, Grafana, Datadog, Splunk) Experience with service mesh (Istio, Linkerd) Knowledge of SRE practices (SLIs, SLOs, error budgets) Strong communication and documentation skills
Sr. SRE / DevOps Engineer - Sunnyvale, CA (Only Local candidate)
Donato Technologies Inc Sunnyvale, California
Job Description Job Description Greetings from Donato Technologies Inc. We have an immediate opening with my client. If you are looking for a new project, please send me a copy of your updated resumes Title: Sr. SRE / DevOps Engineer Location: Sunnyvale, CA (Only Local candidate) Client Interview - In-Person Job Summary - For this role, we are looking for a Sr. SRE / DevOps Engineer at Sunnyvale, California location. As Site Reliability Engineer, the individual will work closely with multi-functional teams, automate operations, optimize infrastructure, implement security and solve issues in an exciting, fast-paced environment. The individual will play a vital role in ensuring that the systems are reliable, scalable, and high performing. Responsibilities - • Ensure system reliability and availability - Monitor system issues, create strategies to detect issues, address those issues, design automated systems to troubleshoot, write and review post-mortems. • Mitigate Operational risks - Collaborate with development teams and other stakeholders to identify potential risks, perform risk assessments, implement risk mitigation strategies, continuously monitor and review the effectiveness of risk strategies. • Monitor system health. • Minimize emergency response (MTTR). • Maintain CI/CD pipelines, etc. • Continuous improvement by collaborating with various teams. • Automation of processes. Must have/required experience and skills: • 8+ years of experience on DevOps and Site Reliability Engineering. • Hands-on with containerization and orchestration: Docker, Kubernetes/EKS. • Proficiency in infrastructure as code tools: Terraform, Ansible, or CloudFormation. • Experience setting up and managing services running on Kubernetes. • In-depth understanding of SRE principals including monitoring, alerting, error budgets, fault analysis, and automation. • In-depth knowledge of monitoring and observability tools: Apache Splunk • Knowledge of Linux operating system principles, networking fundamentals, and systems management • Demonstrable fluency in at least one of the following languages: Java or Python • Ability to identify and communicate technical and architectural problems, while working with partners and their team to iteratively find solutions. • Building and managing CI/CD pipeline - gatekeeping production deployments, develop and implement GIT branching strategies, branch protection rules, network policies, scale up/ scale down the load on AWS. • Strong problem-solving and analytical skills • Solve performance issues and scalability issues in the system. Technical Skills: • DevOps and SRE • AWS Kubernetes/EKS, Docker • Terraform, Ansible, or CloudFormation • Apache Splunk, Apache Flink • Programming/Scripting using Java or Python • CI/CD • Database - Vertica, Snowflake. Behavioral Skills: • Excellent Communication skills and collaboration skills • Ability to propose and implement improvements in the system • Ability to work with cross-functional stakeholders • Adaptability and a willingness to learn new technologies and techniques. • Proactive approach to issues, ability to provide prompt resolution/work Jennifer Sampson Technical Recruiter . DONATO TECHNOLOGIES, INC 12100 Ford Rd Dallas, TX 75234 Direct : (469)- Email: Web:
08/05/2026
Full time
Job Description Job Description Greetings from Donato Technologies Inc. We have an immediate opening with my client. If you are looking for a new project, please send me a copy of your updated resumes Title: Sr. SRE / DevOps Engineer Location: Sunnyvale, CA (Only Local candidate) Client Interview - In-Person Job Summary - For this role, we are looking for a Sr. SRE / DevOps Engineer at Sunnyvale, California location. As Site Reliability Engineer, the individual will work closely with multi-functional teams, automate operations, optimize infrastructure, implement security and solve issues in an exciting, fast-paced environment. The individual will play a vital role in ensuring that the systems are reliable, scalable, and high performing. Responsibilities - • Ensure system reliability and availability - Monitor system issues, create strategies to detect issues, address those issues, design automated systems to troubleshoot, write and review post-mortems. • Mitigate Operational risks - Collaborate with development teams and other stakeholders to identify potential risks, perform risk assessments, implement risk mitigation strategies, continuously monitor and review the effectiveness of risk strategies. • Monitor system health. • Minimize emergency response (MTTR). • Maintain CI/CD pipelines, etc. • Continuous improvement by collaborating with various teams. • Automation of processes. Must have/required experience and skills: • 8+ years of experience on DevOps and Site Reliability Engineering. • Hands-on with containerization and orchestration: Docker, Kubernetes/EKS. • Proficiency in infrastructure as code tools: Terraform, Ansible, or CloudFormation. • Experience setting up and managing services running on Kubernetes. • In-depth understanding of SRE principals including monitoring, alerting, error budgets, fault analysis, and automation. • In-depth knowledge of monitoring and observability tools: Apache Splunk • Knowledge of Linux operating system principles, networking fundamentals, and systems management • Demonstrable fluency in at least one of the following languages: Java or Python • Ability to identify and communicate technical and architectural problems, while working with partners and their team to iteratively find solutions. • Building and managing CI/CD pipeline - gatekeeping production deployments, develop and implement GIT branching strategies, branch protection rules, network policies, scale up/ scale down the load on AWS. • Strong problem-solving and analytical skills • Solve performance issues and scalability issues in the system. Technical Skills: • DevOps and SRE • AWS Kubernetes/EKS, Docker • Terraform, Ansible, or CloudFormation • Apache Splunk, Apache Flink • Programming/Scripting using Java or Python • CI/CD • Database - Vertica, Snowflake. Behavioral Skills: • Excellent Communication skills and collaboration skills • Ability to propose and implement improvements in the system • Ability to work with cross-functional stakeholders • Adaptability and a willingness to learn new technologies and techniques. • Proactive approach to issues, ability to provide prompt resolution/work Jennifer Sampson Technical Recruiter . DONATO TECHNOLOGIES, INC 12100 Ford Rd Dallas, TX 75234 Direct : (469)- Email: Web:
Site Reliability Engineer II
Restaurant365
Job Description Job Description Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering team members to produce top-notch results while elevating their skills. We're constantly evolving and improving to make sure we are and always will be "Best in Class" and we want that for you too! This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365's cloud infrastructure and applications. Qualified candidates will demonstrate growing expertise in site reliability practices, with skills in incident response, system monitoring, automation, and performance troubleshooting. You will collaborate with DevOps, development, and infrastructure teams to resolve moderately complex issues, propose improvements, and strengthen the reliability, scalability, and security of our SaaS platform. How you'll add value: Execution & Collaboration Respond to production incidents, perform triage and troubleshooting, and contribute to post-incident analysis. Identify and automate manual processes to improve efficiency and reduce risk. Enhance and evolve monitoring tools and platforms to improve observability. Promote and apply best practices for reliability, scalability, and performance across engineering. Implement and support cloud automation using Terraform, Ansible, or CloudFormation. Work within change management protocols to provide maximum uptime for production systems. Participate in on-call rotation, providing 24x7 support for incidents and contributing to root cause analysis. Partner with developers, architects, vendors, and IT teams to ensure reliable system operations. Research and remediate vulnerabilities in coordination with security teams. Maintain documentation of infrastructure, monitoring, runbooks, and incident response procedures. Standards & Process Apply company policies and procedures when handling operational tasks and incidents. Suggest and implement improvements to operational processes and monitoring practices. Contribute to technical diagrams, documentation, and runbooks for system reliability. Learning & Growth Expand expertise in cloud services (Azure, AWS, or GCP) and container platforms (EKS, ECS, AKS). Build proficiency with observability and monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios). Develop scripting and automation skills using Python, Bash, PowerShell, or similar. Participate in planning discussions by contributing technical input on system stability and reliability. What you'll need to be successful in this role: BS in Computer Science, Information Systems, or related field (or equivalent experience). 2-4 years of experience in site reliability engineering, DevOps, or cloud operations. Experience with cloud platforms (Azure or AWS), including services such as AKS, ECS/EKS, Functions/Lambda, S3, and Blob storage. Proficiency with infrastructure-as-code and automation (Terraform, Ansible, YAML, Python, Bash, PowerShell). Strong Linux engineering skills; working knowledge of Windows administration. Experience supporting production environments and participating in on-call rotations. Familiarity with web servers and middleware (Nginx, Apache Tomcat). Experience with CI/CD tools (GitLab, Git, or similar). Strong written, oral, and interpersonal communication skills. Preferred Qualifications Experience with monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios). Knowledge of performance analysis and system vulnerability remediation. Cloud certification (AWS or Azure) preferred. Familiarity with restaurant industry SaaS platforms and customer-facing applications. R365 Team Member Benefits & Compensation This position has a salary range of $98,583-$138,016 annually. The above range represents the expected salary range for this position. The actual salary may vary based upon several factors, including, but not limited to, relevant skills/experience, time in the role, business line, and geographic location. Restaurant365 focuses on equitable pay for our team and aims for transparency with our pay practices. Comprehensive medical benefits, 100% paid for employee 401k + matching Equity Option Grant Unlimited PTO + Company holidays Wellness initiatives DYN365, Inc d/b/a Restaurant365 is an equal opportunity employer. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
08/05/2026
Full time
Job Description Job Description Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering team members to produce top-notch results while elevating their skills. We're constantly evolving and improving to make sure we are and always will be "Best in Class" and we want that for you too! This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365's cloud infrastructure and applications. Qualified candidates will demonstrate growing expertise in site reliability practices, with skills in incident response, system monitoring, automation, and performance troubleshooting. You will collaborate with DevOps, development, and infrastructure teams to resolve moderately complex issues, propose improvements, and strengthen the reliability, scalability, and security of our SaaS platform. How you'll add value: Execution & Collaboration Respond to production incidents, perform triage and troubleshooting, and contribute to post-incident analysis. Identify and automate manual processes to improve efficiency and reduce risk. Enhance and evolve monitoring tools and platforms to improve observability. Promote and apply best practices for reliability, scalability, and performance across engineering. Implement and support cloud automation using Terraform, Ansible, or CloudFormation. Work within change management protocols to provide maximum uptime for production systems. Participate in on-call rotation, providing 24x7 support for incidents and contributing to root cause analysis. Partner with developers, architects, vendors, and IT teams to ensure reliable system operations. Research and remediate vulnerabilities in coordination with security teams. Maintain documentation of infrastructure, monitoring, runbooks, and incident response procedures. Standards & Process Apply company policies and procedures when handling operational tasks and incidents. Suggest and implement improvements to operational processes and monitoring practices. Contribute to technical diagrams, documentation, and runbooks for system reliability. Learning & Growth Expand expertise in cloud services (Azure, AWS, or GCP) and container platforms (EKS, ECS, AKS). Build proficiency with observability and monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios). Develop scripting and automation skills using Python, Bash, PowerShell, or similar. Participate in planning discussions by contributing technical input on system stability and reliability. What you'll need to be successful in this role: BS in Computer Science, Information Systems, or related field (or equivalent experience). 2-4 years of experience in site reliability engineering, DevOps, or cloud operations. Experience with cloud platforms (Azure or AWS), including services such as AKS, ECS/EKS, Functions/Lambda, S3, and Blob storage. Proficiency with infrastructure-as-code and automation (Terraform, Ansible, YAML, Python, Bash, PowerShell). Strong Linux engineering skills; working knowledge of Windows administration. Experience supporting production environments and participating in on-call rotations. Familiarity with web servers and middleware (Nginx, Apache Tomcat). Experience with CI/CD tools (GitLab, Git, or similar). Strong written, oral, and interpersonal communication skills. Preferred Qualifications Experience with monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios). Knowledge of performance analysis and system vulnerability remediation. Cloud certification (AWS or Azure) preferred. Familiarity with restaurant industry SaaS platforms and customer-facing applications. R365 Team Member Benefits & Compensation This position has a salary range of $98,583-$138,016 annually. The above range represents the expected salary range for this position. The actual salary may vary based upon several factors, including, but not limited to, relevant skills/experience, time in the role, business line, and geographic location. Restaurant365 focuses on equitable pay for our team and aims for transparency with our pay practices. Comprehensive medical benefits, 100% paid for employee 401k + matching Equity Option Grant Unlimited PTO + Company holidays Wellness initiatives DYN365, Inc d/b/a Restaurant365 is an equal opportunity employer. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Cloud SRE Engineer - Mandarin Bilingual
IntelliPro Group Inc. Palo Alto, California
Job Description Job Description Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type: Contract (12 months) Location: Palo Alto, CA Salary Rate: $70-$100 per hour (USD) Job ID#: Job Description: North America cloud operations team is looking for a skilled Cloud SRE Engineer to own the reliability, stability, and continuous improvement of core cloud services - spanning compute infrastructure (CVM/VMs), networking, and cloud security products. You'll work in a production-critical environment where operational excellence, deep technical expertise, and a self-directed mindset are essential. Since the North America team operates independently from teams in China and Singapore with no overlapping hours, we're looking for someone who can hit the ground running with minimal ramp-up time. Responsibilities: Monitor and maintain cloud compute (CVM), networking, and security products in the North America region to ensure high availability and system stability Respond to and resolve production incidents, customer-reported issues, and system-level outages with urgency and ownership Perform deep troubleshooting across network, compute, security, and platform layers Participate in on-call rotation and handle live production issues independently Deploy new features, bug fixes, and enhancements into production environments using CI/CD pipelines and internal tooling Develop scripts and automation tools to improve operational efficiency and reduce toil Build and improve monitoring, alerting, and disaster recovery systems for 24/7 operations Document operational workflows, runbooks, and best practices Work closely with R&D, security, and platform teams across time zones to drive service reliability Communicate technical issues clearly to internal teams and B2B customers Requirements: Some SRE, DevOps, or cloud operations experience - ability to maintain application stability independently is essential given timezone constraints Mandarin/English bilingual preferred - ability to communicate with teams in China and Singapore is a plus Strong networking fundamentals (TCP/IP, DNS, HTTP, ICMP, load balancing, firewalls, VPC) OR deep Linux/CVM knowledge - ability to own either the networking or compute side of operations Hands-on experience with cloud platforms (AWS, GCP, Azure, or equivalent) - deployment, usage, and high availability Familiarity with Kubernetes and container-based deployments Proficiency in at least one scripting language (Python, Shell, or Go) with automation experience Strong troubleshooting and debugging skills across infrastructure layers Experience with monitoring and alerting tools (Grafana, Prometheus, CloudWatch, or equivalent) Bachelor's degree or above in Computer Science or a related field Strong self-directed work ethic - able to operate independently with minimal supervision across time zones About Us: Founded in 2009, IntelliPro is a global leader in talent acquisition and HR solutions. Our commitment to delivering unparalleled service to clients, fostering employee growth, and building enduring partnerships sets us apart. We continue leading global talent solutions with a dynamic presence in over 160 countries, including the USA, China, Canada, Singapore, Japan, Philippines, UK, India, Netherlands, and the EU. IntelliPro, a global leader connecting individuals with rewarding employment opportunities, is dedicated to understanding your career aspirations. As an Equal Opportunity Employer, IntelliPro values diversity and does not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, genetic information, disability, or any other legally protected group status. Moreover, our Inclusivity Commitment emphasizes embracing candidates of all abilities and ensures that our hiring and interview processes accommodate the needs of all applicants. Learn more about our commitment to diversity and inclusivity at Compensation: The pay offered to a successful candidate will be determined by various factors, including education, work experience, location, job responsibilities, certifications, and more. Additionally, IntelliPro provides a comprehensive benefits package, all subject to eligibility. Powered by JazzHR R4PfTsWHf3
08/05/2026
Full time
Job Description Job Description Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type: Contract (12 months) Location: Palo Alto, CA Salary Rate: $70-$100 per hour (USD) Job ID#: Job Description: North America cloud operations team is looking for a skilled Cloud SRE Engineer to own the reliability, stability, and continuous improvement of core cloud services - spanning compute infrastructure (CVM/VMs), networking, and cloud security products. You'll work in a production-critical environment where operational excellence, deep technical expertise, and a self-directed mindset are essential. Since the North America team operates independently from teams in China and Singapore with no overlapping hours, we're looking for someone who can hit the ground running with minimal ramp-up time. Responsibilities: Monitor and maintain cloud compute (CVM), networking, and security products in the North America region to ensure high availability and system stability Respond to and resolve production incidents, customer-reported issues, and system-level outages with urgency and ownership Perform deep troubleshooting across network, compute, security, and platform layers Participate in on-call rotation and handle live production issues independently Deploy new features, bug fixes, and enhancements into production environments using CI/CD pipelines and internal tooling Develop scripts and automation tools to improve operational efficiency and reduce toil Build and improve monitoring, alerting, and disaster recovery systems for 24/7 operations Document operational workflows, runbooks, and best practices Work closely with R&D, security, and platform teams across time zones to drive service reliability Communicate technical issues clearly to internal teams and B2B customers Requirements: Some SRE, DevOps, or cloud operations experience - ability to maintain application stability independently is essential given timezone constraints Mandarin/English bilingual preferred - ability to communicate with teams in China and Singapore is a plus Strong networking fundamentals (TCP/IP, DNS, HTTP, ICMP, load balancing, firewalls, VPC) OR deep Linux/CVM knowledge - ability to own either the networking or compute side of operations Hands-on experience with cloud platforms (AWS, GCP, Azure, or equivalent) - deployment, usage, and high availability Familiarity with Kubernetes and container-based deployments Proficiency in at least one scripting language (Python, Shell, or Go) with automation experience Strong troubleshooting and debugging skills across infrastructure layers Experience with monitoring and alerting tools (Grafana, Prometheus, CloudWatch, or equivalent) Bachelor's degree or above in Computer Science or a related field Strong self-directed work ethic - able to operate independently with minimal supervision across time zones About Us: Founded in 2009, IntelliPro is a global leader in talent acquisition and HR solutions. Our commitment to delivering unparalleled service to clients, fostering employee growth, and building enduring partnerships sets us apart. We continue leading global talent solutions with a dynamic presence in over 160 countries, including the USA, China, Canada, Singapore, Japan, Philippines, UK, India, Netherlands, and the EU. IntelliPro, a global leader connecting individuals with rewarding employment opportunities, is dedicated to understanding your career aspirations. As an Equal Opportunity Employer, IntelliPro values diversity and does not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, genetic information, disability, or any other legally protected group status. Moreover, our Inclusivity Commitment emphasizes embracing candidates of all abilities and ensures that our hiring and interview processes accommodate the needs of all applicants. Learn more about our commitment to diversity and inclusivity at Compensation: The pay offered to a successful candidate will be determined by various factors, including education, work experience, location, job responsibilities, certifications, and more. Additionally, IntelliPro provides a comprehensive benefits package, all subject to eligibility. Powered by JazzHR R4PfTsWHf3
Software Engineer - Site Reliability Engineering
Zoox San Mateo, California
Job Description Job Description Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services-from designing fault-tolerant, maintainable systems to deploying, operating, and continuously improving them in production. As a robotics company, Zoox embraces automation at every layer of our infrastructure, and you'll help drive that ethos forward. You'll work hands-on with systems that process massive volumes of data and support compute-intensive pipelines running on both CPUs and GPUs. In this role, you will: Architect and optimize scalable systems: You will design, implement, and continuously improve highly reliable infrastructure, directly impacting the success and safety of Zoox's autonomous vehicle platform. Build proactive monitoring solutions: You will develop advanced monitoring, alerting, and reporting tools to ensure potential issues are identified and resolved before they affect production. Collaborate across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline deployment processes, and drive automation initiatives. Lead incident resolution: You will conduct thorough root cause analyses on production issues and rapidly deploy corrective actions to maintain a resilient and stable environment. Ensure business continuity: You will safeguard the company's operations by designing and implementing robust disaster recovery plans to keep the Zoox fleet running smoothly under any circumstances. Qualifications SRE & Distributed Systems Experience: 5+ years of experience in site reliability engineering or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure as Code (IaC): Proven experience operating within major cloud platforms (AWS, GCP, or Azure) and utilizing IaC tools like Terraform, Ansible, Salt, or CloudFormation. Container Orchestration: Technical expertise in deploying, managing, and scaling systems using container orchestration technologies such as Kubernetes. Core Infrastructure Knowledge: Deep, foundational understanding of networking protocols, storage solutions, and database technologies. Programming Proficiency: Strong, demonstrable programming and scripting skills in languages such as Python, Go, C/C++, or Java. Bonus Qualifications Experience in the automotive or autonomous vehicle industry. Knowledge of security best practices and compliance requirements. About Zoox Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We're looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team. Follow us on LinkedIn Accommodations If you need an accommodation to participate in the application or interview process please reach out to or your assigned recruiter. A Final Note: You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
08/05/2026
Full time
Job Description Job Description Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services-from designing fault-tolerant, maintainable systems to deploying, operating, and continuously improving them in production. As a robotics company, Zoox embraces automation at every layer of our infrastructure, and you'll help drive that ethos forward. You'll work hands-on with systems that process massive volumes of data and support compute-intensive pipelines running on both CPUs and GPUs. In this role, you will: Architect and optimize scalable systems: You will design, implement, and continuously improve highly reliable infrastructure, directly impacting the success and safety of Zoox's autonomous vehicle platform. Build proactive monitoring solutions: You will develop advanced monitoring, alerting, and reporting tools to ensure potential issues are identified and resolved before they affect production. Collaborate across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline deployment processes, and drive automation initiatives. Lead incident resolution: You will conduct thorough root cause analyses on production issues and rapidly deploy corrective actions to maintain a resilient and stable environment. Ensure business continuity: You will safeguard the company's operations by designing and implementing robust disaster recovery plans to keep the Zoox fleet running smoothly under any circumstances. Qualifications SRE & Distributed Systems Experience: 5+ years of experience in site reliability engineering or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure as Code (IaC): Proven experience operating within major cloud platforms (AWS, GCP, or Azure) and utilizing IaC tools like Terraform, Ansible, Salt, or CloudFormation. Container Orchestration: Technical expertise in deploying, managing, and scaling systems using container orchestration technologies such as Kubernetes. Core Infrastructure Knowledge: Deep, foundational understanding of networking protocols, storage solutions, and database technologies. Programming Proficiency: Strong, demonstrable programming and scripting skills in languages such as Python, Go, C/C++, or Java. Bonus Qualifications Experience in the automotive or autonomous vehicle industry. Knowledge of security best practices and compliance requirements. About Zoox Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We're looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team. Follow us on LinkedIn Accommodations If you need an accommodation to participate in the application or interview process please reach out to or your assigned recruiter. A Final Note: You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Site Reliability Engineer
Obsidian Security Palo Alto, California
Job Description Job Description Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. We protect 200+ organizations across North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand, including many of the world's largest Fortune 1000 and Global 2000 companies. Founded in 2017 and backed by top investors like Greylock, Obsidian was built to close a critical gap: securing SaaS apps where business happens-Microsoft 365, Salesforce, and hundreds more. The company does this by offering a complete SaaS security platform to reduce risk, detect and respond to threats, and prevent breaches at the source. Obsidian was built by leaders who redefined endpoint and identity security at CrowdStrike, Okta, Cylance, and Carbon Black. Now, they're transforming how SaaS is secured. With AI driving rapid SaaS growth and complexity, agentic AI tools gain privileged access to sensitive data through integrations, creating new risks most security tools miss. Obsidian uniquely detects anomalous OAuth token activity and manages integration risks. Major announcements are on the horizon. Recognizing that SaaS security needs to evolve, Obsidian enables growing organizations to start with a lightweight, prevention-focused browser extension and expand coverage over time. With global momentum, a growing partner ecosystem including SentinelOne, Databricks, and Google Cloud, and a major fundraise ahead, Obsidian is scaling rapidly toward long-term growth and IPO readiness. About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-performing production systems. We work closely with Engineering, Quality Engineering, and Customer Support to deliver end-to-end services that bring code to life and maintain our world-class SaaS security platform. What You'll Do Support and maintain the service quality of our customer-facing SaaS security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with Engineering teams to maintain and enhance Helm charts, application deployment, monitoring and CI/CD pipelines Embed into the engineering team so that you understand the application deeply. Define service verification strategies and implement them as part of the CI/CD process to meet SLAs Improve developer experience by optimizing CI/CD workflows and performance Participate in the on-call rotation, providing 24/7 support in coordination with our global SRE team Monitor, debug, and optimize production infrastructure and services on AWS/GCP What We're Looking For 3+ years of experience in a DevOps or SRE role supporting SaaS services on GCP and/or AWS Bachelor's degree in Computer Science or related field Strong proficiency in Kubernetes, microservices architecture, Helm, GitLab CI/CD, and ArgoCD, Prometheus, Grafana. Programming experience in at least one language; Golang or Python preferred Deep understanding of autoscaling, version upgrades, and cloud service optimization Bonus if you're familiar with technologies like Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong Employee Benefits Our competitive benefits packages are designed to support our employees' well-being, both at work and at home. Our US based employees enjoy: Competitive compensation with equity and 401k Comprehensive healthcare with dental and vision coverage Flexible paid time off and paid holiday time off 12 weeks of new parent or family leave Personal and professional development resources For more details on our US benefits, or for information on our international benefits, please see here. Pay Transparancy Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company. At Obsidian, we are proud to be an equal-opportunity employer. We value diversity and hire for talent, passion, and compassion. In compliance with federal law, all persons hired will be required to submit satisfactory proof of identity and legal authorization. If you have a need that requires accommodation, please contact Information collected and processed as part of any job applications you choose to submit is subject to Obsidian's Applicant Privacy Policy. Base Salary Range $165,000-$190,000 USD
08/05/2026
Full time
Job Description Job Description Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. We protect 200+ organizations across North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand, including many of the world's largest Fortune 1000 and Global 2000 companies. Founded in 2017 and backed by top investors like Greylock, Obsidian was built to close a critical gap: securing SaaS apps where business happens-Microsoft 365, Salesforce, and hundreds more. The company does this by offering a complete SaaS security platform to reduce risk, detect and respond to threats, and prevent breaches at the source. Obsidian was built by leaders who redefined endpoint and identity security at CrowdStrike, Okta, Cylance, and Carbon Black. Now, they're transforming how SaaS is secured. With AI driving rapid SaaS growth and complexity, agentic AI tools gain privileged access to sensitive data through integrations, creating new risks most security tools miss. Obsidian uniquely detects anomalous OAuth token activity and manages integration risks. Major announcements are on the horizon. Recognizing that SaaS security needs to evolve, Obsidian enables growing organizations to start with a lightweight, prevention-focused browser extension and expand coverage over time. With global momentum, a growing partner ecosystem including SentinelOne, Databricks, and Google Cloud, and a major fundraise ahead, Obsidian is scaling rapidly toward long-term growth and IPO readiness. About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-performing production systems. We work closely with Engineering, Quality Engineering, and Customer Support to deliver end-to-end services that bring code to life and maintain our world-class SaaS security platform. What You'll Do Support and maintain the service quality of our customer-facing SaaS security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with Engineering teams to maintain and enhance Helm charts, application deployment, monitoring and CI/CD pipelines Embed into the engineering team so that you understand the application deeply. Define service verification strategies and implement them as part of the CI/CD process to meet SLAs Improve developer experience by optimizing CI/CD workflows and performance Participate in the on-call rotation, providing 24/7 support in coordination with our global SRE team Monitor, debug, and optimize production infrastructure and services on AWS/GCP What We're Looking For 3+ years of experience in a DevOps or SRE role supporting SaaS services on GCP and/or AWS Bachelor's degree in Computer Science or related field Strong proficiency in Kubernetes, microservices architecture, Helm, GitLab CI/CD, and ArgoCD, Prometheus, Grafana. Programming experience in at least one language; Golang or Python preferred Deep understanding of autoscaling, version upgrades, and cloud service optimization Bonus if you're familiar with technologies like Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong Employee Benefits Our competitive benefits packages are designed to support our employees' well-being, both at work and at home. Our US based employees enjoy: Competitive compensation with equity and 401k Comprehensive healthcare with dental and vision coverage Flexible paid time off and paid holiday time off 12 weeks of new parent or family leave Personal and professional development resources For more details on our US benefits, or for information on our international benefits, please see here. Pay Transparancy Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company. At Obsidian, we are proud to be an equal-opportunity employer. We value diversity and hire for talent, passion, and compassion. In compliance with federal law, all persons hired will be required to submit satisfactory proof of identity and legal authorization. If you have a need that requires accommodation, please contact Information collected and processed as part of any job applications you choose to submit is subject to Obsidian's Applicant Privacy Policy. Base Salary Range $165,000-$190,000 USD

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board