it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

10 jobs found

Email me jobs like this
Refine Search
Current Search
site reliability engineer sre
GenAI Engineer - SRE
Dimension Consulting Phoenix, Arizona
Job Titel: GenAI Engineer Location: Phoenix, AZ (Client Site - No Remote, No Virtual) Work Model: Hybrid (4 days onsite Mon-Thu, Friday remote) Interviews: 2-3 rounds (including client interview) Role Summary Seeking a GenAI Engineer to support large-scale Site Reliability Engineering (SRE), Platform Health, and Engineering Productivity initiatives. The candidate will leverage Generative AI technologies, automation frameworks, and cloud-native tooling to improve operational efficiency, incident reduction, observability, code quality, and developer productivity across enterprise platforms. The role requires close collaboration with SRE teams, platform engineering teams, application development teams, and business stakeholders to identify automation opportunities and implement AI-driven solutions. Key Responsibilities GenAI & Automation Technical Skills Python development and automation scripting Generative AI concepts and LLM integrations GitHub Copilot, Devin AI, OpenAI/Azure OpenAI ecosystem REST APIs and automation frameworks Git, CI/CD pipelines, GitHub Actions Linux/Unix administration Cloud Platforms (AWS, Azure, GCP) Terraform and Infrastructure as Code Monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, New Relic SRE Knowledge Incident Management Problem Management Service Reliability Availability Management Operational Excellence Observability Principles Production Support
08/04/2026
Full time
Job Titel: GenAI Engineer Location: Phoenix, AZ (Client Site - No Remote, No Virtual) Work Model: Hybrid (4 days onsite Mon-Thu, Friday remote) Interviews: 2-3 rounds (including client interview) Role Summary Seeking a GenAI Engineer to support large-scale Site Reliability Engineering (SRE), Platform Health, and Engineering Productivity initiatives. The candidate will leverage Generative AI technologies, automation frameworks, and cloud-native tooling to improve operational efficiency, incident reduction, observability, code quality, and developer productivity across enterprise platforms. The role requires close collaboration with SRE teams, platform engineering teams, application development teams, and business stakeholders to identify automation opportunities and implement AI-driven solutions. Key Responsibilities GenAI & Automation Technical Skills Python development and automation scripting Generative AI concepts and LLM integrations GitHub Copilot, Devin AI, OpenAI/Azure OpenAI ecosystem REST APIs and automation frameworks Git, CI/CD pipelines, GitHub Actions Linux/Unix administration Cloud Platforms (AWS, Azure, GCP) Terraform and Infrastructure as Code Monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, New Relic SRE Knowledge Incident Management Problem Management Service Reliability Availability Management Operational Excellence Observability Principles Production Support
Principal Cloud Engineer - Remote or Hybrid
Genesis10 Saint Paul, Minnesota
Genesis10 is currently seeking a Principal Cloud Engineer - Remote or Hybrid position with a Major Healthcare Company located in Eagan, MN. This is a 9+ month contract opportunity. Compensation: $72.50 - $82.50 per hour, W2, based on qualifications Position Overview: We are seeking an experienced Principal Cloud Engineer to support the expansion of an AWS-based cloud platform supporting an enterprise Facets environment. This is a hands-on engineering role responsible for implementing, configuring, and optimizing cloud infrastructure using AWS EKS, Kubernetes, and Terraform, while working within established enterprise architecture, security standards, and regulatory requirements. Key Responsibilities: Deploy, configure, and support AWS infrastructure using Terraform Build and maintain reusable Terraform modules and Infrastructure as Code standards Design and support Kubernetes (Amazon EKS) environments supporting enterprise applications Assist with the implementation and expansion of the organization's Facets platform within AWS Support cloud platform reliability, scalability, and operational excellence Support day-to-day Cloud Operations activities alongside the Cloud Engineering team Troubleshoot infrastructure issues across AWS environments Participate in operational support, incident response, and root cause analysis Improve automation, operational efficiency, and platform stability Engineer solutions within established enterprise cloud standards and governance Implement secure IAM, networking, encryption, and infrastructure best practices Ensure solutions comply with healthcare security and regulatory requirements Partner with security and architecture teams to maintain compliance and operational standards Work closely with Cloud Operations, Infrastructure Engineering, Security, and Application teams Evaluate how new infrastructure integrates into the broader enterprise environment Provide technical leadership while remaining hands-on with implementation Contribute engineering expertise while adhering to established enterprise architecture Primary Requirements: 5+ years of cloud infrastructure or platform engineering experience Strong hands-on experience with AWS, Amazon EKS, Kubernetes, and Terraform Experience deploying and managing Infrastructure as Code in production environments Strong understanding of VPC networking, IAM, CloudTrail, CloudWatch, and Sumo Logic Experience supporting cloud infrastructure in enterprise environments Strong troubleshooting and operational support skills Ability to think strategically while delivering hands-on engineering solutions Experience working within established engineering standards and governance Desired Skills: Healthcare or other highly regulated industry experience Exposure to Facets or other enterprise claims platforms AWS Certifications (Solutions Architect, DevOps Engineer, or similar) Experience with CI/CD pipelines and cloud automation Experience supporting Site Reliability Engineering (SRE) practices Knowledge of enterprise cloud security, compliance, and governance Only candidates available and ready to work directly as Genesis10 employees will be considered for this position. If you have the described qualifications and are interested in this exciting opportunity, please apply! Ranked a Top Staffing Firm in the U.S. by Staffing Industry Analysts for six consecutive years, Genesis10 puts thousands of consultants and employees to work across the United States every year in contract, contract-for-hire, and permanent placement roles. With more than 300 active clients, Genesis10 provides access to many of the Fortune 100 firms and a variety of mid-market organizations across the full spectrum of industry verticals. For contract roles, Genesis10 offers the benefits listed below. If this is a perm-placement opportunity, our recruiter can talk you through the unique benefits offered for that particular client. Benefits of Working with Genesis10: Access to hundreds of clients, most who have been working with Genesis10 for 5-20+ years The opportunity to have a career-home in Genesis10; many of our consultants have been working exclusively with Genesis10 for years Access to an experienced, caring recruiting team (more than 7 years of experience, on average) Behavioral Health Platform Medical, Dental, Vision Health Savings Account Voluntary Hospital Indemnity (Critical Illness & Accident) Voluntary Term Life Insurance 401K Sick Pay (for applicable states/municipalities) Commuter Benefits (Dallas, NYC, SF, and Illinois) For multiple years running, Genesis10 has been recognized as a Top Staffing Firm in the U.S., as a Best Company for Work-Life Balance, as a Best Company for Career Growth, for Diversity, and for Leadership, amongst others. To learn more and to view all of our available career opportunities, please visit us at our website. Genesis10 is an Equal Opportunity Employer. Candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
08/04/2026
Full time
Genesis10 is currently seeking a Principal Cloud Engineer - Remote or Hybrid position with a Major Healthcare Company located in Eagan, MN. This is a 9+ month contract opportunity. Compensation: $72.50 - $82.50 per hour, W2, based on qualifications Position Overview: We are seeking an experienced Principal Cloud Engineer to support the expansion of an AWS-based cloud platform supporting an enterprise Facets environment. This is a hands-on engineering role responsible for implementing, configuring, and optimizing cloud infrastructure using AWS EKS, Kubernetes, and Terraform, while working within established enterprise architecture, security standards, and regulatory requirements. Key Responsibilities: Deploy, configure, and support AWS infrastructure using Terraform Build and maintain reusable Terraform modules and Infrastructure as Code standards Design and support Kubernetes (Amazon EKS) environments supporting enterprise applications Assist with the implementation and expansion of the organization's Facets platform within AWS Support cloud platform reliability, scalability, and operational excellence Support day-to-day Cloud Operations activities alongside the Cloud Engineering team Troubleshoot infrastructure issues across AWS environments Participate in operational support, incident response, and root cause analysis Improve automation, operational efficiency, and platform stability Engineer solutions within established enterprise cloud standards and governance Implement secure IAM, networking, encryption, and infrastructure best practices Ensure solutions comply with healthcare security and regulatory requirements Partner with security and architecture teams to maintain compliance and operational standards Work closely with Cloud Operations, Infrastructure Engineering, Security, and Application teams Evaluate how new infrastructure integrates into the broader enterprise environment Provide technical leadership while remaining hands-on with implementation Contribute engineering expertise while adhering to established enterprise architecture Primary Requirements: 5+ years of cloud infrastructure or platform engineering experience Strong hands-on experience with AWS, Amazon EKS, Kubernetes, and Terraform Experience deploying and managing Infrastructure as Code in production environments Strong understanding of VPC networking, IAM, CloudTrail, CloudWatch, and Sumo Logic Experience supporting cloud infrastructure in enterprise environments Strong troubleshooting and operational support skills Ability to think strategically while delivering hands-on engineering solutions Experience working within established engineering standards and governance Desired Skills: Healthcare or other highly regulated industry experience Exposure to Facets or other enterprise claims platforms AWS Certifications (Solutions Architect, DevOps Engineer, or similar) Experience with CI/CD pipelines and cloud automation Experience supporting Site Reliability Engineering (SRE) practices Knowledge of enterprise cloud security, compliance, and governance Only candidates available and ready to work directly as Genesis10 employees will be considered for this position. If you have the described qualifications and are interested in this exciting opportunity, please apply! Ranked a Top Staffing Firm in the U.S. by Staffing Industry Analysts for six consecutive years, Genesis10 puts thousands of consultants and employees to work across the United States every year in contract, contract-for-hire, and permanent placement roles. With more than 300 active clients, Genesis10 provides access to many of the Fortune 100 firms and a variety of mid-market organizations across the full spectrum of industry verticals. For contract roles, Genesis10 offers the benefits listed below. If this is a perm-placement opportunity, our recruiter can talk you through the unique benefits offered for that particular client. Benefits of Working with Genesis10: Access to hundreds of clients, most who have been working with Genesis10 for 5-20+ years The opportunity to have a career-home in Genesis10; many of our consultants have been working exclusively with Genesis10 for years Access to an experienced, caring recruiting team (more than 7 years of experience, on average) Behavioral Health Platform Medical, Dental, Vision Health Savings Account Voluntary Hospital Indemnity (Critical Illness & Accident) Voluntary Term Life Insurance 401K Sick Pay (for applicable states/municipalities) Commuter Benefits (Dallas, NYC, SF, and Illinois) For multiple years running, Genesis10 has been recognized as a Top Staffing Firm in the U.S., as a Best Company for Work-Life Balance, as a Best Company for Career Growth, for Diversity, and for Leadership, amongst others. To learn more and to view all of our available career opportunities, please visit us at our website. Genesis10 is an Equal Opportunity Employer. Candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Site Reliability Engineer (SRE)
BC Forward Chandler, Arizona
Job Title: Site Reliability Engineer (SRE) / Cloud Automation Engineer Location: Chandler, AZ (Hybrid, 3 days onsite per week required). Local candidates only. Duration: Contract - 12 months Pay Range: $70.23/hr (W2) Job ID: 407394 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Site Reliability Engineer III to join our team. The ideal candidate will have strong experience in cloud computing, infrastructure automation, and observability tooling and a proven ability to implement reliable, automated, and measurable service operations across complex environments. Responsibilities: Establish and maintain partnerships with Application Development and Production Support teams to drive SRE practices. Ensure instrumentation, tooling, ticketing, alerting, and on-call routines are in place for key services. Lead automation efforts using Infrastructure as Code, Terraform, and Ansible to improve reliability and speed of delivery. Design and implement solutions across AWS, Azure, and OpenShift with a focus on compute, storage, network, and security services. Develop and maintain code and automation using Python, Golang, and shell scripting. Implement monitoring and observability with Prometheus, Dynatrace, Azure Monitor, and Log Analytics. Contribute to CI/CD pipelines using Git, Jenkins, and GitOps practices. Decompose complex objectives into units of work for team execution and mentor peers on best practices. Required Skills & Qualifications: Hands-on experience with Terraform Enterprise, HashiCorp Consul, and Red Hat OpenShift. Proficiency in AWS and Azure cloud services across compute, storage, network, and security. Programming and scripting with Python, Golang, shell, and familiarity with Java. Experience with Ansible Automation and CI/CD tools including Git, Jenkins, and GitOps models. Monitoring expertise with Prometheus and Dynatrace, and cloud-native tools such as Azure Monitor and Log Analytics. Strong systems background across RHEL and Windows Servers, containers, and virtual resources in Azure and AWS. Experience with workload management tools such as Jira and Remedy. BS/MS in Computer Science or related field, or equivalent practical experience. 5+ years relevant experience. Preferred Skills: Linux System Administration and Windows Administration. Splunk Administration, Dynatrace Administration, and Grafana. OpenShift containers and Ansible Automation. Horizon CI/CD ecosystem (Jenkins, XLR, Artifactory, Bitbucket). Experience with Azure, AWS, or GCP. HashiCorp Terraform and Red Hat Ansible certifications. Strong problem-solving approach, ownership, and ability to manage competing priorities. Additional Details: Primary Skill: Cloud Computing. Secondary: Robotic Process Automation. Tertiary: Automation Testing. Maximum submissions per supplier: 4. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to cutting-edge technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
08/04/2026
Full time
Job Title: Site Reliability Engineer (SRE) / Cloud Automation Engineer Location: Chandler, AZ (Hybrid, 3 days onsite per week required). Local candidates only. Duration: Contract - 12 months Pay Range: $70.23/hr (W2) Job ID: 407394 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Site Reliability Engineer III to join our team. The ideal candidate will have strong experience in cloud computing, infrastructure automation, and observability tooling and a proven ability to implement reliable, automated, and measurable service operations across complex environments. Responsibilities: Establish and maintain partnerships with Application Development and Production Support teams to drive SRE practices. Ensure instrumentation, tooling, ticketing, alerting, and on-call routines are in place for key services. Lead automation efforts using Infrastructure as Code, Terraform, and Ansible to improve reliability and speed of delivery. Design and implement solutions across AWS, Azure, and OpenShift with a focus on compute, storage, network, and security services. Develop and maintain code and automation using Python, Golang, and shell scripting. Implement monitoring and observability with Prometheus, Dynatrace, Azure Monitor, and Log Analytics. Contribute to CI/CD pipelines using Git, Jenkins, and GitOps practices. Decompose complex objectives into units of work for team execution and mentor peers on best practices. Required Skills & Qualifications: Hands-on experience with Terraform Enterprise, HashiCorp Consul, and Red Hat OpenShift. Proficiency in AWS and Azure cloud services across compute, storage, network, and security. Programming and scripting with Python, Golang, shell, and familiarity with Java. Experience with Ansible Automation and CI/CD tools including Git, Jenkins, and GitOps models. Monitoring expertise with Prometheus and Dynatrace, and cloud-native tools such as Azure Monitor and Log Analytics. Strong systems background across RHEL and Windows Servers, containers, and virtual resources in Azure and AWS. Experience with workload management tools such as Jira and Remedy. BS/MS in Computer Science or related field, or equivalent practical experience. 5+ years relevant experience. Preferred Skills: Linux System Administration and Windows Administration. Splunk Administration, Dynatrace Administration, and Grafana. OpenShift containers and Ansible Automation. Horizon CI/CD ecosystem (Jenkins, XLR, Artifactory, Bitbucket). Experience with Azure, AWS, or GCP. HashiCorp Terraform and Red Hat Ansible certifications. Strong problem-solving approach, ownership, and ability to manage competing priorities. Additional Details: Primary Skill: Cloud Computing. Secondary: Robotic Process Automation. Tertiary: Automation Testing. Maximum submissions per supplier: 4. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to cutting-edge technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
DevOps Platform Development Engineer (Bilingual Mandarin)
Comrise Palo Alto, California
DevOps Platform Development Engineer (Bilingual Mandarin & English) Location: Palo Alto, CA (Onsite) Employment Type: Full-Time, Permanent Compensation: $150,000 - $220,000 annually (12 months base + 2 months bonus) Job Overview: Responsible for deploying, integrating, adapting, and operating domestic DevOps and High Availability (HA) platforms across overseas regions, ensuring successful rollout according to project timelines. Evaluate differences between overseas public cloud, private cloud, and on-premises data center environments in areas such as compute, networking, storage, Kubernetes, security, and compliance, and develop and execute localization and adaptation plans. On the DevOps side, lead the overseas rollout of service deployment platforms, change management systems, and large-scale operations platforms, while optimizing CI/CD pipelines, infrastructure automation, and release management processes. On the High Availability side, implement disaster recovery capabilities, failover strategies, DR validation, and chaos engineering practices, drive issue remediation, and continuously improve the reliability of overseas production environments. Participate in the day-to-day operations and incident response of overseas production environments, quickly troubleshoot deployment, infrastructure, and platform integration issues, and drive root cause analysis and long-term corrective actions. Develop and maintain deployment documentation, operational runbooks, disaster recovery plans, and standard operating procedures (SOPs) to improve the efficiency of global platform delivery and operations. Collaborate closely with platform engineering teams in China as well as overseas infrastructure, networking, security, and business teams to manage implementation schedules, dependencies, and risks, while providing feedback to improve platform capabilities for international deployment scenarios. Required Qualifications: Bachelor's degree or above in Computer Science, Software Engineering, or a related field preferred. Hands-on experience in DevOps, Site Reliability Engineering (SRE), cloud platforms, or infrastructure engineering, with strong delivery and problem-solving skills. Solid understanding of Linux, computer networking, and common infrastructure components, with the ability to independently deploy, configure, and troubleshoot systems. Experience with cloud-native technologies and Infrastructure as Code (IaC), including Kubernetes, Helm, and Terraform. Experience with KubeVela is a plus. Proficiency in one or more programming/scripting languages such as Go, Java, Python, or Shell, with the ability to develop automation tools and read/debug backend service code. Familiarity with at least one major public cloud platform or enterprise private cloud environment, with a solid understanding of multi-region deployment, network connectivity, identity and access management, and security isolation. Experience with monitoring, logging, and alerting systems, including tools such as Prometheus, Grafana, ELK/OpenSearch, or OpenTelemetry. Professional proficiency in both Mandarin and English is required to support collaboration with global teams and stakeholders who primarily communicate in these languages. Strong sense of ownership, excellent execution skills, and the ability to translate centrally designed platform solutions into stable, maintainable local implementations. Preferred Qualifications: Experience deploying internal platforms globally, migrating infrastructure, or supporting multi-region deployments. Experience building or operating DevOps toolchains and large-scale infrastructure at major Internet or technology companies. Experience with disaster recovery, chaos engineering, capacity planning, or managing large-scale production incidents. Familiarity with data residency requirements, privacy protection regulations, and regional infrastructure compliance standards.
08/04/2026
Full time
DevOps Platform Development Engineer (Bilingual Mandarin & English) Location: Palo Alto, CA (Onsite) Employment Type: Full-Time, Permanent Compensation: $150,000 - $220,000 annually (12 months base + 2 months bonus) Job Overview: Responsible for deploying, integrating, adapting, and operating domestic DevOps and High Availability (HA) platforms across overseas regions, ensuring successful rollout according to project timelines. Evaluate differences between overseas public cloud, private cloud, and on-premises data center environments in areas such as compute, networking, storage, Kubernetes, security, and compliance, and develop and execute localization and adaptation plans. On the DevOps side, lead the overseas rollout of service deployment platforms, change management systems, and large-scale operations platforms, while optimizing CI/CD pipelines, infrastructure automation, and release management processes. On the High Availability side, implement disaster recovery capabilities, failover strategies, DR validation, and chaos engineering practices, drive issue remediation, and continuously improve the reliability of overseas production environments. Participate in the day-to-day operations and incident response of overseas production environments, quickly troubleshoot deployment, infrastructure, and platform integration issues, and drive root cause analysis and long-term corrective actions. Develop and maintain deployment documentation, operational runbooks, disaster recovery plans, and standard operating procedures (SOPs) to improve the efficiency of global platform delivery and operations. Collaborate closely with platform engineering teams in China as well as overseas infrastructure, networking, security, and business teams to manage implementation schedules, dependencies, and risks, while providing feedback to improve platform capabilities for international deployment scenarios. Required Qualifications: Bachelor's degree or above in Computer Science, Software Engineering, or a related field preferred. Hands-on experience in DevOps, Site Reliability Engineering (SRE), cloud platforms, or infrastructure engineering, with strong delivery and problem-solving skills. Solid understanding of Linux, computer networking, and common infrastructure components, with the ability to independently deploy, configure, and troubleshoot systems. Experience with cloud-native technologies and Infrastructure as Code (IaC), including Kubernetes, Helm, and Terraform. Experience with KubeVela is a plus. Proficiency in one or more programming/scripting languages such as Go, Java, Python, or Shell, with the ability to develop automation tools and read/debug backend service code. Familiarity with at least one major public cloud platform or enterprise private cloud environment, with a solid understanding of multi-region deployment, network connectivity, identity and access management, and security isolation. Experience with monitoring, logging, and alerting systems, including tools such as Prometheus, Grafana, ELK/OpenSearch, or OpenTelemetry. Professional proficiency in both Mandarin and English is required to support collaboration with global teams and stakeholders who primarily communicate in these languages. Strong sense of ownership, excellent execution skills, and the ability to translate centrally designed platform solutions into stable, maintainable local implementations. Preferred Qualifications: Experience deploying internal platforms globally, migrating infrastructure, or supporting multi-region deployments. Experience building or operating DevOps toolchains and large-scale infrastructure at major Internet or technology companies. Experience with disaster recovery, chaos engineering, capacity planning, or managing large-scale production incidents. Familiarity with data residency requirements, privacy protection regulations, and regional infrastructure compliance standards.
Site Reliability Engineer
BC Forward Buffalo, New York
Job Title: Site Reliability Engineer Location: Buffalo NY Duration: Temp - 12 months Pay Range: $100/hr $110/hr (W2) Job ID: 407436 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex reliability initiatives. Responsibilities: Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure aligned to enterprise standards and SRE best practices. Lead initiatives to improve reliability, availability, performance, and operational maturity through automation and engineering excellence. Define, implement, and monitor SLOs, SLIs, and error budgets for critical services. Develop observability strategies using Dynatrace, OpenTelemetry, distributed tracing, metrics, logs, dashboards, and alerting. Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure, and customer experience health. Analyze production telemetry to identify performance bottlenecks, reliability risks, and capacity constraints proactively. Lead incident response for high-severity events and coordinate cross-functional restoration and communications. Perform and facilitate RCAs with corrective and preventive actions tracked to completion. Automate repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls. Partner with development teams to embed reliability and observability across the SDLC. Design, develop, and execute automated regression testing to validate stability and performance after changes. Review test coverage and reliability validation to ensure comprehensive risk mitigation. Create, maintain, and improve Terraform-based IaC for provisioning, configuration, and standardization. Support and optimize Microsoft Azure environments, including App Services, resource management, scaling, and deployment automation. Use Azure Monitor, Application Insights, and Log Analytics to improve visibility and reliability. Drive performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness. Establish operational readiness standards and enforce requirements before production deployments. Review architectures and recommend improvements for resiliency, efficiency, and cloud optimization. Lead capacity planning, performance tuning, and workload optimization across production environments. Develop and maintain runbooks, incident playbooks, knowledge articles, and SOPs. Partner with engineering, infrastructure, cybersecurity, architecture, and support teams on cross-functional improvements. Communicate system health, reliability trends, risks, and remediation to technical and business stakeholders. Present initiatives, metrics, and recommendations at reviews, forums, and leadership meetings. Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational practices. Adhere to risk and regulatory standards and identify issues requiring escalation. Promote a culture of belonging consistent with company values and maintain internal control standards. Required Skills & Qualifications: Associate's degree with 7+ years in SRE, Cloud, Systems, Infrastructure Engineering, DevOps, or Application Support. Bachelor's degree with 5+ years. Or 9+ years combined education and experience with 5+ years in a technology engineering role. Hands-on observability and monitoring experience with Dynatrace, OpenTelemetry, distributed tracing, metrics, centralized logging, alerting, and dashboards. Proven ability to design and execute automated regression testing frameworks and suites. Strong proficiency with Terraform and Infrastructure as Code practices. Experience with CI/CD, deployment automation, and operational tooling. Expertise in production monitoring, incident management, and troubleshooting of distributed systems. Understanding of application performance management and modern cloud-native architectures. Preferred Skills: Microsoft Azure expertise, including App Services, Resource Groups, networking, scaling and optimization, deployment and release management, and application lifecycle management. Use of Azure Monitor, Application Insights, Log Analytics, dashboards, and alerting. SRE practices such as SLOs, SLIs, error budgets, incident/problem management, RCA, and reliability automation. Experience with performance tuning, capacity planning, proactive issue detection, and observability-driven improvements. Automated recovery mechanisms and self-healing solutions, resiliency patterns, DR planning, and high-availability architectures. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to modern technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
08/04/2026
Full time
Job Title: Site Reliability Engineer Location: Buffalo NY Duration: Temp - 12 months Pay Range: $100/hr $110/hr (W2) Job ID: 407436 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex reliability initiatives. Responsibilities: Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure aligned to enterprise standards and SRE best practices. Lead initiatives to improve reliability, availability, performance, and operational maturity through automation and engineering excellence. Define, implement, and monitor SLOs, SLIs, and error budgets for critical services. Develop observability strategies using Dynatrace, OpenTelemetry, distributed tracing, metrics, logs, dashboards, and alerting. Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure, and customer experience health. Analyze production telemetry to identify performance bottlenecks, reliability risks, and capacity constraints proactively. Lead incident response for high-severity events and coordinate cross-functional restoration and communications. Perform and facilitate RCAs with corrective and preventive actions tracked to completion. Automate repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls. Partner with development teams to embed reliability and observability across the SDLC. Design, develop, and execute automated regression testing to validate stability and performance after changes. Review test coverage and reliability validation to ensure comprehensive risk mitigation. Create, maintain, and improve Terraform-based IaC for provisioning, configuration, and standardization. Support and optimize Microsoft Azure environments, including App Services, resource management, scaling, and deployment automation. Use Azure Monitor, Application Insights, and Log Analytics to improve visibility and reliability. Drive performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness. Establish operational readiness standards and enforce requirements before production deployments. Review architectures and recommend improvements for resiliency, efficiency, and cloud optimization. Lead capacity planning, performance tuning, and workload optimization across production environments. Develop and maintain runbooks, incident playbooks, knowledge articles, and SOPs. Partner with engineering, infrastructure, cybersecurity, architecture, and support teams on cross-functional improvements. Communicate system health, reliability trends, risks, and remediation to technical and business stakeholders. Present initiatives, metrics, and recommendations at reviews, forums, and leadership meetings. Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational practices. Adhere to risk and regulatory standards and identify issues requiring escalation. Promote a culture of belonging consistent with company values and maintain internal control standards. Required Skills & Qualifications: Associate's degree with 7+ years in SRE, Cloud, Systems, Infrastructure Engineering, DevOps, or Application Support. Bachelor's degree with 5+ years. Or 9+ years combined education and experience with 5+ years in a technology engineering role. Hands-on observability and monitoring experience with Dynatrace, OpenTelemetry, distributed tracing, metrics, centralized logging, alerting, and dashboards. Proven ability to design and execute automated regression testing frameworks and suites. Strong proficiency with Terraform and Infrastructure as Code practices. Experience with CI/CD, deployment automation, and operational tooling. Expertise in production monitoring, incident management, and troubleshooting of distributed systems. Understanding of application performance management and modern cloud-native architectures. Preferred Skills: Microsoft Azure expertise, including App Services, Resource Groups, networking, scaling and optimization, deployment and release management, and application lifecycle management. Use of Azure Monitor, Application Insights, Log Analytics, dashboards, and alerting. SRE practices such as SLOs, SLIs, error budgets, incident/problem management, RCA, and reliability automation. Experience with performance tuning, capacity planning, proactive issue detection, and observability-driven improvements. Automated recovery mechanisms and self-healing solutions, resiliency patterns, DR planning, and high-availability architectures. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to modern technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
Site Reliability Engineer- W2 Role
RSA TECH GROUP Dallas, Texas
Site Reliability Engineer- W2 Role Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts / AWS, Terraform, MongoDB, Spring Boot Framework, Kafka, RESTful API, Camunda
08/04/2026
Full time
Site Reliability Engineer- W2 Role Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts / AWS, Terraform, MongoDB, Spring Boot Framework, Kafka, RESTful API, Camunda
Observability Operations Engineer
BC Forward Phoenix, Arizona
Job Title: Observability Operations Engineer Location: Phoenix, AZ (Hybrid) Duration: Temp - 12 months Pay Range: $50/hr to $60/hr (W2) Job ID: 407523 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is desirable. The role will operate at approximately 20% automation and 80% operations and will ensure availability, scalability, operational excellence, and continuous improvement of monitoring and logging platforms that support mission-critical applications. Work Schedule & Coverage: Onsite presence 3 days per week. Standard shifts: 9:00 AM-6:00 PM or 10:00 AM-6:30 PM. Work 1 weekend day every 2-3 weeks to provide coverage. Responsibilities: Administer and optimize Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms. Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions. Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning. Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, RUM, DEM, Davis AI, and APM. Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI. Develop dashboards, alerts, reports, and executive operational metrics. Support Linux infrastructure and Kubernetes environments, including Docker, OpenShift, or Rancher. Implement observability best practices using OpenTelemetry for distributed tracing, metrics, logs, and events. Perform root cause analysis for production incidents using observability platforms. Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams. Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible, including AI-assisted automation. Participate in incident, problem, change, and release management processes. Drive platform upgrades, patching, security compliance, and operational governance. Improve platform reliability through automation, self-healing, and AI-assisted operations. Required Skills & Qualifications: Dynatrace, Splunk Enterprise, OpenSearch, and Elasticsearch administration. Grafana, Prometheus, Kibana, Jaeger, and OpenTelemetry. Kubernetes and Linux administration with Docker, OpenShift, or Rancher. Networking fundamentals including TCP/IP, DNS, load balancers, and firewalls. AWS, Azure, or GCP with CI/CD pipelines, Git, Terraform, Ansible, and REST APIs. Scripting with Python and Bash/Shell. PowerShell preferred. 6-10+ years in IT infrastructure or observability operations with 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch. Strong Linux system administration and production support experience. Excellent troubleshooting, analytical, communication, and stakeholder management skills. Preferred Skills: Kafka exposure. Experience with AI/ML for observability and AIOps. Experience with Dynatrace Davis AI, predictive monitoring, and intelligent alerting. Generative AI tools such as ChatGPT, GitHub Copilot, Amazon Q, or Microsoft Copilot to improve operational efficiency. AI-assisted runbooks, incident summarization, log analysis, and automated ticket enrichment. Knowledge of RAG, vector databases, embeddings, and AI-powered knowledge search. Python with AI frameworks such as LangChain, LangGraph, or OpenAI APIs. Education & Certifications: Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience. Preferred: Dynatrace Associate or Professional, Splunk Enterprise Certified Administrator, Elastic Certified Engineer, Kubernetes CKA/CKAD, AWS/Azure/GCP, ITIL Foundation, and AI/ML or Generative AI certification. Must Haves: Dynatrace, Splunk, OpenSearch/Elasticsearch, and OpenTelemetry. Kubernetes, AI/ML, Grafana, and Sahara. Automation using AI with a focus on operations. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to cutting-edge technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
08/04/2026
Full time
Job Title: Observability Operations Engineer Location: Phoenix, AZ (Hybrid) Duration: Temp - 12 months Pay Range: $50/hr to $60/hr (W2) Job ID: 407523 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is desirable. The role will operate at approximately 20% automation and 80% operations and will ensure availability, scalability, operational excellence, and continuous improvement of monitoring and logging platforms that support mission-critical applications. Work Schedule & Coverage: Onsite presence 3 days per week. Standard shifts: 9:00 AM-6:00 PM or 10:00 AM-6:30 PM. Work 1 weekend day every 2-3 weeks to provide coverage. Responsibilities: Administer and optimize Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms. Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions. Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning. Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, RUM, DEM, Davis AI, and APM. Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI. Develop dashboards, alerts, reports, and executive operational metrics. Support Linux infrastructure and Kubernetes environments, including Docker, OpenShift, or Rancher. Implement observability best practices using OpenTelemetry for distributed tracing, metrics, logs, and events. Perform root cause analysis for production incidents using observability platforms. Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams. Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible, including AI-assisted automation. Participate in incident, problem, change, and release management processes. Drive platform upgrades, patching, security compliance, and operational governance. Improve platform reliability through automation, self-healing, and AI-assisted operations. Required Skills & Qualifications: Dynatrace, Splunk Enterprise, OpenSearch, and Elasticsearch administration. Grafana, Prometheus, Kibana, Jaeger, and OpenTelemetry. Kubernetes and Linux administration with Docker, OpenShift, or Rancher. Networking fundamentals including TCP/IP, DNS, load balancers, and firewalls. AWS, Azure, or GCP with CI/CD pipelines, Git, Terraform, Ansible, and REST APIs. Scripting with Python and Bash/Shell. PowerShell preferred. 6-10+ years in IT infrastructure or observability operations with 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch. Strong Linux system administration and production support experience. Excellent troubleshooting, analytical, communication, and stakeholder management skills. Preferred Skills: Kafka exposure. Experience with AI/ML for observability and AIOps. Experience with Dynatrace Davis AI, predictive monitoring, and intelligent alerting. Generative AI tools such as ChatGPT, GitHub Copilot, Amazon Q, or Microsoft Copilot to improve operational efficiency. AI-assisted runbooks, incident summarization, log analysis, and automated ticket enrichment. Knowledge of RAG, vector databases, embeddings, and AI-powered knowledge search. Python with AI frameworks such as LangChain, LangGraph, or OpenAI APIs. Education & Certifications: Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience. Preferred: Dynatrace Associate or Professional, Splunk Enterprise Certified Administrator, Elastic Certified Engineer, Kubernetes CKA/CKAD, AWS/Azure/GCP, ITIL Foundation, and AI/ML or Generative AI certification. Must Haves: Dynatrace, Splunk, OpenSearch/Elasticsearch, and OpenTelemetry. Kubernetes, AI/ML, Grafana, and Sahara. Automation using AI with a focus on operations. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to cutting-edge technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.
Mastercard
Site Reliability Engineer II
Mastercard O Fallon, Missouri
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team - one that makes better decisions, drives innovation, and delivers better business results. About the Role The Payment Network Business Operations team is seeking a highly motivated and experienced Site Reliability Engineer II (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor. The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle. As part of the Business Operations team, you will: Work independently on elements of projects/processes within the Site Reliability Engineering area by applying intermediate/practical knowledge and area best practices to meet organizational standards of quality and excellence. Support the implementation and maintenance of high-availability systems to ensure operational stability. Assist in evaluating operational needs and developing technical solutions under guidance. Contribute to automation and scripting projects to streamline routine operational tasks. Troubleshoot and resolve basic to moderate system issues, escalating more complex problems as needed. Document operational procedures and shares knowledge with team members. Participate in quality checks and reviews to ensure system stability and reliability. Utilize experience and a comprehensive understanding of area processes and tools to make minor adjustments or enhancements to resolve identifiable issues. May manage smaller project/initiatives as an experienced individual contributor with specialized knowledge within the Site Reliability Engineering area. Role qualifications: The ideal candidate will apply the following skills independently in routine and moderately complex situations, requiring occasional guidance typically only in unfamiliar or highly complex scenarios. They will demonstrate growing consistency and reliability in applying the skills. Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement. Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar. Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability. Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency. Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations. Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations. Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage. IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs. Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability. Specific Tool / Systems: Strong knowledge of ITSM practices, observability, and monitoring using tools such as Splunk and Dynatrace Experience operating and supporting applications on PCF and AWS platforms Proven ability to implement CI/CD pipelines using Jenkins, Bitbucket, and XLR for automated build and release management Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly. Corporate Security Responsibility All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: Abide by Mastercard's security policies and practices; Ensure the confidentiality and integrity of the information being accessed; Report any suspected information security violation or breach, and Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines. In line with Mastercard's total compensation philosophy and assuming that the job will be performed in the US, the successful candidate will be offered a competitive base salary and may be eligible for an annual bonus or commissions depending on the role. The base salary offered may vary depending on multiple factors, including but not limited to location, job-related knowledge, skills, and experience. Mastercard benefits for full time (and certain part time) employees generally include: insurance (including medical, prescription drug, dental, vision, disability, life insurance); flexible spending account and health savings account; paid leaves (including 16 weeks of new parent leave and up to 20 days of bereavement leave); 80 hours of Paid Sick and Safe Time, 25 days of vacation time and 5 personal days . click apply for full job details
08/04/2026
Full time
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team - one that makes better decisions, drives innovation, and delivers better business results. About the Role The Payment Network Business Operations team is seeking a highly motivated and experienced Site Reliability Engineer II (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor. The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle. As part of the Business Operations team, you will: Work independently on elements of projects/processes within the Site Reliability Engineering area by applying intermediate/practical knowledge and area best practices to meet organizational standards of quality and excellence. Support the implementation and maintenance of high-availability systems to ensure operational stability. Assist in evaluating operational needs and developing technical solutions under guidance. Contribute to automation and scripting projects to streamline routine operational tasks. Troubleshoot and resolve basic to moderate system issues, escalating more complex problems as needed. Document operational procedures and shares knowledge with team members. Participate in quality checks and reviews to ensure system stability and reliability. Utilize experience and a comprehensive understanding of area processes and tools to make minor adjustments or enhancements to resolve identifiable issues. May manage smaller project/initiatives as an experienced individual contributor with specialized knowledge within the Site Reliability Engineering area. Role qualifications: The ideal candidate will apply the following skills independently in routine and moderately complex situations, requiring occasional guidance typically only in unfamiliar or highly complex scenarios. They will demonstrate growing consistency and reliability in applying the skills. Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement. Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar. Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability. Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency. Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations. Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations. Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage. IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs. Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability. Specific Tool / Systems: Strong knowledge of ITSM practices, observability, and monitoring using tools such as Splunk and Dynatrace Experience operating and supporting applications on PCF and AWS platforms Proven ability to implement CI/CD pipelines using Jenkins, Bitbucket, and XLR for automated build and release management Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly. Corporate Security Responsibility All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: Abide by Mastercard's security policies and practices; Ensure the confidentiality and integrity of the information being accessed; Report any suspected information security violation or breach, and Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines. In line with Mastercard's total compensation philosophy and assuming that the job will be performed in the US, the successful candidate will be offered a competitive base salary and may be eligible for an annual bonus or commissions depending on the role. The base salary offered may vary depending on multiple factors, including but not limited to location, job-related knowledge, skills, and experience. Mastercard benefits for full time (and certain part time) employees generally include: insurance (including medical, prescription drug, dental, vision, disability, life insurance); flexible spending account and health savings account; paid leaves (including 16 weeks of new parent leave and up to 20 days of bereavement leave); 80 hours of Paid Sick and Safe Time, 25 days of vacation time and 5 personal days . click apply for full job details
Mastercard
Senior Site Reliability Engineer
Mastercard O Fallon, Missouri
Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you'll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth.
08/01/2026
Full time
Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you'll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth.
Mastercard
Lead Site Reliability Engineer
Mastercard O Fallon, Missouri
Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you'll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment.
08/01/2026
Full time
Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you'll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board