Senior Site Reliability Engineer

  • Mastercard
  • O Fallon, Missouri
  • 08/16/2026
Full time Information Technology Telecommunications Python

Job Description

Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you'll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth.

Responsibilities

  • Design and maintain highly available, scalable, and secure cloud infrastructure for Mastercard's IT & Data Management platforms.
  • Build and improve automation for deployments, configuration management, and infrastructure provisioning.
  • Implement and refine monitoring, logging, and alerting to ensure service reliability and rapid incident detection.
  • Lead and participate in incident response, root cause analysis, and post-incident reviews to drive continuous improvement.
  • Partner with development and data teams to embed SRE best practices, including SLIs/SLOs and capacity planning.
  • Optimize system performance and cost efficiency across distributed, cloud-native environments.
  • Develop tools and scripts to reduce toil and improve operational excellence.
  • Contribute to security, compliance, and governance standards within production environments.

Required Skills

  • Site Reliability Engineering (SRE)
  • Cloud platforms (AWS, Azure, or GCP)
  • Kubernetes and containerization (Docker)
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • CI/CD pipelines (Jenkins, Git
  • Hub Actions, Git
  • Lab CI)
  • Linux systems administration
  • Monitoring and observability (Prometheus, Grafana, Datadog, New Relic)
  • Scripting/programming (Python, Go, Bash)
  • Distributed systems and microservices
  • Incident management and on-call operations