As a Site Reliability Engineering Intern at Mastercard, you will support the design, monitoring, and operation of large-scale, cloud-based systems that power secure, global payments. Working alongside experienced SREs and software engineers, you'll help maintain production environments, assist in incident response, and contribute to root-cause analysis to prevent recurrence. You'll aid in building automation for deployments, monitoring, and alerting, and help document runbooks and processes. This internship offers exposure to cutting-edge technologies, modern DevOps practices, and a collaborative culture focused on innovation and continuous learning.
Responsibilities
- Assist in maintaining and monitoring production and cloud infrastructure
- Support incident response, troubleshooting, and root-cause analysis
- Help improve system reliability, scalability, and performance
- Contribute to automation of deployments, monitoring, and alerts
- Work with engineers to implement reliability best practices
- Document processes, runbooks, and technical findings
- Participate in on-call simulations and reliability exercises
- Collaborate with cross-functional teams in an Agile environment
Required Skills
- Linux system administration basics
- Cloud computing fundamentals (AWS/GCP/Azure)
- Scripting/programming (Python, Java, or Go)
- CI/CD pipelines and Dev
- Ops concepts
- Containers and orchestration (Docker, Kubernetes)
- Monitoring and logging tools (Prometheus, Grafana, ELK, etc.)
- Networking fundamentals (TCP/IP, DNS, HTTP)
- Version control with Git
- Automation and Infrastructure as Code concepts
- Troubleshooting and root-cause analysis