Site Reliability / Operations Engineer

  • Vantor
  • Herndon, Virginia
  • 08/25/2026
Full time Information Technology Telecommunications Python

Job Description

Vantor seeks a Site Reliability / Operations Engineer to design, automate, and operate secure, scalable cloud infrastructure for modern web and data platforms in a TS/SCI environment. You'll build and manage CI/CD pipelines, infrastructure as code, containers, and Kubernetes; implement monitoring, logging, and alerting; and lead incident response to keep services reliable and performant. You'll collaborate closely with developers and clients in an agile, low ego culture, champion automation and best practices, and help evolve our SRE and DevOps standards across impactful digital transformation projects.

Responsibilities

  • Design, build, and maintain secure, scalable cloud infrastructure for client applications.
  • Implement and manage infrastructure as code, CI/CD pipelines, and automation tooling.
  • Operate and optimize containerized and Kubernetes-based platforms.
  • Set up and refine monitoring, logging, and alerting for reliability and performance.
  • Lead and participate in incident response, troubleshooting, and on-call rotations.
  • Collaborate with development teams to embed SRE and Dev
  • Ops best practices.
  • Harden systems and configurations to meet TS/SCI security and compliance requirements.
  • Continuously improve reliability, scalability, and cost efficiency of services.
  • Document architectures, runbooks, and operational procedures.
  • Contribute to evolving SRE standards, tooling, and knowledge sharing across teams.

Required Skills

  • Site Reliability Engineering (SRE)
  • Linux/Unix administration
  • Cloud platforms (AWS/Azure/GCP)
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • CI/CD pipelines
  • Containerization (Docker)
  • Kubernetes operations
  • Monitoring and observability (Prometheus/Grafana/ELK)
  • Scripting (Python/Bash)
  • Incident response and on-call operations