Platform Reliability, Availability, Serviceability

  • Qualcomm
  • Santa Clara, California
  • 09/03/2026
Full time Information Technology Telecommunications

Job Description

Qualcomm seeks a Platform Reliability, Availability, Serviceability engineer to ensure always on performance of large scale telecom and media platforms. You will design and implement highly available, fault tolerant services, build monitoring and alerting, and drive incident response for 5G and connectivity solutions. Partner with software, hardware, and SRE teams to improve resiliency, automate recovery, and optimize SLAs. Analyze production issues, perform root cause analysis, and implement long term fixes. This role suits engineers who thrive in fast paced, innovative environments and enjoy solving complex distributed systems challenges.

Responsibilities

  • Design and maintain highly available, fault tolerant telecom and media platforms
  • Implement monitoring, logging, and alerting for large scale distributed systems
  • Lead and participate in incident response, troubleshooting, and post mortems
  • Perform root cause analysis and implement long term reliability fixes
  • Develop automation for deployment, recovery, and scaling of services
  • Collaborate with software, hardware, SRE, and network teams on resiliency improvements
  • Define and track SLAs, SLOs, and SLIs for critical services
  • Optimize performance, capacity, and reliability of 5
  • G and connectivity platforms
  • Contribute to reliability architecture, tooling, and best practices
  • Document systems, runbooks, and reliability standards

Required Skills

  • Site Reliability Engineering (SRE)
  • Distributed systems design
  • Linux administration
  • Cloud platforms (AWS, GCP, or Azure)
  • Kubernetes and containers
  • Monitoring and observability
  • Incident response and troubleshooting
  • Automation and scripting (Python, Bash)
  • Networking and TCP/IPCI/CD pipelines