Site Reliability Engineer JumpCloud is seeking a Site Reliability Engineer (SRE) to join our Infrastructure & Reliability Engineering team. This role focuses on ensuring platform resilience through high availability, performance, and operational maturity of our production systems, primarily in cloud environments like AWS and GCP. The SRE will be responsible for designing automation, building observability frameworks, and implementing reliability best practices across our microservices.
Key Responsibilities: Design, deploy, and maintain the reliability and performance of JumpCloud systems and APIs. Operationalize SLIs, SLOs, and error budgets in collaboration with application teams. Develop end-to-end observability across microservices using tools like Datadog. Manage production Kubernetes clusters using GitOps delivery workflows. Write automation tools and scripts in Python or Go to eliminate operational toil. Required Skills & Qualifications 5+ years of professional experience in SRE, DevOps, or Platform Engineering. Proficiency in Python or Go for developing SRE tools and automation. Experience with Kubernetes operations and GitOps pipelines. Solid knowledge of Infrastructure as Code, especially with Terraform. Direct experience with AWS or GCP cloud workloads.

Also on the board Same function, level within a rung

Level

Senior

Location

Denver, CO

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Apply for this role →