aidoosDenver, CO
Site Reliability Engineer
Site Reliability Engineer
Site Reliability Engineer
aidoosDenver, CO
yesterday
Computer Systems Design ServicesOther Computer Related ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Apply for this role →Site Reliability Engineer JumpCloud is seeking a Site Reliability Engineer (SRE) to join our Infrastructure & Reliability Engineering team. This role focuses on ensuring platform resilience through high availability, performance, and operational maturity of our production systems, primarily in cloud environments like AWS and GCP. The SRE will be responsible for designing automation, building observability frameworks, and implementing reliability best practices across our microservices.
Key Responsibilities:
Design, deploy, and maintain the reliability and performance of JumpCloud systems and APIs.
Operationalize SLIs, SLOs, and error budgets in collaboration with application teams.
Develop end-to-end observability across microservices using tools like Datadog.
Manage production Kubernetes clusters using GitOps delivery workflows.
Write automation tools and scripts in Python or Go to eliminate operational toil.
Required Skills & Qualifications 5+ years of professional experience in SRE, DevOps, or Platform Engineering.
Proficiency in Python or Go for developing SRE tools and automation.
Experience with Kubernetes operations and GitOps pipelines.
Solid knowledge of Infrastructure as Code, especially with Terraform.
Direct experience with AWS or GCP cloud workloads.
Also on the board Same function, level within a rung
Level
Senior
Location
Denver, CO
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday