Our client is seeking a Principal Site Reliability Engineer, Infrastructure Observability to help build and advance a world-class SRE function focused on observability, reliability, scalability, resilience, and automation across complex cloud and on-premises environments. This role will be instrumental in driving operational excellence through modern engineering practices, automation, and best-in-class observability tooling. The ideal candidate brings deep expertise in cloud infrastructure, DevOps, SRE methodologies, incident management, automation, and infrastructure monitoring. They will serve as a strategic leader and hands-on technical expert, helping shape reliability practices across a highly distributed enterprise environment.
Responsibilities: Lead the design and implementation of reliability-focused solutions that improve system availability and minimize service disruptions. Drive SRE best practices across the organization, promoting a culture of automation, operational excellence, and continuous improvement. Champion observability initiatives, including monitoring, alerting, logging, and application performance management (APM).Conduct incident trend analysis and lead initiatives to reduce recurring technology failures. Facilitate blameless post-mortems and reliability reviews to strengthen operational maturity. Develop automated solutions to proactively prevent incidents and accelerate remediation efforts. Create unified visibility across technology platforms to identify risks, redundancies, and optimization opportunities. Partner with engineering, infrastructure, security, and business stakeholders to improve service reliability and operational performance. Contribute to target-state architecture and the long-term evolution of the technology ecosystem. Mentor engineers and help establish standards, frameworks, and operational practices across the organization.
Requirements: Bachelor's degree or equivalent combination of education and experience.10+ years of experience designing, building, and operating enterprise infrastructure solutions with significant organizational impact.5+ years of hands-on experience with Amazon Web Services (AWS).5+ years building, leading, or supporting Site Reliability Engineering (SRE) and/or DevOps functions. Experience implementing and operating chaos engineering practices at scale. Proven success leading strategic technology and transformation initiatives. Strong scripting, systems administration, and infrastructure automation experience. Demonstrated ability to leverage automation to improve reliability and reduce operational risk. Proficiency in multiple programming languages, including Python, Java, Go, Node.js, and/or .NET Core. Strong database experience with SQL Server, PostgreSQL, MySQL, or similar platforms. Deep understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics, and reliability measurement frameworks. Experience implementing and managing Error Budgets. Strong incident response, root cause analysis, and service recovery expertise. Experience standardizing observability, monitoring, logging, and dashboarding across enterprise environments. Hands-on experience with tools such as New Relic, Splunk, Elastic Stack, Prometheus, Grafana, SolarWinds, and cloud-native monitoring platforms. Experience with infrastructure automation and cloud management tools including Terraform, Ansible, Vault, and Vagrant. Ability to influence stakeholders across technical and business teams. Strong communication, leadership, and mentoring skills. Willingness to participate in on-call rotations and support critical production environments.
Preferred Qualifications: Cloud, DevOps, or Site Reliability Engineering certifications. Working knowledge of Microsoft Azure. Experience within highly regulated, large-scale enterprise environments. Why Apply? This is an opportunity to play a pivotal role in shaping the reliability, observability, and operational excellence strategy for a leading enterprise technology organization. You'll work alongside senior engineering leaders, influence large-scale technology initiatives, and drive the adoption of modern SRE practices across a complex global environment.

Also on the board Same function, level within a rung

Level

Lead

Location

Owings Mills, MD

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Apply for this role →