Senior Monitoring and Observability Engineer
Senior Monitoring and Observability Engineer
mastech digitalChicago, IL
yesterday
Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsSoftware Developers
Computer Systems Design ServicesCustom Computer Programming ServicesOther Computer Related Services
Apply for this role →Title: Senior Monitoring and Observability Engineer
Location: Remote
Duration:
Contract to Hire The Senior Monitoring and Observability Engineer will support by engineering, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and container platforms. This hands on role is responsible for monitoring coverage, platform integration, agent deployment, tagging and normalization, dashboards, alerting, logs, APM, synthetic monitoring, automation, and operational integrations. Datadog is the primary enterprise observability platform used in the environment. Strong Datadog experience is preferred, however, candidates with substantial experience engineering and operating other enterprise monitoring or observability platforms will be considered where they demonstrate strong transferable monitoring expertise and the ability to rapidly develop proficiency with new technologies. The engineer partners with Operations and engineering teams to improve visibility, alert quality, incident detection, troubleshooting, performance analysis, and operational reliability across the enterprise.
Primary Responsibilities:
In this Role you will:• Engineer, operate, maintain, and continuously improve the enterprise monitoring and observability platform, including dashboards, monitors, metrics, logs, APM, synthetic monitoring, tagging, integrations, and related capabilities.• Assess monitoring coverage across enterprise systems and applications, identify visibility gaps, and coordinate onboarding or remediation with the appropriate technical teams.• Maintain monitoring coverage across Windows, Linux, cloud, Open Shift/Kubernetes, virtualized, database, network, storage, middleware, and application environments.• Support monitoring and observability for Red Hat Open Shift, Kubernetes, Open Shift Virtualization, and virtual machine workloads running on Open Shift.• Configure and troubleshoot monitoring agents, integrations, collectors, APIs, and related platform components.• Build and maintain consistent tagging, metadata, dashboards, alerts, service health views, and operational reporting.• Automate monitoring deployment, configuration, tagging, onboarding, upgrades, and integrations using Ansible, APIs, scripting, CI/CD, infrastructure-as-code, or similar technologies.• Develop and maintain integrations between observability platforms, Service Now, notification systems, on-call workflows, and other enterprise operational systems.• Use monitoring and observability data to troubleshoot complex performance and availability issues, support incident response and root-cause analysis, and recommend technical remediation.• Correlate infrastructure, application, platform, and dependency telemetry to identify service degradation and recurring technical issues.• Partner with Operations and engineering teams to improve monitoring coverage, alert quality, service visibility, incident detection, escalation, and operational response.• Analyze telemetry and historical trends to identify capacity risks, recurring issues, monitoring gaps, and opportunities for improvement.• Develop actionable performance, availability, capacity, and monitoring coverage reporting for technical and leadership stakeholders.• Maintain monitoring standards, technical documentation, configuration guidance, and operational procedures.
Basic Qualifications:
• BS degree and 8-12 years of prior relevant experience, or Master's degree with 6-10 years of prior relevant experience. Additional relevant experience may be considered in lieu of degree requirements where permitted by contract.• Strong hands-on experience engineering and operating enterprise monitoring or observability platforms.• Strong Datadog experience is preferred; however, substantial experience with Science Logic SL1, Solar Winds, Dynatrace, New Relic, Splunk Observability, Logic Monitor, Prometheus/Grafana, or comparable enterprise platforms will be considered based on demonstrated monitoring and observability engineering expertise.• Demonstrated ability to apply monitoring and observability engineering principles across technologies and rapidly develop proficiency with new platforms.• Production experience monitoring Windows and Linux infrastructure and Kubernetes or Red Hat Open Shift environments.• Experience deploying, configuring, upgrading, and troubleshooting monitoring agents, integrations, dashboards, alerts, tagging, and operational reporting.• Experience automating monitoring deployment or administration using Ansible, APIs, scripting, CI/CD pipelines, infrastructure-as-code, or similar technologies.• Experience integrating monitoring or observability platforms with ITSM systems such as Service Now.• Strong troubleshooting and dependency-analysis skills across infrastructure, applications, networks, platforms, and services.• Ability to analyze technical telemetry, identify monitoring or performance gaps, and translate findings into actionable recommendations.• Ability to communicate technical findings and recommendations to technical teams, project leadership, and customer stakeholders.• Must meet applicable contract citizenship and work authorization requirements and be able to obtain and maintain SEC Public Trust or other required clearance.
Also on the board Same function, level within a rung
Level
Lead
Location
Chicago, IL
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday