Remote Senior Monitoring and Observability Engineer

Everforth Apex
Washington, DC, United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Application Programming Interfaces (APIs) Amazon Web Services Application Performance Management Microsoft Azure Cloud Computing Linux Monitoring of Systems Intrusion Detection and Prevention Metadata OpenShift Ansible
+14 more
Systems Integration Virtual Machines Virtualization Technology Datadog Scripting Software Troubleshooting Containerization Kubernetes SolarWinds (Software) Terraform Splunk New Relic (SaaS) Dynatrace Servicenow

Job description

The Senior Monitoring and Observability Engineer will support engineering, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and container platforms. This role is responsible for monitoring coverage, platform integration, agent deployment, tagging, dashboards, alerting, logs, APM, and synthetic monitoring. The engineer partners with Operations and engineering teams to improve visibility, alert quality, incident detection, and operational reliability., * Engineer, operate, maintain, and continuously improve the enterprise monitoring and observability platform.

  • Assess monitoring coverage across enterprise systems and applications to identify visibility gaps.
  • Maintain monitoring coverage across Windows, Linux, cloud, OpenShift/Kubernetes, and other environments.
  • Support monitoring and observability for Red Hat OpenShift and virtual machine workloads.
  • Configure and troubleshoot monitoring agents, integrations, collectors, and APIs.
  • Build and maintain consistent tagging, metadata, dashboards, and alerts.
  • Automate monitoring deployment and configuration using Ansible, APIs, or similar technologies.
  • Develop and maintain integrations between observability platforms and ServiceNow.
  • Use monitoring data to troubleshoot complex performance and availability issues.
  • Correlate telemetry to identify service degradation and recurring technical issues.
  • Partner with teams to improve monitoring coverage, alert quality, and operational response.
  • Analyze telemetry and trends to identify capacity risks and opportunities for improvement.
  • Develop performance, availability, and capacity reporting for stakeholders.
  • Maintain monitoring standards, technical documentation, and operational procedures.

Requirements

  • BS degree and 8-12 years of prior relevant experience, or Master’s degree with 6-10 years of prior relevant experience. Additional relevant experience may be considered in lieu of a degree.
  • Hands-on experience engineering and operating enterprise monitoring or observability platforms.
  • Strong Datadog experience is preferred; however, experience with platforms like ScienceLogic SL1, SolarWinds, Dynatrace, New Relic, or Splunk Observability will be considered.
  • Production experience monitoring Windows, Linux, and Kubernetes or Red Hat OpenShift environments is required.
  • Experience deploying and troubleshooting monitoring agents, integrations, and dashboards. Experience automating monitoring deployment using Ansible, APIs, or scripting.
  • Experience integrating monitoring platforms with ITSM systems such as ServiceNow.
  • Strong troubleshooting and dependency-analysis skills are necessary.
  • Must have the ability to analyze technical telemetry and communicate findings to stakeholders.

Preferred Qualifications

  • Direct experience engineering or administering Datadog in a large enterprise environment.
  • Experience with application performance monitoring, distributed tracing, or OpenTelemetry.
  • Experience with Datadog APM, Log Management, or Synthetic Monitoring.
  • Experience with Red Hat OpenShift Virtualization or related technologies.
  • Experience monitoring Microsoft Azure or AWS environments.
  • Experience with Terraform or monitoring-as-code approaches.
  • Experience supporting federal agency IT environments.
  • Relevant technical certifications such as Datadog, AWS, Microsoft Azure, Red Hat OpenShift, Terraform, or ITIL are preferred.

About the company

Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:39 min

Streamlining application deployment with agentless Ansible

Goetz Rieger Goetz Rieger · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all