Remote Observability Engineer

Insight Global
Burlington, United States of America
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate

Job location

Burlington, United States of America

Tech stack

Java
.NET
Amazon Web Services (AWS)
Application Performance Management
Bash
Health Informatics
Cloud Computing
System Configuration
DevOps
Distributed Systems
Python
Reliability Engineering
Datadog
Scripting (Bash/Python/Go/Ruby)
Grafana
Cloudformation
Cloudwatch
Terraform
Splunk
New Relic (SaaS)
Dynatrace
Microservices

Job description

We are seeking a detail-oriented and technically skilled Observability Engineer to join our Reliability Engineering & Operations (REO) team. In this role, you will be responsible for configuring, maintaining, and continuously improving the observability stack - spanning monitoring, dashboarding, alerting, and application performance management (APM) across our AWS-hosted production environments. You will lead the standup of our observability capabilities during an active proof-of-concept (PoC) period, establishing the tooling, instrumentation patterns, and alerting standards that the broader REO organization will rely on. This is a practitioner role for someone who is passionate about making complex systems legible - turning raw telemetry into actionable insight.

Requirements

2-4 years of experience in an observability, monitoring, SRE, or DevOps engineering role with a strong focus on instrumentation and telemetry

  • Healthcare experience

  • Hands-on experience with AWS observability tooling including CloudWatch (metrics, logs, alarms, dashboards) and X-Ray (distributed tracing)

  • Experience configuring and administering at least one third-party APM or observability platform (e.g., Datadog, Dynatrace, New Relic, Grafana, or similar)

  • Working knowledge of log management and aggregation (e.g., CloudWatch Logs, OpenSearch, Splunk, or similar)

  • Experience building operational dashboards that serve diverse audiences including NOC, engineering, and leadership

  • Familiarity with distributed systems, microservices architectures, and the instrumentation challenges they present

  • Scripting proficiency in Python, Bash, or similar for automation and telemetry configuration

  • Strong analytical skills and a methodical approach to signal-to-noise optimization in alerting

  • Experience working in a HIPAA-regulated or compliance-driven environment - AWS certifications (Cloud Practitioner, SysOps Administrator, or equivalent)

  • Experience with OpenTelemetry (OTel) for vendor-agnostic instrumentation

  • Familiarity with infrastructure-as-code tooling (Terraform, CloudFormation) for managing observability configuration as code

  • Experience supporting Java or .NET-based enterprise SaaS application observability

  • Exposure to SLI/SLO frameworks and reliability engineering practices

  • Experience in healthcare IT or payer technology environments

Apply for this position