Remote Observability Engineer
Role details
Job location
Tech stack
Job description
We are seeking a detail-oriented and technically skilled Observability Engineer to join our Reliability Engineering & Operations (REO) team. In this role, you will be responsible for configuring, maintaining, and continuously improving the observability stack - spanning monitoring, dashboarding, alerting, and application performance management (APM) across our AWS-hosted production environments. You will lead the standup of our observability capabilities during an active proof-of-concept (PoC) period, establishing the tooling, instrumentation patterns, and alerting standards that the broader REO organization will rely on. This is a practitioner role for someone who is passionate about making complex systems legible - turning raw telemetry into actionable insight.
Requirements
2-4 years of experience in an observability, monitoring, SRE, or DevOps engineering role with a strong focus on instrumentation and telemetry
-
Healthcare experience
-
Hands-on experience with AWS observability tooling including CloudWatch (metrics, logs, alarms, dashboards) and X-Ray (distributed tracing)
-
Experience configuring and administering at least one third-party APM or observability platform (e.g., Datadog, Dynatrace, New Relic, Grafana, or similar)
-
Working knowledge of log management and aggregation (e.g., CloudWatch Logs, OpenSearch, Splunk, or similar)
-
Experience building operational dashboards that serve diverse audiences including NOC, engineering, and leadership
-
Familiarity with distributed systems, microservices architectures, and the instrumentation challenges they present
-
Scripting proficiency in Python, Bash, or similar for automation and telemetry configuration
-
Strong analytical skills and a methodical approach to signal-to-noise optimization in alerting
-
Experience working in a HIPAA-regulated or compliance-driven environment - AWS certifications (Cloud Practitioner, SysOps Administrator, or equivalent)
-
Experience with OpenTelemetry (OTel) for vendor-agnostic instrumentation
-
Familiarity with infrastructure-as-code tooling (Terraform, CloudFormation) for managing observability configuration as code
-
Experience supporting Java or .NET-based enterprise SaaS application observability
-
Exposure to SLI/SLO frameworks and reliability engineering practices
-
Experience in healthcare IT or payer technology environments