Operational Data & Observability Engineer

NSCALE, LLC
Seattle, WA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$145,000.0 - $180,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Application Performance Management Audit Trail Microsoft Azure Bash Shell Databases System Configuration Extract Transform Load (ETL) Linux DevOps Elasticsearch
+27 more
Python (Programming Language) Log Analysis Operational Data Store Reliability Engineering Ansible Prometheus Software Engineering Datadog Data Logging Scripting Google Cloud System Availability Grafana Mttr Build Management Kubernetes Data Analytics Performance Monitor Cloudwatch Terraform Splunk New Relic (SaaS) Dynatrace Elk Stack Golang Programming Languages Microservices

Job description

We’re looking for an Operational Data & Observability Engineer to build and evolve the monitoring, logging, and observability capabilities that power our production environments. In this role, you’ll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.

You’ll partner closely with DevOps, Site Reliability Engineering (SRE), platform, and software engineering teams to develop modern observability solutions, improve incident response, and enable data-driven operational excellence., * Design and implement enterprise observability strategies across infrastructure, services, and applications.

  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.
  • Build and maintain centralized logging and log analysis pipelines.
  • Implement distributed tracing to improve visibility across microservices and complex application workflows.
  • Establish performance baselines and develop anomaly detection strategies.

Operational Data Engineering

  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.
  • Design and manage operational data pipelines that support monitoring and analytics.
  • Develop APIs and integrations that enable operational data consumption across teams.
  • Ensure data quality, consistency, retention, and cost-efficient storage practices.

Reliability & Operations

  • Troubleshoot production issues using monitoring, logging, and tracing data.
  • Participate in an on-call rotation and support incident response activities.
  • Create and maintain operational documentation, runbooks, and troubleshooting guides.
  • Partner with engineering teams to improve platform reliability, scalability, and operational readiness.
  • Continuously optimize observability infrastructure for performance and resilience., * Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies.
  • Evaluate emerging observability tools and recommend improvements.
  • Automate monitoring deployments, instrumentation, and platform configuration.
  • Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure., Success in this role will be measured by your ability to:
  • Improve platform visibility and operational health.
  • Reduce Mean Time to Resolution (MTTR) during incidents.
  • Increase alert quality while reducing unnecessary noise.
  • Deliver highly available, scalable observability platforms.
  • Improve engineering productivity through actionable monitoring and operational insights.
  • Optimize observability infrastructure performance and cost efficiency.

Work Environment

  • Participate in a rotating on-call schedule to support production environments.
  • Support mission-critical systems with occasional after-hours or incident response responsibilities.
  • Hybrid or remote work arrangements available, depending on business needs.

Requirements

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.
  • Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions.
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.
  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring.
  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem-solving., * Experience supporting microservices-based architectures.
  • Expertise across multiple observability platforms.
  • Experience with incident management, root cause analysis, and post-incident reviews.
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools.
  • Familiarity with eBPF or low-level Linux performance monitoring.
  • Experience building custom telemetry, ETL, or operational data pipelines.
  • Understanding of security monitoring, audit logging, and compliance requirements., You’ll play a critical role in building the operational intelligence that keeps our platforms running at scale. If you’re passionate about observability, automation, reliability, and empowering engineering teams with meaningful operational insights, we’d love to hear from you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all