Data Dog Cloud Engineer (Observability)

System One
Washington, DC, United States
9 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$131,040.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Amazon Web Services Microsoft Azure Cloud Computing Python (Programming Language) Node.Js OpenShift Datadog ServiceNow IT Service Management Istio Kubernetes
+8 more
Bicep Linkerd (Service Mesh) Terraform Splunk New Relic (SaaS) Dynatrace Servicenow Golang

Job description

  • Build and run observability tooling across metrics, logs, traces/APM, RUM, synthetics, and network monitoring
  • Create and maintain dashboards, monitors, alerts, SLOs/SLIs that teams actually use (high signal, low noise)
  • Instrument applications and services using agents, OpenTelemetry, and language-specific APM tooling (Java, .NET, Python, Node.js, Go)
  • Improve production performance and reliability using telemetry to troubleshoot:
  • latency, saturation, capacity, errors, and dependency issues
  • Partner with cloud/platform/app teams to embed observability into:
  • AWS + Azure
  • Kubernetes/OpenShift
  • Integrate monitoring workflows with ServiceNow, CI/CD pipelines, and on-call/paging processes
  • Establish and enforce telemetry standards (tagging strategy, governance, cost control), System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

Requirements

  • 8+ years in infrastructure/platform engineering (or similar), including 5+ years focused on observability/performance/SRE-type work
  • Hands-on experience operating an observability platform:
  • Datadog preferred, but Dynatrace / New Relic / Splunk Observability / Grafana+Prometheus are also relevant
  • Strong experience with APM + distributed tracing in production (including instrumentation and service mapping)
  • Production monitoring experience with Kubernetes/OpenShift
  • Cloud experience supporting monitoring/telemetry in AWS and/or Azure
  • Bachelor’s degree (or equivalent experience)
  • Ability to support an on-call rotation in a 24x7 environment

Nice-to-haves (stand out if you have them)

  • Experience in federal/government environments (FISMA, FedRAMP, NIST-aligned)
  • Datadog certifications (or comparable)
  • eBPF observability, service mesh telemetry (Istio/Linkerd)
  • Terraform/Bicep/ARM for deploying monitoring as code
  • AWS/Azure/OpenShift/Terraform certifications

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all