DevOps Engineer IV- 4P/702

4P Consulting Inc.
Atlanta, GA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Continuous Integration DevOps Reliability Engineering Prometheus Grafana Software Troubleshooting Kubernetes Microservices

Job description

We are seeking an experienced DevOps Engineer IV / Site Reliability Engineer (SRE) with strong hands-on experience in observability, telemetry, monitoring, and service reliability . The ideal candidate will have deep knowledge of Grafana, OpenTelemetry (OTEL), PromQL, and application/system instrumentation .

This role will partner with engineering, operations, and application teams to improve service reliability, telemetry quality, alerting maturity, and operational visibility across complex environments., * Design, implement, and support monitoring and observability solutions.

  • Build dashboards, alerts, and telemetry solutions using Grafana and related tools.
  • Implement OpenTelemetry standards for application and system instrumentation.
  • Write and optimize PromQL queries for monitoring and reliability insights.
  • Improve alerting quality, reduce noise, and create actionable alerts.
  • Troubleshoot application and infrastructure issues using logs, metrics, and traces.
  • Support incident response, root cause analysis, and reliability improvements.
  • Collaborate with engineering, operations, and application teams.

Requirements

Do you have experience in Technical troubleshooting support?, * Strong experience as a DevOps Engineer, SRE, Observability Engineer, or similar role.

  • Hands-on experience with Grafana, OpenTelemetry, and PromQL.
  • Experience with application and system instrumentation.
  • Strong understanding of logs, metrics, traces, alerting, and service reliability.
  • Ability to design monitoring solutions across complex environments.
  • Strong troubleshooting, analytical, communication, and collaboration skills., * Experience with Prometheus, Loki, Tempo, Kubernetes, containers, cloud platforms, or microservices.
  • Familiarity with CI/CD, automation, infrastructure-as-code, incident response, SLIs, SLOs, and reliability metrics., DevOps, SRE, Observability, Grafana, OpenTelemetry, OTEL, PromQL, Prometheus, Monitoring, Alerting, Logs, Metrics, Traces, Instrumentation, Incident Response, Root Cause Analysis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

Videos

See all

Related articles

See all