Observability Cloud) Engineer

Drevol LLC
Malvern, PA, United States
about 1 month ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Microsoft Azure Cloud Computing Distributed Systems Monitoring of Systems Reliability Engineering Data Logging Google Cloud Spring Cloud System Availability Deployment Automation
+2 more
Splunk Microservices

Job description

We are seeking an experienced Observability Engineer to support a strategic

observability platform migration initiative.

This role will be responsible for playing a core part in executing the migration of observability artifacts: SLOs, dashboards, alerts, and

monitoring workflows from Honeycomb to Splunk Observability Cloud., * Assist in the migration of observability assets and telemetry workflows from

Honeycomb to Splunk Observability Cloud.

  • Assess, document, and translate existing dashboards, alerts, SLOs, and other

monitoring configurations.

  • Design and implement observability solutions utilizing OpenTelemetry standards

and best practices.

  • Partner with enterprise engineering, SRE, and platform teams to validate

telemetry quality, monitoring coverage, and operational readiness.

  • Identify opportunities to improve observability maturity, reliability, and operational

efficiency during the migration.

  • Develop documentation, knowledge transfer materials, and migration playbooks.
  • Communicate progress, risks, dependencies, and recommendations to technical

and non-technical stakeholders., * The successful candidate will help ensure a seamless migration from Honeycomb ton Splunk Observability Cloud while maintaining observability coverage, improving telemetry quality, and enabling engineering teams to effectively monitor and support production systems.

  • This role offers an opportunity to drive a highly visible observability transformation initiative and help establish the foundation for future reliability and operational excellence efforts.

Requirements

The ideal candidate combines deep technical expertise with robust communication and

collaboration skills and has hands-on experience designing, implementing, and

operating modern observability solutions using OpenTelemetry, Honeycomb, and

Splunk Observability Cloud., * Robust technical, analytical, and communication skills.

  • Experience working in Site Reliability Engineering (SRE), Observability

Engineering, Platform Engineering, or a related discipline.

  • Hands-on experience with Splunk Observability Cloud, including dashboards,

detectors, APM, infrastructure monitoring, and related capabilities.

  • Hands-on experience with Honeycomb, including telemetry analysis,

dashboards, and operational workflows.

  • Experience implementing and supporting OpenTelemetry instrumentation,

collection, and telemetry pipelines.

  • Experience working with distributed systems, cloud-native applications, APIs,

and microservices environments.

  • Ability to collaborate effectively across engineering, operations, and leadership

teams.

  • Experience documenting technical solutions and communicating complex
  • concepts to diverse audiences.

Preferred Qualifications

  • Experience leading or participating in enterprise-scale observability platform

migrations.

  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Familiarity with CI/CD pipelines and automated deployment practices.
  • Knowledge of modern monitoring, tracing, logging, and metrics best practices.
  • Experience supporting production environments with high availability and

reliability requirements.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

6:10 min

Unlocking free learning credits via Google Cloud Innovators

Asrar Asrar · World Congress 2024

1:10 min

Exposing sensitive information through partial search logs

Dennis Schulz Dennis Schulz +1 · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

4:42 min

Container hosting options available on Google Cloud Platform

Federico Fregosi · World Congress 2022

1:56 min

Discovering incidents using logs, metrics, and traces

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all