Principal, Cloud Engineer - Observability in Derry Village

Energy Jobline
Derry, NH, United States
2 months ago
Apply on www.energyjobline.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Computing Platforms Microsoft Azure Cloud Engineering Continuous Integration Software Debugging Distributed Systems Identity and Access Management Python (Programming Language) Machine Learning
+11 more
Open Source Technology Performance Tuning Prometheus Software Engineering Datadog System Availability Grafana Multi-Cloud Cloudformation Kubernetes Terraform

Job description

You will work in a collaborative, transparent, and innovation-driven environment where engineering excellence, continuous learning, and open-source contribution are core to how we operate. This is a high-impact role where your expertise will influence platform architecture, engineering practices, and the developer experience across the organization.

What You’ll Do

  • Lead the design and implementation of cloud- observability platforms across AWS and Azure environments
  • Develop and operate highly scalable systems on AWS (EKS, core services) with strong focus on reliability, performance, and automation
  • Own the end-to-end lifecycle of observability tooling, including hosting, maintenance, scaling, and optimization
  • Drive adoption of OpenTelemetry (OTel) standards for metrics, traces, logs, and profiling
  • Build and enhance platform capabilities using Python and/or Go
  • Architect and optimize CI/CD pipelines enabling rapid, secure, and reliable deployments
  • Collaborate with cross-functional teams to improve system visibility, debugging capabilities, and performance insights
  • Define and promote best practices for cloud engineering, observability, and platform reliability
  • Mentor engineers and provide technical leadership across squads

Requirements

We are seeking a highly motivated Principal Cloud Engineer to join our Observability Platform team within Fidelity Architecture and Engineering. In this role, you will help design, build, and operate scalable, cloud- observability solutions that support our most critical digital services., * 10+ years of software engineering or cloud engineering experience

  • Deep expertise in AWS cloud stack, especially:
  • EKS (Kubernetes on AWS)
  • Core services (IAM, EC2, networking, storage, etc.)
  • Strong experience working with Kubernetes and cloud- ecosystems
  • Hands-on experience with observability tools/platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry, etc.), including hosting and operational ownership
  • Proficiency in Python and/or Go for platform and tooling development
  • Strong understanding of CI/CD practices and tools
  • Experience working with Infrastructure as Code (Terraform, CloudFormation, etc.)
  • Familiarity with the Azure cloud stack and hybrid/multi-cloud environments
  • Working knowledge of OpenTelemetry (OTel) concepts and implementation

Bonus Skills

  • Experience building or operating large-scale internal platforms
  • Exposure to eBPF-based observability or advanced profiling solutions
  • Experience integrating observability across multi-region / multi-cloud environments
  • Experience or interest in applying AI/ML techniques to observability (e.g., anomaly detection, predictive insights, intelligent alerting, or AIOps)
  • Active participation in or contributions to open-source projects
  • Strong background in performance optimization and distributed systems

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all