> Markdown version of [/jobs/ext/3018163-expert-automation-observability-engineer](https://www.wearedevelopers.com/jobs/ext/3018163-expert-automation-observability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Expert Automation & Observability Engineer - **Company:** Ensono - **Location:** Olympia, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Databases, Continuous Integration, Data Retention, Linux, DevOps, Github, Tivoli Management Framework, Python (Programming Language), Windows Servers, Citrix Systems, OpenShift, Role-Based Access Control, Red Hat Enterprise Linux, Reliability Engineering, Ansible, Prometheus, Load Balancing, Grafana, Mttr, Multi-Cloud, HybridCloud, Gitlab, Containerization, Git Flow, Kubernetes, Information Technology, Influxdb, SolarWinds (Software), Terraform, Splunk, Webhooks, Docker, Jenkins, Servicenow, Vmware, Microservices - **Published:** September 20, 2026 - **Apply:** https://dejobs.org/x/x/767725A542F34DD7BF9B6BA987337EC3/job/ ## About the Role * 12+ years of total IT experience , with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment. * Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms. * Hands-on expertise in building scalable, secure telemetry pipelines and time-series databases. * Extensive experience leading FOAK rollouts and complex vendor/operations transition (KT) programs. * Preferred Certifications: CKA (Certified Kubernetes Administrator), Cloud Architect (AWS/Azure), or specific APM/Observability vendor certifications. ## Description We are seeking an Expert Observability Engineer to serve as the strategic technical lead and architect for our enterprise Observability, APM, and Telemetry ecosystems. You will lead the transformation from decentralized, reactive monitoring to a unified, automated, and proactive observability framework. Operating across hybrid cloud, Kubernetes, and legacy environments, you will design scalable architectures, drive Site Reliability Engineering (SRE) practices, and lead First-of-a-Kind (FOAK) technology implementations to ensure maximum service reliability., * Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB . * Lead First-of-a-Kind (FOAK) implementations-evaluating new observability tech and converting them into secure, repeatable, production-ready patterns. * Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management. 2. SRE & Service Reliability * Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets. * Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA). * Drastically reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping. 3. Platform Engineering & Automation * Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps . * Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors. * Integrate observability platforms seamlessly with ITSM (ServiceNow), Netcool, and CI/CD pipelines. 4. Cloud-Native & Kubernetes Observability * Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP). * Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies. * Ensure secure-by-design telemetry pipelines (RBAC, TLS, secrets management, and image scanning). 5. Technical Leadership & Transition Management * Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24×7 teams. * Mentor cross-functional engineering teams and influence enterprise technology roadmaps. Required Technical Stack * Observability & APM: IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, InfluxDB. * Legacy/Traditional Monitoring: SolarWinds, Netcool, Elastic/Splunk. * Cloud & Containerization: Kubernetes, Docker, OpenShift, AWS/Azure/GCP. * Infrastructure: Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies. * Automation & DevOps: Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins). * ITSM/Operations: ServiceNow, ITIL 4, advanced Major Incident Management. ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Generating code with Angular schematics](https://www.wearedevelopers.com/videos/129-generating-code-with-angular-schematics) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)