> Markdown version of [/jobs/ext/1345665-observability-engineer-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1345665-observability-engineer-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Observability Engineer / Site Reliability Engineer - **Company:** Ontrac Solutions - **Location:** New York, NY, United States (Remote available) - **Salary:** $187,200.0 - $208,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Unix, Cloud Computing, Cloud Engineering, Profiling, Computer Programming, Continuous Integration, Linux, Distributed Systems, Perl (Programming Language), Github, Python (Programming Language), OpenShift, Performance Tuning, Reliability Engineering, Ansible, Prometheus, Shell Script, Datadog, Data Logging, Scripting, Google Cloud, Cloud Platform System, Cloud Monitoring, System Availability, Delivery Pipeline, Grafana, Multi-Cloud, Kubernetes, Deployment Automation, Terraform, Golang - **Published:** July 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b3d98d00e0587a8e ## About the Role * Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources. * Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable. * OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management. * Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell). * Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift. ## Description We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools., * GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler). * Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations. * Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling. * CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines. * System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments. * SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability. ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)