> Markdown version of [/jobs/ext/960780-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/960780-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** GREETINGS FROM, LLC - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Applications Architecture, Audit Trail, Bash Shell, BigQuery, Cloud Computing, Cloud Engineering, Cloud Storage, Data Warehousing, DevOps, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Identity and Access Management, Python (Programming Language), Performance Tuning, Windows PowerShell, Reliability Engineering, Prometheus, Software Engineering, Data Logging, Google Cloud, Load Balancing, Delivery Pipeline, Grafana, Reliability of Systems, Amazon Virtual Private Cloud (VPC), Containerization, Kubernetes, Terraform - **Published:** June 12, 2026 - **Apply:** https://www.dice.com/job-detail/af7ebcda-ae2d-4940-ba43-9ee617d10810 ## About the Role * 7 or more years of experience in SRE, DevOps, cloud engineering, or infrastructure engineering. * Strong experience with Google Cloud Platform architecture, networking, identity, and managed services. * Expertise with Kubernetes and container platforms. * Hands on experience implementing infrastructure as code using Terraform. * Strong proficiency with modern observability stacks. * Experience in Python, PowerShell, or similar languages. * Experience with orchestration platforms such as Harness. * Proven ability to diagnose and solve complex reliability problems in distributed systems. * Experience leveraging AI tools to enhance workflow automation, experimentation, and problem solving. * Excellent communication skills and the ability to influence outcomes across teams. Preferred Skills * Experience in regulated or high-availability environments (e.g., financial services, healthcare). * Familiarity with chaos engineering, performance optimization, and capacity planning. * Software development background using languages such as Python or Go. * Experience designing multi-region fault tolerant architectures in Google Cloud Platform. ## Description We are seeking a highly skilled and proactive Senior Specialist, Site Reliability Engineering (SRE) to help drive reliability, scalability, and performance of our critical platforms while bringing deep technical expertise in Google Cloud Platform. This role is ideal for a senior-level engineer who combines deep technical expertise with a passion for automation, observability, and operational excellence and who is highly technical, thrives in distributed systems, and is passionate about operational excellence and modern cloud practices., As a Senior Specialist, you ll work on complex reliability challenges, lead technical initiatives, and collaborate across engineering, product, and infrastructure teams to ensure our systems are resilient and efficient., Architect and implement solutions that improve system reliability, scalability, and performance across Google Cloud Platform based services. Define and manage SLIs, SLOs, and error budgets for critical systems. Automate operational tasks, reduce toil, and improve the reliability posture of our environments. Influence system and application architecture to ensure reliability is designed from the beginning. * Incident Management and Root Cause Analysis Serve as the technical lead during major incidents and drive restoration efforts. Conduct detailed root cause analysis and deliver long term corrective actions. Champion and facilitate blameless postmortems and continuous improvement practices. * Cloud Architecture and Operations (Google Cloud Platform Focused) Design Architect and improve Google Cloud Platform infrastructure including VPC design, Cloud DNS, load balancing, Cloud Armor equivalents for WAF and filtering, cloud storage patterns, managed compute platforms such as GKE and Cloud Run, and data warehouse platforms such as BigQuery. Collaborate with teams to implement resilient multi zone and multi-region cloud architectures. Lead the design and implementation of disaster recovery strategies and automated failover patterns within Google Cloud Platform. Manage and optimize core Google Cloud Platform services such as IAM, service accounts, logging, and network controls. Apply governance guardrails for secure multi project environments using tools such as Google Cloud Platform Organization policies, Cloud Identity, and related controls. * Automation and Infrastructure as Code Build infrastructure using Terraform and maintain consistent, scalable IaC patterns. Create automation using Python, Bash, PowerShell, or similar languages. Participate in CI and CD pipeline improvements and ensure high quality deployments into Google Cloud Platform environments. * Monitoring & Tooling Enhance observability through metrics, logs, and tracing using tools such as Prometheus, Grafana, Google Cloud Operations Suite, or similar solutions. Build dashboards, alerts, and automated remediation systems that support reliability and performance goals. Analyze cloud level logs such as VPC Flow Logs and Cloud Audit Logs to strengthen security and performance. * Technical Leadership Collaborate with security and software engineering teams to drive reliability and cloud excellence. Influence system design and architecture to embed reliability from the ground up. Stay current with Google Cloud Platform capabilities and recommend improvements to enhance performance, security, and efficiency. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)