Site Reliability Engineer

GREETINGS FROM, LLC
Atlanta, GA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Applications Architecture Audit Trail Bash Shell BigQuery Cloud Computing Cloud Engineering Cloud Storage Data Warehousing DevOps Disaster Recovery Distributed Systems
+19 more
Domain Name System (DNS) Fault Tolerance Identity and Access Management Python (Programming Language) Performance Tuning Windows PowerShell Reliability Engineering Prometheus Software Engineering Data Logging Google Cloud Load Balancing Delivery Pipeline Grafana Reliability of Systems Amazon Virtual Private Cloud (VPC) Containerization Kubernetes Terraform

Job description

We are seeking a highly skilled and proactive Senior Specialist, Site Reliability Engineering (SRE) to help drive reliability, scalability, and performance of our critical platforms while bringing deep technical expertise in Google Cloud Platform. This role is ideal for a senior-level engineer who combines deep technical expertise with a passion for automation, observability, and operational excellence and who is highly technical, thrives in distributed systems, and is passionate about operational excellence and modern cloud practices., As a Senior Specialist, you ll work on complex reliability challenges, lead technical initiatives, and collaborate across engineering, product, and infrastructure teams to ensure our systems are resilient and efficient., Architect and implement solutions that improve system reliability, scalability, and performance across Google Cloud Platform based services.

Define and manage SLIs, SLOs, and error budgets for critical systems.

Automate operational tasks, reduce toil, and improve the reliability posture of our environments.

Influence system and application architecture to ensure reliability is designed from the beginning.

  • Incident Management and Root Cause Analysis

Serve as the technical lead during major incidents and drive restoration efforts.

Conduct detailed root cause analysis and deliver long term corrective actions.

Champion and facilitate blameless postmortems and continuous improvement practices.

  • Cloud Architecture and Operations (Google Cloud Platform Focused)

Design Architect and improve Google Cloud Platform infrastructure including VPC design, Cloud DNS, load balancing, Cloud Armor equivalents for WAF and filtering, cloud storage patterns, managed compute platforms such as GKE and Cloud Run, and data warehouse platforms such as BigQuery.

Collaborate with teams to implement resilient multi zone and multi-region cloud architectures.

Lead the design and implementation of disaster recovery strategies and automated failover patterns within Google Cloud Platform.

Manage and optimize core Google Cloud Platform services such as IAM, service accounts, logging, and network controls.

Apply governance guardrails for secure multi project environments using tools such as Google Cloud Platform Organization policies, Cloud Identity, and related controls.

  • Automation and Infrastructure as Code

Build infrastructure using Terraform and maintain consistent, scalable IaC patterns.

Create automation using Python, Bash, PowerShell, or similar languages.

Participate in CI and CD pipeline improvements and ensure high quality deployments into Google Cloud Platform environments.

  • Monitoring & Tooling

Enhance observability through metrics, logs, and tracing using tools such as Prometheus, Grafana, Google Cloud Operations Suite, or similar solutions.

Build dashboards, alerts, and automated remediation systems that support reliability and performance goals.

Analyze cloud level logs such as VPC Flow Logs and Cloud Audit Logs to strengthen security and performance.

  • Technical Leadership

Collaborate with security and software engineering teams to drive reliability and cloud excellence.

Influence system design and architecture to embed reliability from the ground up.

Stay current with Google Cloud Platform capabilities and recommend improvements to enhance performance, security, and efficiency.

Requirements

  • 7 or more years of experience in SRE, DevOps, cloud engineering, or infrastructure engineering.
  • Strong experience with Google Cloud Platform architecture, networking, identity, and managed services.
  • Expertise with Kubernetes and container platforms.
  • Hands on experience implementing infrastructure as code using Terraform.
  • Strong proficiency with modern observability stacks.
  • Experience in Python, PowerShell, or similar languages.
  • Experience with orchestration platforms such as Harness.
  • Proven ability to diagnose and solve complex reliability problems in distributed systems.
  • Experience leveraging AI tools to enhance workflow automation, experimentation, and problem solving.
  • Excellent communication skills and the ability to influence outcomes across teams.

Preferred Skills

  • Experience in regulated or high-availability environments (e.g., financial services, healthcare).
  • Familiarity with chaos engineering, performance optimization, and capacity planning.
  • Software development background using languages such as Python or Go.
  • Experience designing multi-region fault tolerant architectures in Google Cloud Platform.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:27 min

Explaining query execution overhead and caching limitations in BigQuery

Adnan Rahic · JS Congress

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all