GCP Site Reliability Engineer

TekCommands Inc
Buffalo Grove, IL, United States
29 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$110,000.0 - $130,000.0
Working hours
Regular working hours

Tech stack

Agile Methodology Airflow BigQuery Cloud Computing Cloud Computing Security Cloud Storage Github Python (Programming Language) Machine Learning Microsoft Visual Studio Reliability Engineering Prometheus
+13 more
Cloudera Tableau (Software) Data Logging Google Cloud Microsoft Power Automate Grafana Reliability of Systems Ab Initio Pyspark Kubernetes Splunk Pagerduty Servicenow

Job description

We are seeking a Google Site Reliability Engineer (SRE) to build, operate, and support highly available, scalable, and secure cloud services on Google Cloud Platform (GCP). The ideal candidate will have strong experience in incident management, observability, automation, and cloud operations with expertise in GCP technologies. Key Responsibilities Monitor production systems and manage incident detection, logging, and resolution while meeting SLA targets. Lead bridge calls and communications for P1/P2 incidents. Perform root cause analysis (RCA) and prepare postmortem reports. Build and maintain monitoring dashboards, alerts, and observability using Prometheus, Grafana, and Splunk. Automate operational tasks to improve reliability and reduce manual effort. Define and manage SLIs, SLOs, and error budgets. Participate in on-call rotations and ensure timely incident mitigation and recovery. Collaborate with development and operations teams to improve system reliability and performance. Support problem management and continuous service improvements. Required Skills Strong experience with Google Cloud Platform (GCP) services including: BigQuery Cloud Storage Dataproc GKE (Google Kubernetes Engine) Airflow/Cloud Composer Pub/Sub Cloud Functions, The Ammonia Refrigeration Plant Engineer/HVAC Stationary Engineer Technician performs scheduled maintenance, safety inspections and repairs to varying types of equipment and report…

  • 21 days ago, The Site Reliability and Performance Engineer designs implement and maintains monitoring and observability solutions and supports the performance and scalability of IT infrastructu…
  • 1 month ago

Requirements

Experience with Prometheus, Grafana, and Splunk. Proficiency with GitHub and Visual Studio Code. Familiarity with Microsoft Copilot. Strong understanding of Incident Management, Problem Management, and Agile methodologies. Excellent communication, analytical, and troubleshooting skills. Nice to Have Python, PySpark, or Machine Learning experience. Experience with Tidal, ServiceNow, xMatters, Ab Initio, Tableau, Opsgenie, and Zeke. Top 3 Skills Google Cloud Platform (GCP) Site Reliability Engineering (SRE) & Incident Management Prometheus, Grafana & Splunk Monitoring Work Location: Remote or Hybrid (Buffalo Grove, IL)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

Videos

See all

Related articles

See all