GCP Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
We are seeking a Google Site Reliability Engineer (SRE) to build, operate, and support highly available, scalable, and secure cloud services on Google Cloud Platform (GCP). The ideal candidate will have strong experience in incident management, observability, automation, and cloud operations with expertise in GCP technologies. Key Responsibilities Monitor production systems and manage incident detection, logging, and resolution while meeting SLA targets. Lead bridge calls and communications for P1/P2 incidents. Perform root cause analysis (RCA) and prepare postmortem reports. Build and maintain monitoring dashboards, alerts, and observability using Prometheus, Grafana, and Splunk. Automate operational tasks to improve reliability and reduce manual effort. Define and manage SLIs, SLOs, and error budgets. Participate in on-call rotations and ensure timely incident mitigation and recovery. Collaborate with development and operations teams to improve system reliability and performance. Support problem management and continuous service improvements. Required Skills Strong experience with Google Cloud Platform (GCP) services including: BigQuery Cloud Storage Dataproc GKE (Google Kubernetes Engine) Airflow/Cloud Composer Pub/Sub Cloud Functions, The Ammonia Refrigeration Plant Engineer/HVAC Stationary Engineer Technician performs scheduled maintenance, safety inspections and repairs to varying types of equipment and report…
- 21 days ago, The Site Reliability and Performance Engineer designs implement and maintains monitoring and observability solutions and supports the performance and scalability of IT infrastructu…
- 1 month ago
Requirements
Experience with Prometheus, Grafana, and Splunk. Proficiency with GitHub and Visual Studio Code. Familiarity with Microsoft Copilot. Strong understanding of Incident Management, Problem Management, and Agile methodologies. Excellent communication, analytical, and troubleshooting skills. Nice to Have Python, PySpark, or Machine Learning experience. Experience with Tidal, ServiceNow, xMatters, Ab Initio, Tableau, Opsgenie, and Zeke. Top 3 Skills Google Cloud Platform (GCP) Site Reliability Engineering (SRE) & Incident Management Prometheus, Grafana & Splunk Monitoring Work Location: Remote or Hybrid (Buffalo Grove, IL)
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
The Best X (Twitter) Accounts for Developers
Fully Remote Software Engineer Jobs
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence