> Markdown version of [/jobs/ext/1256565-gcp-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1256565-gcp-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GCP Site Reliability Engineer - **Company:** TekCommands Inc - **Location:** Buffalo Grove, IL, United States (Remote available) - **Salary:** $110,000.0 - $130,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Airflow, BigQuery, Cloud Computing, Cloud Computing Security, Cloud Storage, Github, Python (Programming Language), Machine Learning, Microsoft Visual Studio, Reliability Engineering, Prometheus, Cloudera, Tableau (Software), Data Logging, Google Cloud, Microsoft Power Automate, Grafana, Reliability of Systems, Ab Initio, Pyspark, Kubernetes, Splunk, Pagerduty, Servicenow - **Published:** July 13, 2026 - **Apply:** https://www.careerjet.com/jobad/us497047b63dd4c263e147590ada24cdb1 ## About the Role Experience with Prometheus, Grafana, and Splunk. Proficiency with GitHub and Visual Studio Code. Familiarity with Microsoft Copilot. Strong understanding of Incident Management, Problem Management, and Agile methodologies. Excellent communication, analytical, and troubleshooting skills. Nice to Have Python, PySpark, or Machine Learning experience. Experience with Tidal, ServiceNow, xMatters, Ab Initio, Tableau, Opsgenie, and Zeke. Top 3 Skills Google Cloud Platform (GCP) Site Reliability Engineering (SRE) & Incident Management Prometheus, Grafana & Splunk Monitoring Work Location: Remote or Hybrid (Buffalo Grove, IL) ## Description We are seeking a Google Site Reliability Engineer (SRE) to build, operate, and support highly available, scalable, and secure cloud services on Google Cloud Platform (GCP). The ideal candidate will have strong experience in incident management, observability, automation, and cloud operations with expertise in GCP technologies. Key Responsibilities Monitor production systems and manage incident detection, logging, and resolution while meeting SLA targets. Lead bridge calls and communications for P1/P2 incidents. Perform root cause analysis (RCA) and prepare postmortem reports. Build and maintain monitoring dashboards, alerts, and observability using Prometheus, Grafana, and Splunk. Automate operational tasks to improve reliability and reduce manual effort. Define and manage SLIs, SLOs, and error budgets. Participate in on-call rotations and ensure timely incident mitigation and recovery. Collaborate with development and operations teams to improve system reliability and performance. Support problem management and continuous service improvements. Required Skills Strong experience with Google Cloud Platform (GCP) services including: BigQuery Cloud Storage Dataproc GKE (Google Kubernetes Engine) Airflow/Cloud Composer Pub/Sub Cloud Functions, The Ammonia Refrigeration Plant Engineer/HVAC Stationary Engineer Technician performs scheduled maintenance, safety inspections and repairs to varying types of equipment and report… + 21 days ago, The Site Reliability and Performance Engineer designs implement and maintains monitoring and observability solutions and supports the performance and scalability of IT infrastructu… + 1 month ago ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)