> Markdown version of [/jobs/ext/1652983-google-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1652983-google-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Google Site Reliability Engineer (SRE) - **Company:** Spectraforce - **Location:** United States (Remote available) - **Contract:** Temporary to permanent - **Skills:** Agile Modeling, Airflow, BigQuery, Cloud Storage, Github, Intrusion Detection and Prevention, Python (Programming Language), Machine Learning, Microsoft Visual Studio, Reliability Engineering, Prometheus, Cloudera, Software Engineering, Tableau (Software), Data Logging, Grafana, Ab Initio, Pyspark, Low Latency, Splunk, Servicenow - **Published:** July 14, 2026 - **Apply:** http://leoforce.us/Careers/Spectraforce/JobDetails.html?jobid=f9d18a26-ab1e-4954-8d27-724f7ab44d56&OrgId=1&UserId=1298 ## About the Role Knowledge/experience in GCP (BigQuery, Cloud Storage, Dataproc, GKE, Airflow/Composer, Pub-sub, Cloud Functions, Cloud SQL, etc.) Knowledge/experience in GitHub & Visual Studio Code. Knowledge/experience in MS Copilot. Knowledge/experience in Prometheus, Grafana & Splunk. Knowledge in Python/Pyspark/Machine learning is an added advantage Soft skills: Clear written and verbal communication, particularly under pressure (e.g., during incidents). Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority. Strong communication, analytical skills, knowledge of the entire Incident management life cycle process, Agile model experience, and problem-solving skills ## Description Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services. Role focused on: * Responsible for Incident Detection & Logging and meeting the agreed SLA for incident tickets. * Responsible for Bridge Activation & Communication (P1-P2). * Postmortem Preparation (Within 24-72 Hours) & Root Cause Analysis. * Responsible for critical monitoring activities, Problem Management & Grafana Integration. * Participate in on-call rotations, handle incidents, and drive timely mitigation and recovery. * Automating operational work so services can scale without manual toil, also operating highly available, low latency & secure systems. * Defining and measuring reliability through SLIs/SLOs and error budgets. * Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services. * Tune alerting to reduce noise while ensuring rapid detection of user impacting issues. * Lead or contribute to post incident reviews and root cause analysis, and ensure follow-up actions are implemented to prevent recurrence. * Added Advantage if resource is familiar with Tools Tidal, ServiceNow, Xmatters, Abinitio, Tableau, Opsgenie&Zeke. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)