Google Site Reliability Engineer (SRE)

Spectraforce
United States
28 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Agile Modeling Airflow BigQuery Cloud Storage Github Intrusion Detection and Prevention Python (Programming Language) Machine Learning Microsoft Visual Studio Reliability Engineering Prometheus Cloudera
+9 more
Software Engineering Tableau (Software) Data Logging Grafana Ab Initio Pyspark Low Latency Splunk Servicenow

Job description

Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services.

Role focused on:

  • Responsible for Incident Detection & Logging and meeting the agreed SLA for incident tickets.
  • Responsible for Bridge Activation & Communication (P1-P2).
  • Postmortem Preparation (Within 24-72 Hours) & Root Cause Analysis.
  • Responsible for critical monitoring activities, Problem Management & Grafana Integration.
  • Participate in on-call rotations, handle incidents, and drive timely mitigation and recovery.
  • Automating operational work so services can scale without manual toil, also operating highly available, low latency & secure systems.
  • Defining and measuring reliability through SLIs/SLOs and error budgets.
  • Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services.
  • Tune alerting to reduce noise while ensuring rapid detection of user impacting issues.
  • Lead or contribute to post incident reviews and root cause analysis, and ensure follow-up actions are implemented to prevent recurrence.
  • Added Advantage if resource is familiar with Tools Tidal, ServiceNow, Xmatters, Abinitio, Tableau, Opsgenie&Zeke.

Requirements

Knowledge/experience in GCP (BigQuery, Cloud Storage, Dataproc, GKE, Airflow/Composer, Pub-sub, Cloud Functions, Cloud SQL, etc.) Knowledge/experience in GitHub & Visual Studio Code. Knowledge/experience in MS Copilot. Knowledge/experience in Prometheus, Grafana & Splunk. Knowledge in Python/Pyspark/Machine learning is an added advantage

Soft skills: Clear written and verbal communication, particularly under pressure (e.g., during incidents). Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority. Strong communication, analytical skills, knowledge of the entire Incident management life cycle process, Agile model experience, and problem-solving skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on leoforce.us

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all