Site Reliability Engineer
Spectraforce
Woonsocket, RI, United States
7 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on leoforce.us
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
Artificial Intelligence
Airflow
Batch Processing
Computer Programming
Software Debugging
DevOps
Distributed Systems
Elasticsearch
Python (Programming Language)
PostgreSQL
Reliability Engineering
+19 more
Ansible
Prometheus
Software Engineering
SQL Databases
Data Streaming
Google Cloud
Cloud Platform System
ReactJS
Istio
Large Language Models
Grafana
HybridCloud
Kubernetes
Rancher
Google Bigquery
Apache Kafka
Terraform
Splunk
Data Pipelines
Job description
- Own and drive the end-to-end reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and on-premises environments.
- Establish and maintain SLI/SLO health, alerting strategies, observability standards, and business-aligned monitoring for the assigned application domain.
- Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
- Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
- Build and optimize automation, self-service capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
- Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
- Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement.
- Influence organizational adoption of SLO-driven engineering, observability, incident management, and reliability practices through collaboration, credibility, and measurable outcomes.
Requirements
- 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility.
- Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents - structured leadership updates, not just participant involvement.
- Experience tuning and validating time-series anomaly detection models in a production observability context anomaly-based detection is a core function of this role.
- Strong programming proficiency in Python, React, and Java at production quality - capable of writing operational tooling that other engineers will rely on.
- Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services.
- Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solutions (Loki, Splunk, Elasticsearch).
- Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments.
- Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
- Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes.
- Experience with AI-assisted tooling and development.
- Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipelines: Apache Airflow and Tidal., * Experience owning Production Readiness Reviews or service launch gates.
- Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics and reporting frameworks: Google BigQuery, PostgreSQL.
- Hands-on chaos or fault injection experience.
- TIC (Technical Incident Commander) certification or equivalent structured incident command training.
- Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact.
- LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) - design or implementation experience.
- Experience with streaming data platforms: Kafka.
- Experience with service mesh and traffic management: Istio, Envoy.
- Infrastructure-as-code proficiency at production scale: Terraform or Ansible. Position Summary
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on leoforce.us
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
about 2 months ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
CH
Chris Heilmann
The Geometry of Incidents: Connecting User Impact to Architecture
21 days ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago