Site Reliability Engineer

Procal Technologies
Peru, IN, United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Bash Shell Cloud Computing Cloud Engineering DevOps Distributed Systems Python (Programming Language) Performance Tuning Reliability Engineering Prometheus Datadog
+12 more
Data Logging Scripting Google Cloud Enterprise Software Applications Cloud Monitoring Istio Grafana Kubernetes Cloudwatch Terraform Splunk Devsecops

Job description

We are looking for a Senior Site Reliability Engineer to build, operate, and improve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a hands-on technical leadership role focused on AWS/Google Cloud Platform, Kubernetes, observability, automation, and reliability engineering., * Design and implement reliability strategies for distributed systems across AWS and Google Cloud Platform.

  • Define and monitor SLIs, SLOs, error budgets, and reliability metrics.
  • Build and enhance monitoring, logging, tracing, alerting, and observability solutions.
  • Lead incident response, root cause analysis, and postmortems.
  • Improve system performance, scalability, resiliency, and operational readiness.
  • Automate operational processes and reduce manual toil.
  • Guide engineering teams on reliability architecture, capacity planning, and non-functional requirements.

Requirements

  • 7+ years in SRE, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong production experience with AWS and/or Google Cloud Platform.
  • Hands-on expertise with Kubernetes (EKS/GKE).
  • Strong understanding of SRE principles, SLIs, SLOs, error budgets, and incident management.
  • Experience with observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, or Splunk.
  • Strong Terraform/IaC experience.
  • Proficiency in Python, Bash, or similar scripting languages.
  • Strong understanding of cloud networking, distributed systems, security, and performance optimization.

Preferred

  • Large-scale cloud migration or modernization experience.
  • Chaos engineering/resilience testing experience.
  • Istio/service mesh knowledge.
  • AWS/Google Cloud Platform certifications.
  • Experience in Agile, DevOps, or DevSecOps environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all