Site Reliability Engineer
Procal Technologies
Peru, IN, United States
4 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source
Tech stack
Agile Methodology
Amazon Web Services
Bash Shell
Cloud Computing
Cloud Engineering
DevOps
Distributed Systems
Python (Programming Language)
Performance Tuning
Reliability Engineering
Prometheus
Datadog
+12 more
Data Logging
Scripting
Google Cloud
Enterprise Software Applications
Cloud Monitoring
Istio
Grafana
Kubernetes
Cloudwatch
Terraform
Splunk
Devsecops
Job description
We are looking for a Senior Site Reliability Engineer to build, operate, and improve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a hands-on technical leadership role focused on AWS/Google Cloud Platform, Kubernetes, observability, automation, and reliability engineering., * Design and implement reliability strategies for distributed systems across AWS and Google Cloud Platform.
- Define and monitor SLIs, SLOs, error budgets, and reliability metrics.
- Build and enhance monitoring, logging, tracing, alerting, and observability solutions.
- Lead incident response, root cause analysis, and postmortems.
- Improve system performance, scalability, resiliency, and operational readiness.
- Automate operational processes and reduce manual toil.
- Guide engineering teams on reliability architecture, capacity planning, and non-functional requirements.
Requirements
- 7+ years in SRE, Cloud Engineering, DevOps, or Platform Engineering.
- Strong production experience with AWS and/or Google Cloud Platform.
- Hands-on expertise with Kubernetes (EKS/GKE).
- Strong understanding of SRE principles, SLIs, SLOs, error budgets, and incident management.
- Experience with observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, or Splunk.
- Strong Terraform/IaC experience.
- Proficiency in Python, Bash, or similar scripting languages.
- Strong understanding of cloud networking, distributed systems, security, and performance optimization.
Preferred
- Large-scale cloud migration or modernization experience.
- Chaos engineering/resilience testing experience.
- Istio/service mesh knowledge.
- AWS/Google Cloud Platform certifications.
- Experience in Agile, DevOps, or DevSecOps environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
25 days ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago