AWS Site Reliability Engineer (SRE)
Savvyan Technologies
Columbus, OH, United States
6 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Backup Devices
Bash Shell
Cloud Engineering
Continuous Integration
Linux
DevOps
Disaster Recovery
Distributed Systems
Domain Name System (DNS)
+37 more
Fault Tolerance
Github
Identity and Access Management
Subnetting
Python (Programming Language)
Key Management
Reliability Engineering
Ansible
Prometheus
Software Deployment
Datadog
Data Logging
Load Balancing
Cloud Platform System
Autoscaling
Istio
System Availability
Grafana
Reliability of Systems
Infrastructure as Code (IaC)
Amazon Virtual Private Cloud (VPC)
Cloudformation
Gitlab-ci
Git Flow
Kubernetes
Infrastructure Automation Frameworks
Deployment Automation
Route53
Functional Programming
Cloudwatch
Terraform
Splunk
New Relic (SaaS)
Dynatrace
Docker
Jenkins
Microservices
Job description
The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, CI/CD, monitoring/observability, Linux, Python/Bash scripting, and production incident management. The engineer will work closely with development, infrastructure, security, and operations teams to improve system reliability and automate operational processes., * Design, implement, and maintain highly available and scalable infrastructure on AWS.
- Manage AWS services including EC2, EKS, ECS, S3, RDS, Lambda, VPC, IAM, CloudWatch, Route 53, and Load Balancers.
- Build and maintain infrastructure using Terraform/Terragrunt and Infrastructure as Code (IaC) best practices.
- Design, administer, and troubleshoot Kubernetes/EKS environments.
- Develop and maintain CI/CD pipelines using Jenkins, GitLab CI/CD, GitHub Actions, or similar tools.
- Automate infrastructure provisioning, application deployments, monitoring, and operational processes.
- Develop scripts and automation using Python, Bash, or Go.
- Implement monitoring, alerting, logging, and observability using tools such as CloudWatch, Splunk, Datadog, Dynatrace, Prometheus, and Grafana.
- Participate in 24x7 on-call rotation and respond to production incidents.
- Troubleshoot complex infrastructure, networking, application, and cloud-related issues.
- Lead incident response, root-cause analysis (RCA), and post-incident remediation activities.
- Define and improve SLOs, SLIs, SLAs, availability, latency, and reliability metrics.
- Implement auto-scaling, fault tolerance, disaster recovery, backup, and high-availability solutions.
- Identify opportunities to eliminate manual operational tasks through automation.
- Work with development teams to ensure applications are designed and deployed with reliability, scalability, and performance in mind.
- Implement AWS security best practices involving IAM, encryption, networking, secrets management, and least-privilege access.
- Monitor AWS infrastructure for performance and cost optimization opportunities.
- Create and maintain technical documentation, runbooks, operational procedures, and troubleshooting guides.
Requirements
- 6+ years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or related roles.
- Strong hands-on experience with AWS.
- Strong experience with Terraform/Terragrunt.
- Hands-on experience with Kubernetes/EKS.
- Strong understanding of Linux/Unix administration.
- Experience with CI/CD pipelines and deployment automation.
- Proficiency in Python, Bash, or Go.
- Strong understanding of AWS networking, VPC, subnets, security groups, IAM, DNS, and load balancing.
- Experience with monitoring, logging, alerting, and observability.
- Strong production troubleshooting and incident-management experience.
- Experience with distributed systems, microservices, and cloud-native architectures.
- Strong understanding of high availability, scalability, resiliency, and disaster recovery.
Preferred Skills
- AWS Certified DevOps Engineer or AWS Certified Solutions Architect.
- Experience with Helm, ArgoCD, GitOps, Docker, Ansible, or CloudFormation.
- Experience with Prometheus and Grafana.
- Experience with Splunk, Datadog, Dynatrace, or New Relic.
- Experience implementing blue/green, canary, and rolling deployments.
- Experience with service mesh technologies such as Istio.
- Experience with AWS cost optimization and FinOps.
- Experience working in financial services, healthcare, telecommunications, or other highly regulated environments.
- Strong knowledge of security and compliance requirements for cloud environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
LM
Luis Minvielle
Where To Find Software Engineering Jobs
over 2 years ago
EM
Eli McGarvie
React Developer Salary [2023]
over 3 years ago