Sr. Cloud/Site Reliability Engineer

RCG, Inc.
Woodbridge Township, United States of America
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 220K

Job location

Remote
Woodbridge Township, United States of America

Tech stack

Amazon Web Services (AWS)
Azure
Bash
Cloud Computing
Computer Engineering
Python
Reliability Engineering
Ansible
Prometheus
Systems Architecture
Datadog
Google Cloud Platform
Grafana
Infrastructure as Code (IaC)
Gitlab
Cloudformation
Containerization
Gitlab-ci
Kubernetes
Deployment Automation
Terraform
Splunk
Docker
ELK
Jenkins
Microservices

Job description

Design, implement, and maintain cloud infrastructure on AWS, Azure, or Google Cloud to support scalable, highly available applications. Manage and optimize Kubernetes clusters and Docker containers for seamless deployment and orchestration of microservices. Develop and maintain Infrastructure as Code (IaC) using tools such as Terraform, or CloudFormation to automate cloud resource provisioning and management. Implement, monitor, and tune performance of production systems using tools like Splunk, Prometheus, Grafana, Datadog, or the ELK stack to ensure uptime and system health. Utilize CI/CD pipelines using tools like Jenkins, GitLab CI, or Azure DevOps, ensuring smooth automated deployments with rollback strategies and monitoring. Write and maintain custom automation scripts in Python, Bash, or Go to streamline operational tasks and system maintenance. Collaborate with development teams to ensure alignment on system architecture, reliability goals, and incident management strategies. Drive the continuous improvement of incident response protocols, including the creation of SLOs, SLIs, and SLAs to meet operational goals. Troubleshoot and resolve incidents in production, ensuring root cause analysis and implementation of preventive measures. Stay up-to-date with industry trends, cloud technologies, and best practices in Site Reliability Engineering.

Requirements

Bachelor's degree in Computer Engineering, Electrical Engineering or a related field, Position requires a Bachelor's degree in Computer Engineering, Electrical Engineering or a related field plus five (5) years of experience. Position also requires any experience in the following: Cloud Platforms - at least one of the following: AWS, GCP, Azure; Infrastructure as Code (IaC): At least one of the following: Terraform, Ansible or CloudFormation; Monitoring and Incident Response; Scripting and Automation: At least one of the following: Python, Bash or Go; CI/CD Pipelines: At least one of the following: Jenkins, GitLab, Azure DevOps; and Containerization and Orchestration: such as Kubernetes or Docker.

Benefits & conditions

Salary range: from $190.000 to $220.000/year; benefits include PTO, medical/dental/vision insurance, 401(k)

Apply for this position