Sr. Cloud/Site Reliability Engineer

RCG, Inc.
Woodbridge Township, NJ, United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$190,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Computer Engineering Python (Programming Language) Reliability Engineering Ansible Prometheus Systems Architecture Datadog Google Cloud
+14 more
Grafana Infrastructure as Code (IaC) Gitlab Cloudformation Containerization Gitlab-ci Kubernetes Deployment Automation Terraform Splunk Docker Elk Stack Jenkins Microservices

Job description

Design, implement, and maintain cloud infrastructure on AWS, Azure, or Google Cloud to support scalable, highly available applications. Manage and optimize Kubernetes clusters and Docker containers for seamless deployment and orchestration of microservices. Develop and maintain Infrastructure as Code (IaC) using tools such as Terraform, or CloudFormation to automate cloud resource provisioning and management. Implement, monitor, and tune performance of production systems using tools like Splunk, Prometheus, Grafana, Datadog, or the ELK stack to ensure uptime and system health. Utilize CI/CD pipelines using tools like Jenkins, GitLab CI, or Azure DevOps, ensuring smooth automated deployments with rollback strategies and monitoring. Write and maintain custom automation scripts in Python, Bash, or Go to streamline operational tasks and system maintenance. Collaborate with development teams to ensure alignment on system architecture, reliability goals, and incident management strategies. Drive the continuous improvement of incident response protocols, including the creation of SLOs, SLIs, and SLAs to meet operational goals. Troubleshoot and resolve incidents in production, ensuring root cause analysis and implementation of preventive measures. Stay up-to-date with industry trends, cloud technologies, and best practices in Site Reliability Engineering.

Requirements

Bachelor’s degree in Computer Engineering, Electrical Engineering or a related field, Position requires a Bachelor’s degree in Computer Engineering, Electrical Engineering or a related field plus five (5) years of experience. Position also requires any experience in the following: Cloud Platforms - at least one of the following: AWS, GCP, Azure; Infrastructure as Code (IaC): At least one of the following: Terraform, Ansible or CloudFormation; Monitoring and Incident Response; Scripting and Automation: At least one of the following: Python, Bash or Go; CI/CD Pipelines: At least one of the following: Jenkins, GitLab, Azure DevOps; and Containerization and Orchestration: such as Kubernetes or Docker.

Benefits & conditions

Salary range: from $190.000 to $220.000/year; benefits include PTO, medical/dental/vision insurance, 401(k)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all