Cloud Infrastructure Site Reliability Engineer

THE JUDGE GROUP, INC.
Berkeley Heights, NJ, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$145,600.0 - $166,400.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Automation of Tests Microsoft Azure Business Process Modeling C++ (Programming Language) Cloud Computing Cloud Engineering Continuous Delivery Continuous Integration Linux DevOps
+21 more
File Systems Distributed Systems Python (Programming Language) Networking Basics Reliability Engineering Cloud Services Ansible Software Engineering Data Logging Google Cloud Cloud Platform System Reliability of Systems Cloudformation Containerization Infrastructure Automation Frameworks Information Technology Data Management Terraform Dynatrace Serverless Computing Golang

Job description

As a Cloud Infrastructure Site Reliability Engineer (SRE) with expertise in multiple public cloud service provider platforms, you will be responsible for operating infrastructure solutions, following the principles and practices pioneered by Google’s SRE model. Your work will ensure our cloud services meet uptime, reliability, and performance targets, and you will drive automation and continuous improvement across our production environments. This role will involve collaborating with cross-functional teams to enhance our cloud reliability posture and streamline processes through automation., * Design, build, and maintain highly available, scalable, and secure cloud infrastructure on platforms such as AWS, Google Cloud Platform, or Azure.

  • Develop and implement automation for provisioning, monitoring, scaling, and incident response using Infrastructure-as-Code tools (e.g., Terraform, CloudFormation, Ansible).
  • Monitor system reliability, capacity, and performance; proactively detect and address issues before they impact users.
  • Respond to production incidents, participate in on-call rotations, and lead post-incident reviews to drive root cause analysis and reliability improvements.
  • Collaborate with software engineering and security teams to ensure new services and features are production-ready and meet reliability standards.
  • Build and maintain tools for deployment, monitoring, and operations; automate manual processes to reduce toil.
  • Document operational processes and system architectures to ensure knowledge sharing and repeatability.
  • Continuously evaluate and implement new technologies to improve system reliability, security, and efficiency.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of experience in software development with proficiency in at least one programming language (e.g., Python, Go, Java, C++).
  • Experience administrating cloud platforms (AWS, Google Cloud Platform, Azure), including networking, security, containerization, storage, data management, and serverless technologies.
  • Solid understanding of Linux systems, networking fundamentals, virtualized, and distributed systems, file systems, system processes and configurations.
  • Deep understanding of observability (monitoring, alerting, and logging) tools in cloud environments. Ability to set up and maintain monitoring dashboards, alerts, and logs.
  • Familiarity with Continuous Integration/Continuous Deployment (CI/CD) tools for automated testing, deployments, provisioning, and observability.
  • Ability to manage and respond to incidents, perform root cause analysis, and implement post-mortem reviews.
  • Understanding of setting, monitoring, and maintaining Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) for system reliability.

Needs experience with Terraform and Dynatrace

  • Additional Qualifications a Plus: Experience working with enterprise-scale financial services or other regulated industries
  • 5+ years of experience in SRE, DevOps, infrastructure, or cloud engineering roles, preferably supporting large-scale, distributed systems.
  • Excellent problem-solving, troubleshooting, and communication skills.
  • Experience leading technical projects or mentoring junior engineers.
  • Certifications: Certified Engineer, DevOps, SRE, CSREF

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:13 min

Defining cloud proficiency by technical role

Piet Van Dongen · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all