> Markdown version of [/jobs/ext/2039098-cloud-infrastructure-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2039098-cloud-infrastructure-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Infrastructure Site Reliability Engineer - **Company:** THE JUDGE GROUP, INC. - **Location:** Berkeley Heights, NJ, United States - **Experience:** Experienced - **Salary:** $145,600.0 - $166,400.0 - **Contract:** Temporary to permanent - **Skills:** Java (Programming Language), Amazon Web Services, Automation of Tests, Microsoft Azure, Business Process Modeling, C++ (Programming Language), Cloud Computing, Cloud Engineering, Continuous Delivery, Continuous Integration, Linux, DevOps, File Systems, Distributed Systems, Python (Programming Language), Networking Basics, Reliability Engineering, Cloud Services, Ansible, Software Engineering, Data Logging, Google Cloud, Cloud Platform System, Reliability of Systems, Cloudformation, Containerization, Infrastructure Automation Frameworks, Information Technology, Data Management, Terraform, Dynatrace, Serverless Computing, Golang - **Published:** August 12, 2026 - **Apply:** https://www.dice.com/job-detail/9552715e-294a-4b60-bd10-9746a2c9d7e1 ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. * 3+ years of experience in software development with proficiency in at least one programming language (e.g., Python, Go, Java, C++). * Experience administrating cloud platforms (AWS, Google Cloud Platform, Azure), including networking, security, containerization, storage, data management, and serverless technologies. * Solid understanding of Linux systems, networking fundamentals, virtualized, and distributed systems, file systems, system processes and configurations. * Deep understanding of observability (monitoring, alerting, and logging) tools in cloud environments. Ability to set up and maintain monitoring dashboards, alerts, and logs. * Familiarity with Continuous Integration/Continuous Deployment (CI/CD) tools for automated testing, deployments, provisioning, and observability. * Ability to manage and respond to incidents, perform root cause analysis, and implement post-mortem reviews. * Understanding of setting, monitoring, and maintaining Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) for system reliability. Needs experience with Terraform and Dynatrace * Additional Qualifications a Plus: Experience working with enterprise-scale financial services or other regulated industries * 5+ years of experience in SRE, DevOps, infrastructure, or cloud engineering roles, preferably supporting large-scale, distributed systems. * Excellent problem-solving, troubleshooting, and communication skills. * Experience leading technical projects or mentoring junior engineers. * Certifications: Certified Engineer, DevOps, SRE, CSREF ## Description As a Cloud Infrastructure Site Reliability Engineer (SRE) with expertise in multiple public cloud service provider platforms, you will be responsible for operating infrastructure solutions, following the principles and practices pioneered by Google's SRE model. Your work will ensure our cloud services meet uptime, reliability, and performance targets, and you will drive automation and continuous improvement across our production environments. This role will involve collaborating with cross-functional teams to enhance our cloud reliability posture and streamline processes through automation., * Design, build, and maintain highly available, scalable, and secure cloud infrastructure on platforms such as AWS, Google Cloud Platform, or Azure. * Develop and implement automation for provisioning, monitoring, scaling, and incident response using Infrastructure-as-Code tools (e.g., Terraform, CloudFormation, Ansible). * Monitor system reliability, capacity, and performance; proactively detect and address issues before they impact users. * Respond to production incidents, participate in on-call rotations, and lead post-incident reviews to drive root cause analysis and reliability improvements. * Collaborate with software engineering and security teams to ensure new services and features are production-ready and meet reliability standards. * Build and maintain tools for deployment, monitoring, and operations; automate manual processes to reduce toil. * Document operational processes and system architectures to ensure knowledge sharing and repeatability. * Continuously evaluate and implement new technologies to improve system reliability, security, and efficiency. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Hosting a modern justice system](https://www.wearedevelopers.com/videos/332-hosting-a-modern-justice-system) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)