> Markdown version of [/jobs/ext/1154746-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1154746-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Specialty Cores Inc - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Microsoft Azure, Cloud Computing Security, Cloud Engineering, DevOps, Fault Tolerance, Reliability Engineering, Software Engineering, Datadog, Data Logging, Cloud Monitoring, System Availability, Delivery Pipeline, Mttr, Reliability of Systems, AWS Lambda, Infrastructure as Code (IaC), Bicep, Cloudwatch, Terraform, Serverless Computing, Microservices - **Published:** July 2, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=2a6abb99b1e07425 ## About the Role * Strong expertise in cloud platforms (Azure, AWS) * Deep understanding of cloud-native architecture patterns (microservices, containers (Azure Container Apps/AKS/EKS), serverless (Azure Functions/AWS Lambda)) * Proficiency in Infrastructure as Code (Terraform, ARM/Bicep) * Experience with observability platforms (Datadog, Azure Monitor, Azure Application Insights) * Knowledge of CI/CD pipelines and GitOps practices * Expertise in system reliability concepts: + SLI / SLO / SLA management + Chaos engineering + High availability & fault isolationFamiliarity with security, compliance, and regulatory controls (SOC, ISO, cloud security frameworks) Experience: * 5+ years experience in Site Reliability Engineering, DevOps, or Cloud Engineering * Proven experience supporting mission-critical production systems at scale * Hands-on experience with incident management and on-call operations * Experience implementing automated monitoring, alerting, and remediation frameworks * Exposure to regulated environments (insurance, financial services) preferred * Demonstrated ability to work across cross-functional architecture, engineering, and operations teams, Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over work authorization sponsorship now or in the future for this position. ## Description The Site Reliability Engineer (SRE) is responsible for ensuring the availability, scalability, performance, and resiliency of enterprise cloud platforms across Azure, and AWS environments. This role combines software engineering, automation, and infrastructure expertise to operationalize reliability engineering practices, drive cloud-native resiliency patterns, and enable business-critical applications to meet defined SLAs, SLOs, and compliance requirements. The SRE partners with engineering, security, and operations teams to implement observability, incident response frameworks, and reliability automation, aligning with enterprise architecture standards and regulatory expectations. Key Accountabilities/Deliverables: * Design and implement highly available, fault-tolerant architectures using cloud-native services (microservices, containers, serverless) * Define and operationalize SLOs, SLIs, and error budgets for critical applications and platforms * Build and maintain Infrastructure as Code (IaC) (Terraform) to ensure repeatable and compliant deployments * Develop automated remediation and self-healing capabilities to reduce MTTR and improve system resilience * Establish enterprise-level monitoring, logging, and observability frameworks (Datadog, Azure Monitor, CloudWatch, OpenTelemetry, Azure Application Insights) * Drive cost optimization (FinOps) initiatives, including resource utilization tracking and rightsizing recommendations * Support DR/BCP strategy execution, including failover testing and regional isolation validation * Collaborate with application teams to embed reliability engineering practices into CI/CD pipelines ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)