Site Reliability Engineer

Specialty Cores Inc
United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Microsoft Azure Cloud Computing Security Cloud Engineering DevOps Fault Tolerance Reliability Engineering Software Engineering Datadog Data Logging Cloud Monitoring
+11 more
System Availability Delivery Pipeline Mttr Reliability of Systems AWS Lambda Infrastructure as Code (IaC) Bicep Cloudwatch Terraform Serverless Computing Microservices

Job description

The Site Reliability Engineer (SRE) is responsible for ensuring the availability, scalability, performance, and resiliency of enterprise cloud platforms across Azure, and AWS environments.

This role combines software engineering, automation, and infrastructure expertise to operationalize reliability engineering practices, drive cloud-native resiliency patterns, and enable business-critical applications to meet defined SLAs, SLOs, and compliance requirements.

The SRE partners with engineering, security, and operations teams to implement observability, incident response frameworks, and reliability automation, aligning with enterprise architecture standards and regulatory expectations.

Key Accountabilities/Deliverables:

  • Design and implement highly available, fault-tolerant architectures using cloud-native services (microservices, containers, serverless)
  • Define and operationalize SLOs, SLIs, and error budgets for critical applications and platforms
  • Build and maintain Infrastructure as Code (IaC) (Terraform) to ensure repeatable and compliant deployments
  • Develop automated remediation and self-healing capabilities to reduce MTTR and improve system resilience
  • Establish enterprise-level monitoring, logging, and observability frameworks (Datadog, Azure Monitor, CloudWatch, OpenTelemetry, Azure Application Insights)
  • Drive cost optimization (FinOps) initiatives, including resource utilization tracking and rightsizing recommendations
  • Support DR/BCP strategy execution, including failover testing and regional isolation validation
  • Collaborate with application teams to embed reliability engineering practices into CI/CD pipelines

Requirements

  • Strong expertise in cloud platforms (Azure, AWS)
  • Deep understanding of cloud-native architecture patterns (microservices, containers (Azure Container Apps/AKS/EKS), serverless (Azure Functions/AWS Lambda))
  • Proficiency in Infrastructure as Code (Terraform, ARM/Bicep)
  • Experience with observability platforms (Datadog, Azure Monitor, Azure Application Insights)
  • Knowledge of CI/CD pipelines and GitOps practices
  • Expertise in system reliability concepts:
  • SLI / SLO / SLA management
  • Chaos engineering
  • High availability & fault isolationFamiliarity with security, compliance, and regulatory controls (SOC, ISO, cloud security frameworks)

Experience:

  • 5+ years experience in Site Reliability Engineering, DevOps, or Cloud Engineering
  • Proven experience supporting mission-critical production systems at scale
  • Hands-on experience with incident management and on-call operations
  • Experience implementing automated monitoring, alerting, and remediation frameworks
  • Exposure to regulated environments (insurance, financial services) preferred
  • Demonstrated ability to work across cross-functional architecture, engineering, and operations teams, Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over work authorization sponsorship now or in the future for this position.

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Health savings account
  • Dental insurance
  • Flexible spending account
  • Employee assistance program, At Core Specialty, you will receive a competitive salary and opportunities for professional development and advancement. We offer medical, dental, vision, and life insurances; short and long-term disability; a Company-match of 100% of a 6% contribution 401(k) plan; an Employee Assistance Plan; Health Savings Account, Flexible Spending Account, Health Reimbursement Account, and a wellness program

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

Videos

See all

Related articles

See all