Site Reliability Engineer

Collabera
Westbrook, ME, United States
3 days ago
Apply on www.collabera.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$83,200.0 - $124,800.0
Working hours
Regular working hours

Tech stack

.NET Framework Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Continuous Integration Relational Databases Software Debugging DevOps Java Platform Enterprise Edition (J2EE) Python (Programming Language) Microsoft SQL Server
+18 more
MySQL Oracle (Applications) Reliability Engineering SQL Databases Datadog ReactJS Software Troubleshooting Cloudformation Containerization AngularJS Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Amazon Simple Queue Service (SQS) Terraform Docker Jenkins

Job description

We are seeking a Site Reliability Engineer (SRE) to join the team. This role owns the reliability, availability, and performance of production systems-monitoring, incident response, automation, and infrastructure-so our applications stay fast, stable, and available for customers around the clock. What will you do? o Own the reliability, availability, and performance of hosted diagnostic imaging and telemedicine applications running on AWS. o Design, build, and maintain monitoring, alerting, and observability using Datadog and the AWS Console to detect and resolve issues before they affect customers. o Manage and optimize core AWS infrastructure (EC2, RDS, ECS, SQS, S3, and related services), applying infrastructure-as-code practices for repeatable, auditable deployments. o Build and maintain CI/CD pipelines in Jenkins for automated, low-risk deployment of releases and hotfixes. o Lead incident response and root cause analysis for production issues; write and maintain runbooks and postmortems to prevent recurrence. o Triage and resolve production escalations, fixing defects and shipping mini-releases with minimal customer disruption. o Define and track SLIs/SLOs and error budgets for critical services, and use that data to help prioritize engineering work. o Partner with development and product teams to build reliability, security, and scalability into new features from design through deployment. o Participate in an on-call rotation, providing timely response to production incidents.

Requirements

o 3+ years of experience in site reliability engineering, DevOps, or production systems support, ideally for cloud-hosted, customer-facing applications. o Hands-on experience with AWS services (EC2, RDS, ECS, SQS, S3) and infrastructure-as-code tools (CloudFormation or Terraform). o Experience with CI/CD tooling (Jenkins or similar) and automated deployment pipelines. o Proficiency with monitoring and observability platforms (Datadog or similar) and building actionable alerting. o Scripting and automation skills (Python, Bash, or similar) to reduce manual toil. o Working knowledge of relational databases (Oracle, MySQL, SQL Server) and SQL. o Strong troubleshooting and incident-response skills, with the ability to stay calm and methodical under production pressure. o Excellent communication skills, both verbal and written, including the ability to translate technical issues to non-technical audiences. o Ability to work independently and within cross-functional teams in an agile environment. Preferred Experience: o BA/BS in Computer Science, Engineering, or a related field, or equivalent work experience. o Experience with containerization and orchestration (Docker, ECS, or Kubernetes). o Familiarity with Java/J2EE, .NET, or similar application stacks to support root-cause debugging. o Experience handling raw image data or image files, or working in healthcare or other regulated environments. o AWS certification (SysOps Administrator, DevOps Engineer, or Solutions Architect). o Knowledge of Angular or React front-end stacks, useful for full-stack incident triage.

Benefits & conditions

The Company offers the following benefits for this position, subject to applicable eligibility requirements: medical insurance, dental insurance, vision insurance, 401(k) retirement plan, life insurance, long-term disability insurance, short-term disability insurance, paid parking/public transportation, paid time off, paid sick and safe time, hours of paid vacation time, weeks of paid parental leave, and paid holidays annually - as applicable.

Job Requirement o Datadog o aws o SRE o site reliability engineer

Reach Out to a Recruiter o Recruiter

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collabera.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · World Congress 2026 Europe

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all