Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
We are seeking a Site Reliability Engineer (SRE) to join the team. This role owns the reliability, availability, and performance of production systems-monitoring, incident response, automation, and infrastructure-so our applications stay fast, stable, and available for customers around the clock. What will you do? o Own the reliability, availability, and performance of hosted diagnostic imaging and telemedicine applications running on AWS. o Design, build, and maintain monitoring, alerting, and observability using Datadog and the AWS Console to detect and resolve issues before they affect customers. o Manage and optimize core AWS infrastructure (EC2, RDS, ECS, SQS, S3, and related services), applying infrastructure-as-code practices for repeatable, auditable deployments. o Build and maintain CI/CD pipelines in Jenkins for automated, low-risk deployment of releases and hotfixes. o Lead incident response and root cause analysis for production issues; write and maintain runbooks and postmortems to prevent recurrence. o Triage and resolve production escalations, fixing defects and shipping mini-releases with minimal customer disruption. o Define and track SLIs/SLOs and error budgets for critical services, and use that data to help prioritize engineering work. o Partner with development and product teams to build reliability, security, and scalability into new features from design through deployment. o Participate in an on-call rotation, providing timely response to production incidents.
Requirements
o 3+ years of experience in site reliability engineering, DevOps, or production systems support, ideally for cloud-hosted, customer-facing applications. o Hands-on experience with AWS services (EC2, RDS, ECS, SQS, S3) and infrastructure-as-code tools (CloudFormation or Terraform). o Experience with CI/CD tooling (Jenkins or similar) and automated deployment pipelines. o Proficiency with monitoring and observability platforms (Datadog or similar) and building actionable alerting. o Scripting and automation skills (Python, Bash, or similar) to reduce manual toil. o Working knowledge of relational databases (Oracle, MySQL, SQL Server) and SQL. o Strong troubleshooting and incident-response skills, with the ability to stay calm and methodical under production pressure. o Excellent communication skills, both verbal and written, including the ability to translate technical issues to non-technical audiences. o Ability to work independently and within cross-functional teams in an agile environment. Preferred Experience: o BA/BS in Computer Science, Engineering, or a related field, or equivalent work experience. o Experience with containerization and orchestration (Docker, ECS, or Kubernetes). o Familiarity with Java/J2EE, .NET, or similar application stacks to support root-cause debugging. o Experience handling raw image data or image files, or working in healthcare or other regulated environments. o AWS certification (SysOps Administrator, DevOps Engineer, or Solutions Architect). o Knowledge of Angular or React front-end stacks, useful for full-stack incident triage.
Benefits & conditions
The Company offers the following benefits for this position, subject to applicable eligibility requirements: medical insurance, dental insurance, vision insurance, 401(k) retirement plan, life insurance, long-term disability insurance, short-term disability insurance, paid parking/public transportation, paid time off, paid sick and safe time, hours of paid vacation time, weeks of paid parental leave, and paid holidays annually - as applicable.
Job Requirement o Datadog o aws o SRE o site reliability engineer
Reach Out to a Recruiter o Recruiter
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs