Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
We are seeking a Senior Site Reliability Engineer (SRE) to join a high-impact Platform Engineering team focused on building scalable cloud infrastructure, reusable platform capabilities, and automation frameworks that enable software engineers to develop and deploy applications at scale. This is not a traditional operations role-the team focuses on platform engineering, Infrastructure as Code, developer enablement, and reliability through software and automation., You will design and build reusable cloud infrastructure across AWS and Azure, develop Infrastructure as Code from the ground up, create automation and self-service platform capabilities, enhance observability using Datadog, and contribute to modern CI/CD practices. The ideal candidate has a strong software engineering mindset and enjoys building platforms that improve engineering productivity while supporting highly available production environments.
Requirements
- 6+ years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, or Cloud Engineering.
- Hands-on experience designing, building, and supporting production cloud environments across both AWS and Azure.
- Strong experience building reusable Infrastructure as Code using Terraform, including custom modules, shared frameworks, and cloud platform components (AWS CDK and/or CloudFormation preferred).
- Experience designing and supporting Kubernetes-based platforms in production environments.
- Strong software engineering mindset with experience building automation using languages such as Go, Python, PowerShell, TypeScript, or similar.
- Strong experience with Datadog (or similar observability platforms), including monitoring, troubleshooting production issues, and improving platform reliability.
- Experience building modern CI/CD pipelines and developer enablement capabilities through platform engineering.
Nice to Have Skills & Experience
- Experience leveraging AI tools, AI agents, or LLM-powered workflows to improve engineering productivity or platform automation.
- Experience working in regulated industries such as healthcare or medical devices.
- Experience with serverless technologies (AWS Lambda, Azure Functions).
- AWS, Azure, Terraform, or Kubernetes certifications.
Benefits & conditions
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on jobs.insightglobal.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
Why Upskilling And Reskilling is Important For Developers
Find a Developer Job: 12 Best Job Sites For Developers
Fully Remote Software Engineer Jobs