> Markdown version of [/jobs/ext/2694308-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2694308-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Blu Omega - **Location:** Atlanta, GA, United States - **Experience:** Experienced - **Salary:** $130,000.0 - $145,000.0 - **Contract:** Permanent contract - **Skills:** Active Directory, Artificial Intelligence, Business Analytics Applications, Systems Engineering, Microsoft Azure, Cloud Computing, DevOps, Reliability Engineering, Power BI, Azure Active Directory, Prometheus, Runbook, Systems Integration, GitHub Copilot, Grafana, Kubernetes, Azure AKS, Terraform, Splunk, Azure Resource Manager - **Published:** September 3, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9138521/site-reliability-engineer ## About the Role Ability to obtain and maintain a Public Trust or Suitability/Fitness determination, * 4+ years of experience in cloud infrastructure, systems engineering, DevOps, or SRE roles * 2+ years of hands-on Microsoft Azure experience supporting production environments * 2+ years of Terraform experience, including designing modules and managing production infrastructure * Experience administering, operating, or supporting Kubernetes environments, preferably AKS * Experience automating cloud infrastructure and troubleshooting cloud networking, Kubernetes, and reliability issues * Strong analytical and problem-solving skills with the ability to collaborate across teams * Ability to work on-site in Atlanta, GA * Ability to obtain and maintain a Public Trust or Suitability/Fitness determination * Bachelor's degree in a related field or equivalent professional experience Nice to Have * Experience with certificate lifecycle management and renewal processes * Experience developing operational dashboards using Power BI or similar tools * Familiarity with observability platforms such as Grafana, Prometheus, Elastic, or Splunk * Experience integrating Azure resources with Microsoft Entra ID / Active Directory * Use of AI-assisted coding tools (e.g., GitHub Copilot) to support automation and development * Microsoft Azure, Kubernetes, or Terraform certifications ## Description Blu Omega is seeking a Site Reliability Engineer to support a federal program focused on enterprise cloud modernization. This role operates within a hybrid environment and is responsible for ensuring the reliability, scalability, and operational health of mission-critical cloud-based data and analytics platforms. The position requires experience supporting mission-focused cloud infrastructure environments., * Design, implement, and troubleshoot production Azure infrastructure using Terraform * Develop and maintain reusable Terraform modules, manage Terraform state, and troubleshoot deployment failures * Support the reliability, performance, scalability, and availability of workloads in Azure Kubernetes Service (AKS) * Automate cloud infrastructure and operational processes to improve consistency and platform reliability * Partner with development, platform, security, and operations teams to improve deployment and incident response processes * Implement and enhance monitoring, alerting, and operational dashboards to track platform health and performance * Identify infrastructure and reliability risks, recommending architecture and process improvements * Support incident investigation, root-cause analysis, and remediation of production issues * Develop and maintain documentation, operational procedures, troubleshooting guides, and runbooks ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)