> Markdown version of [/jobs/ext/2874344-sr-site-reliability-engineer-kk000095-95kk](https://www.wearedevelopers.com/jobs/ext/2874344-sr-site-reliability-engineer-kk000095-95kk). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Site Reliability Engineer - KK000095/ 95KK - **Company:** Sumeru INC - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Salary:** $96,800.0 - $145,200.0 - **Contract:** Permanent contract - **Skills:** Microsoft Azure, Bash Shell, Cloud Computing, Continuous Integration, DevOps, Identity and Access Management, Python (Programming Language), Windows PowerShell, Reliability Engineering, Prometheus, Data Logging, Scripting, Cloud Platform System, Cloud Monitoring, Grafana, Software Troubleshooting, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Deployment Automation, Bicep, Terraform, Docker, Jenkins - **Published:** September 13, 2026 - **Apply:** https://www.careerjet.com/jobad/us536e4616d536db5d548ab8393920bd82 ## About the Role Education: Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience. Experience: 4+ years in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, including hands-on production experience with Microsoft Azure (compute, networking, storage, IAM). Experience with Tencent Kubernetes Engine (TKE) or comparable Kubernetes platforms strongly preferred, as the platform will migrate to TKE. Technical Skills: Proficiency with CI/CD tooling (e.g., Azure DevOps, Jenkins, GitLab CI), containerization and orchestration (Kubernetes, Docker, TKE), infrastructure-as-code (Terraform, ARM/Bicep), scripting/automation (Python, Bash, PowerShell), and monitoring/observability platforms (e.g., Grafana, Prometheus, Azure Monitor). Other: Strong troubleshooting and incident-response skills; experience supporting cloud platform migrations a plus; ability to work effectively as an embedded vendor resource within a client engineering team; on-call availability as required. Azure certifications (e.g., AZ-104, AZ-400, AZ-500) and/or Kubernetes certifications (CKA/CKAD). Prior experience migrating workloads from a cloud VM-based platform to Kubernetes/TKE, including containerization of legacy services. ## Description This role serves as a vendor-provided Site Reliability Engineer responsible for improving and protecting the reliability, scalability, and performance of a platform currently hosted on Microsoft Azure that will be migrated to Tencent Kubernetes Engine (TKE) in the future. It manages availability, latency, performance, security, and capacity while enabling efficient, automated software delivery across both the current Azure environment and the upcoming TKE platform. The role differentiates by combining deep Azure cloud infrastructure expertise with Kubernetes-based container orchestration, positioning the team for a smooth cloud-to-TKE migration. Success is measured by improved system uptime, faster incident resolution, and a reliable, well-supported migration path. The work directly impacts IT service quality, operational resilience, and customer experience through the transition and beyond. Responsibility Approx. % of Time Monitor, troubleshoot, and resolve incidents affecting availability, latency, and performance of current Azure-hosted workloads 20% Provision, configure, and manage Azure infrastructure (VMs, networking, storage, IAM) to support production and non-production environments 20% Support planning and execution of the platform's migration from Azure to Tencent Kubernetes Engine (TKE), including workload containerization and cutover activities 20% Design, build, and maintain CI/CD pipelines that support automated deployment and testing today on Azure and going forward on TKE 15% Build and maintain observability tooling - dashboards, alerts, logging, and health checks - to proactively identify and address system risks across both environments 10% Drive automation and infrastructure-as-code practices to reduce manual toil and improve deployment consistency 10% Collaborate with internal engineering teams and stakeholders to support incident response, capacity planning, and migration readiness 5% ## Related Videos - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)