Sr. Site Reliability Engineer - KK000095/ 95KK

Sumeru INC
Charlotte, NC, United States
10 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$96,800.0 - $145,200.0
Working hours
Regular working hours

Tech stack

Microsoft Azure Bash Shell Cloud Computing Continuous Integration DevOps Identity and Access Management Python (Programming Language) Windows PowerShell Reliability Engineering Prometheus Data Logging Scripting
+15 more
Cloud Platform System Cloud Monitoring Grafana Software Troubleshooting Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Deployment Automation Bicep Terraform Docker Jenkins

Job description

This role serves as a vendor-provided Site Reliability Engineer responsible for improving and protecting the reliability, scalability, and performance of a platform currently hosted on Microsoft Azure that will be migrated to Tencent Kubernetes Engine (TKE) in the future. It manages availability, latency, performance, security, and capacity while enabling efficient, automated software delivery across both the current Azure environment and the upcoming TKE platform. The role differentiates by combining deep Azure cloud infrastructure expertise with Kubernetes-based container orchestration, positioning the team for a smooth cloud-to-TKE migration. Success is measured by improved system uptime, faster incident resolution, and a reliable, well-supported migration path. The work directly impacts IT service quality, operational resilience, and customer experience through the transition and beyond. Responsibility Approx. % of Time Monitor, troubleshoot, and resolve incidents affecting availability, latency, and performance of current Azure-hosted workloads 20% Provision, configure, and manage Azure infrastructure (VMs, networking, storage, IAM) to support production and non-production environments 20% Support planning and execution of the platform’s migration from Azure to Tencent Kubernetes Engine (TKE), including workload containerization and cutover activities 20% Design, build, and maintain CI/CD pipelines that support automated deployment and testing today on Azure and going forward on TKE 15% Build and maintain observability tooling - dashboards, alerts, logging, and health checks - to proactively identify and address system risks across both environments 10% Drive automation and infrastructure-as-code practices to reduce manual toil and improve deployment consistency 10% Collaborate with internal engineering teams and stakeholders to support incident response, capacity planning, and migration readiness 5%

Requirements

Education: Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent practical experience. Experience: 4+ years in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, including hands-on production experience with Microsoft Azure (compute, networking, storage, IAM). Experience with Tencent Kubernetes Engine (TKE) or comparable Kubernetes platforms strongly preferred, as the platform will migrate to TKE. Technical Skills: Proficiency with CI/CD tooling (e.g., Azure DevOps, Jenkins, GitLab CI), containerization and orchestration (Kubernetes, Docker, TKE), infrastructure-as-code (Terraform, ARM/Bicep), scripting/automation (Python, Bash, PowerShell), and monitoring/observability platforms (e.g., Grafana, Prometheus, Azure Monitor). Other: Strong troubleshooting and incident-response skills; experience supporting cloud platform migrations a plus; ability to work effectively as an embedded vendor resource within a client engineering team; on-call availability as required. Azure certifications (e.g., AZ-104, AZ-400, AZ-500) and/or Kubernetes certifications (CKA/CKAD). Prior experience migrating workloads from a cloud VM-based platform to Kubernetes/TKE, including containerization of legacy services.

Benefits & conditions

  • Plano, TX
  • $96,800-145,200 per year NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organi…

  • 1 day ago +

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all