Lead Site Reliability Engineer - Cloud Platform...

HTC Global Services, Inc.
Burbank, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Engineering DevOps Disaster Recovery Distributed Systems Python (Programming Language) Reliability Engineering Prometheus Google Cloud
+8 more
Cloud Platform System Grafana Kubernetes Helm Charts Multi-Cloud Kubernetes Cloud Migration Terraform Splunk

Job description

We are seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence across a rapidly growing technology ecosystem. This role serves as a technical leader focused on cloud architecture, Kubernetes platforms, infrastructure automation, and highly available distributed systems. The position plays a key role in defining infrastructure strategy, improving platform resiliency, and mentoring engineering teams., * Design and support highly available cloud infrastructure in GCP * Architect and manage Kubernetes environments at scale * Build and maintain Infrastructure-as-Code using Terraform * Develop and manage Helm charts and Kubernetes deployments * Design failover, disaster recovery, and multi-region strategies * Improve platform scalability, reliability, and performance * Implement monitoring, alerting, and observability best practices * Partner with engineering teams on platform architecture and cloud adoption * Mentor engineers and provide technical leadership

Requirements

  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps * Expert-level Kubernetes experience * Strong Google Cloud Platform (GCP) experience * Expertise with Terraform * Experience with Helm * Multi-cloud exposure, including AWS and Azure * Experience with distributed systems * Python or Bash scripting experience * Experience with Prometheus, Grafana, Splunk, or OpenTelemetry

About the company

What Makes HTC A Great Place To Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You’ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

LI-Onsite #LI-DT1 #Hiring

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all