Lead Site Reliability Engineer

Skanda Solutions Llc
Orlando, FL, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$110,000.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Application Lifecycle Management Microsoft Azure Bash Shell Cloud Computing Databases Continuous Integration DevOps Disaster Recovery Distributed Systems Github
+23 more
Python (Programming Language) PostgreSQL Enterprise Messaging Systems MongoDB Redis Reliability Engineering Prometheus Vault (Revision Control System) YAML Data Logging Google Cloud Spring Cloud System Availability Multi-Cloud Gitlab-ci Kubernetes Deployment Automation Apache Kafka Terraform Splunk Appdynamics Dynatrace Jenkins

Job description

The ideal candidate combines deep expertise in Site Reliability Engineering, cloud infrastructure, Kubernetes, and Infrastructure as Code with strong leadership skills. You will work alongside platform engineers, architects, and development teams to build resilient systems, improve automation, and ensure high availability across a multi-cloud environment., * Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure.

  • Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code.
  • Develop scalable platform solutions capable of maintaining 99.99% service availability.
  • Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices.
  • Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams.
  • Design and implement automated deployment pipelines using modern CI/CD tools, including Harness.
  • Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts.
  • Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies.
  • Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements.
  • Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance.
  • Troubleshoot complex production issues across distributed systems and cloud-native applications.
  • Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency.
  • Ensure infrastructure aligns with security, governance, and compliance standards.

Requirements

  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Expert-level experience administering and operating Kubernetes in production environments.
  • Strong experience with Helm for Kubernetes application management.
  • Advanced experience using Terraform for Infrastructure as Code.
  • Hands-on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms.
  • Experience supporting production workloads across Google Cloud Platform, AWS, and Azure.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness.
  • Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Strong troubleshooting skills within distributed systems and cloud-native environments.
  • Experience working within Agile development environments.
  • Excellent communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences.

Benefits & conditions

Pay: $110,000.00 - $140,000.00 per year

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 · WWC 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

1:31 min

Orchestrating generative configurations using standardized YAML files

Han Xiao · WWC 2022

Videos

See all

Related articles

See all