Site Reliability Engineer (Kubernetes & Observability)

THE JUDGE GROUP, INC.
Tampa, FL, United States
12 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Bash Shell Cloud Computing Cloud Engineering Continuous Integration DevOps Distributed Systems Monitoring of Systems HP Systems Insight Manager Python (Programming Language) Linux System Administration Reliability Engineering
+22 more
Prometheus Server Administration Service Discovery Data Storage Management Load Balancing Autoscaling System Availability Grafana IT Architecture Infrastructure as Code (IaC) Containerization Git Flow Kubernetes Infrastructure Automation Frameworks Deployment Automation Performance Monitor Bitbucket Terraform Splunk Dynatrace Docker Jenkins

Job description

We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes infrastructure, observability platforms, and automation to support and improve large-scale production systems. This role requires a hands-on engineer who can build, scale, and maintain Kubernetes environments while driving reliability, monitoring, and operational excellence. You will work closely with development, infrastructure, production support, and platform engineering teams to design scalable solutions, automate operational processes, and ensure high availability across mission-critical applications., * Build and administer Kubernetes clusters across multiple environments

  • Design and implement containerized solutions using Kubernetes and Docker
  • Troubleshoot production issues involving networking, storage, cluster performance, and application reliability
  • Support incident management, root cause analysis, and proactive reliability improvements
  • Develop Infrastructure as Code (IaC) solutions using Terraform
  • Build and maintain CI/CD pipelines using Jenkins, Bitbucket, and GitOps methodologies
  • Automate infrastructure provisioning, deployments, and operational workflows
  • Implement observability and monitoring solutions using Dynatrace, Splunk, Prometheus, Grafana, and OpenTelemetry
  • Create dashboards, alerts, and telemetry solutions to improve application visibility and performance monitoring
  • Collaborate with engineering teams to improve system scalability, availability, and operational efficiency
  • Develop automation tools and scripts using Python, Go, or Bash
  • Support platform modernization initiatives and help establish best practices for SRE and Cloud-Native operations

Requirements

Core Skills

  • 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering
  • Strong hands-on Kubernetes experience with cluster administration and platform build-out
  • Extensive experience supporting production environments and resolving system-level issues
  • Proven expertise in Infrastructure as Code (IaC) using Terraform
  • Strong understanding of reliability engineering principles, system uptime, scalability, and availability

Kubernetes & Container Platforms

  • Hands-on experience building Kubernetes environments from the ground up
  • Expertise with Kubernetes cluster administration, workload orchestration, networking, and storage management
  • Strong understanding of containerization technologies including Docker
  • Experience managing Kubernetes deployments in enterprise environments
  • Knowledge of service discovery, ingress controllers, load balancing, and autoscaling

CI/CD & Tools

  • Experience building and maintaining CI/CD pipelines using Jenkins and Bitbucket
  • Familiarity with GitOps deployment methodologies
  • Experience automating infrastructure provisioning and deployments
  • Ability to create self-service automation solutions that reduce manual intervention

Infrastructure & Cloud

  • Strong understanding of infrastructure architecture and distributed systems
  • Expertise with Terraform and Infrastructure as Code practices
  • Experience provisioning compute instances and automating infrastructure deployments
  • Understanding of Linux environments and server administration
  • Cloud platform experience is beneficial, but AWS expertise is not required

Observability & Monitoring

  • Hands-on experience with Dynatrace, Splunk, Prometheus, Grafana, or similar monitoring platforms
  • Experience building dashboards, alerts, telemetry, and application monitoring solutions
  • Understanding of OpenTelemetry concepts and distributed tracing

Experience troubleshooting applications using performance metrics, logs, and telemetry data By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively “Judge”) to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge’s Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

9:06 min

Questions on career paths and continuous delivery orchestration platforms

Zan Markan Zan Markan · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all