Lead DevOps Engineer

SECURE ROOTS
Chicago, IL, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Amazon Elastic Compute Cloud Cloud Computing Data Stores DevOps Disaster Recovery Distributed Systems Monitoring of Systems Python (Programming Language) NoSQL Performance Tuning
+19 more
Reliability Engineering Ansible Prometheus Scala (Programming Language) SQL Databases Datadog Data Logging Grafana Event Driven Architecture Containerization Kubernetes Hashicorp Apache Kafka Terraform Dynatrace Pagerduty Jenkins Golang Microservices

Job description

  • Lead DevOps and Site Reliability Engineering (SRE) initiatives across enterprise platforms.
  • Design and implement CI/CD pipelines in Kubernetes environments.
  • Manage and support multi-region Kubernetes and Kafka platforms.
  • Define and drive SLOs, SLIs, Error Budgets, and operational excellence practices.
  • Build observability solutions using metrics, logging, and distributed tracing.
  • Drive automation initiatives using Terraform, Ansible, and Infrastructure-as-Code.
  • Lead incident response, post-mortems, runbook creation, and on-call processes.
  • Partner with Product, Infrastructure, Security, and Architecture teams to deliver scalable and secure solutions.
  • Support capacity planning, failover strategies, and performance optimization.

Requirements

  • AWS (EC2), Kubernetes, Kafka, Jenkins, Terraform, Ansible, HashiCorp Vault
  • Prometheus, Grafana, OpenTelemetry, Datadog, or similar monitoring tools
  • Chaos engineering principles and tooling (e.g., Chaos Monkey, Gremlin, LitmusChaos)
  • PagerDuty, OpsGenie, or other incident management platforms
  • Microservices, distributed systems, and event-driven architectures
  • Cloud networking and security
  • SQL, NoSQL, and in-memory data stores
  • Development experience with Java, Python, Scala, or Golang
  • Multi-region/high-availability architecture and disaster recovery
  • Strong understanding of SRE principles, automation, and cloud-native technologies, * 8+ years of experience building large-scale, data-centric solutions
  • 8+ years of recent experience in DevOps, SRE, or Platform Engineering environments
  • Experience working in fast-paced, highly available enterprise environments
  • Strong communication and collaboration skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

4:30 min

Transitioning from software engineering into a DevOps trajectory

Davide Imola Davide Imola · LIVE

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all