TELECOMMUTE Lead Site Reliability Engineer
UKUND INC
United States
14 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Microsoft Azure
Bash Shell
Cloud Computing
Computer Programming
Linux
DevOps
Elasticsearch
Python (Programming Language)
Log Analysis
Reliability Engineering
Ansible
+16 more
Prometheus
Ruby
Google Cloud
Istio
Grafana
Kubernetes
Infrastructure Automation Frameworks
Apache Kafka
Kibana
Terraform
Splunk
Network Server
Dynatrace
Docker
Service Stack
Golang
Job description
- Design, deploy, and operate enterprise observability platforms.
- Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
- Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
- Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
- Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
- Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
- Automate infrastructure using Terraform and configuration management tools.
Requirements
- 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
- Hands-on experience administering Splunk Enterprise or Splunk Cloud.
- Strong knowledge of Splunk SPL.
- Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
- Experience implementing metrics, logs, and traces as part of a modern observability strategy.
- Experience with Terraform and Infrastructure as Code.
- Programming experience in Python, Go, Ruby, or Bash.
Preferred Qualifications
-
Splunk certification.
- Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
- Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
about 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
EM
Eli McGarvie
The Best Job Search Websites of 2025
over 2 years ago
EM
Eli McGarvie
Best Job Boards for Remote Work for Developers
almost 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
DC
Daniel Cranney
Mastering Remote Work: Tips for Developers
over 1 year ago