Lead Site Reliability Engineer Observability

Tata Consultancy Services Limited
San Jose, CA, United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$64,000.0 - $130,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Computer Programming Linux DevOps Elasticsearch Python (Programming Language) Log Analysis Reliability Engineering Ansible
+15 more
Prometheus Ruby Google Cloud Istio Grafana Kubernetes Infrastructure Automation Frameworks Apache Kafka Kibana Terraform Splunk Network Server Dynatrace Docker Golang

Job description

Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul. Roles & Responsibilities:

  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools. Nice to have skills:

  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.

Requirements

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
  • Experience supporting FedRAMP or regulated environments., In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person ====, a U.S. Permanent Resident (i.e., a “”), or a Political Asylee or Refugee.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

Videos

See all

Related articles

See all