> Markdown version of [/jobs/ext/1887171-lead-site-reliability-engineer-observability](https://www.wearedevelopers.com/jobs/ext/1887171-lead-site-reliability-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer Observability - **Company:** Tata Consultancy Services Limited - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $64,000.0 - $130,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Computer Programming, Linux, DevOps, Elasticsearch, Python (Programming Language), Log Analysis, Reliability Engineering, Ansible, Prometheus, Ruby, Google Cloud, Istio, Grafana, Kubernetes, Infrastructure Automation Frameworks, Apache Kafka, Kibana, Terraform, Splunk, Network Server, Dynatrace, Docker, Golang - **Published:** July 31, 2026 - **Apply:** https://www.dice.com/job-detail/a2ca3d16-2629-4bcd-9259-a1ce612c009d ## About the Role * 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps. * Hands-on experience administering Splunk Enterprise or Splunk Cloud. * Strong knowledge of Splunk SPL. * Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka. * Experience implementing metrics, logs, and traces as part of a modern observability strategy. * Experience with Terraform and Infrastructure as Code. * Programming experience in Python, Go, Ruby, or Bash. * Splunk certification. * Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies. * Experience supporting FedRAMP or regulated environments., In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person ====, a U.S. Permanent Resident (i.e., a ""), or a Political Asylee or Refugee. ## Description Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul. Roles & Responsibilities: * Design, deploy, and operate enterprise observability platforms. * Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers. * Deploy and operate large-scale Elasticsearch clusters for log analytics and search. * Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry. * Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies. * Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions. * Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo. * Automate infrastructure using Terraform and configuration management tools. Nice to have skills: * Splunk certification. * Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies. * Experience supporting FedRAMP or regulated environments. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)