Site Reliability Engineer (SRE)

Info Way Solutions LLC
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Remote

Tech stack

Amazon Web Services (AWS)
Data analysis
Azure
Bash
Cloud Computing
Computer Programming
Linux
DevOps
Elasticsearch
Python
Linux System Administration
Reliability Engineering
Ansible
Prometheus
Ruby
Data Logging
Google Cloud Platform
Cloud Platform System
Istio
System Availability
Grafana
Reliability of Systems
Infrastructure as Code (IaC)
GIT
Kubernetes
Infrastructure Automation Frameworks
Information Technology
Kafka
Linkerd (Service Mesh)
Kibana
Terraform
Splunk
Network Server
Dynatrace
Docker
ELK
Programming Languages

Job description

We are seeking a highly skilled Senior Site Reliability Engineer (SRE) - Observability Platform to design, build, and operate enterprise-scale observability solutions across cloud environments. The ideal candidate will have extensive experience with logging, metrics, distributed tracing, monitoring, and infrastructure automation, helping drive reliability, scalability, and operational excellence for mission-critical applications. This role requires expertise in modern observability technologies including Splunk, Elasticsearch, Prometheus, Grafana, OpenTelemetry, Grafana Tempo, Kafka, and Terraform., * Design, deploy, and manage enterprise observability platforms supporting large-scale cloud infrastructure.

  • Administer and maintain Splunk Enterprise and Splunk Cloud, including:

  • Indexers

  • Search Head Clusters (SHC)

  • Heavy Forwarders

  • Deployment Servers

Deploy, optimize, and manage Elasticsearch/ELK clusters for centralized logging and search.

Design and support distributed tracing solutions using Grafana Tempo and OpenTelemetry.

Build and maintain end-to-end observability pipelines for logs, metrics, and traces.

Define instrumentation standards, trace retention policies, and monitoring best practices.

Scale and manage monitoring platforms including Prometheus, Grafana, Kafka, and Tempo.

Develop dashboards, alerts, reports, analytics, and trace visualizations using:

  • Splunk SPL
  • Grafana
  • Kibana
  • Tempo

Automate infrastructure deployment using Terraform and Infrastructure as Code (IaC).

Collaborate with development, DevOps, and platform engineering teams to improve system reliability and performance.

Perform troubleshooting, root cause analysis, and incident response for production environments.

Ensure high availability, security, scalability, and compliance of observability platforms., * Splunk Enterprise

  • Splunk Cloud
  • Splunk SPL
  • Elasticsearch
  • ELK Stack
  • Kibana
  • Prometheus
  • Grafana
  • Grafana Tempo
  • OpenTelemetry
  • Distributed Tracing
  • Kafka

Cloud & Infrastructure

  • AWS
  • Azure
  • Google Cloud Platform (Google Cloud Platform)
  • Kubernetes
  • Docker
  • Linux

Automation & DevOps

  • Terraform
  • Infrastructure as Code (IaC)
  • Ansible
  • CI/CD Pipelines
  • Git

Programming

  • Python
  • Go
  • Bash
  • Ruby

Requirements

  • Bachelor''s degree in Computer Science, Information Technology, or a related field (or equivalent experience).

  • 7+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Cloud Infrastructure.

  • Hands-on administration experience with Splunk Enterprise or Splunk Cloud.

  • Strong expertise in Splunk Search Processing Language (SPL).

  • Experience with:

  • Elasticsearch / ELK Stack

  • Prometheus

  • Grafana

  • Grafana Tempo

  • OpenTelemetry

  • Distributed Tracing

  • Kafka

Strong understanding of modern observability practices involving metrics, logs, and traces.

Experience with Terraform and Infrastructure as Code (IaC).

Proficiency in one or more scripting/programming languages:

  • Python
  • Go
  • Ruby
  • Bash

Strong Linux administration and troubleshooting skills., * Splunk Certified Administrator or Splunk Architect certification.

  • Experience with Kubernetes and Docker.
  • Experience with AWS, Azure, or Google Cloud Platform (Google Cloud Platform).
  • Experience using Ansible, Consul, and CI/CD pipelines.
  • Knowledge of Service Mesh technologies (Istio, Linkerd, etc.).
  • Experience supporting FedRAMP High, IL-5, or other regulated environments.
  • Strong understanding of security, compliance, and cloud governance., * Strong analytical and troubleshooting skills.
  • Experience managing enterprise-scale monitoring platforms.
  • Excellent communication and cross-functional collaboration.
  • Ability to automate operational processes.
  • Strong incident management and root cause analysis capabilities.
  • Experience designing scalable, highly available cloud-native observability solutions., * Must be a U.S. Citizen or U.S. National (U.S. Person).
  • Work must be performed from within the United States.
  • Ability to support environments subject to FedRAMP High and Impact Level 5 (IL-5) security requirements.

Apply for this position