Site Reliability Engineer (OpenSearch) - Remote

Information Consulting Services
Herndon, VA, United States
3 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Shift work
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Application Layers Automation of Tests Ubuntu (Operating System) Software as a Service Cloud Computing Cloud Foundry Databases Linux DevOps
+29 more
Disaster Recovery Distributed Systems Amazon DynamoDB Identity and Access Management Internet Protocol Performance Tuning Reliability Engineering Cloud Services Prometheus Runbook Virtualization Technology Web Services Apache Zookeeper SUSE Linux Data Ingestion Cloud Monitoring System Availability Grafana Indexer Amazon Virtual Private Cloud (VPC) Git Concourse Amazon Relational Database Service Kubernetes Apache Kafka Route53 Cloudwatch Terraform Jenkins

Job description

Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team. Key Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
  • Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
  • Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
  • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
  • Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
  • Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
  • Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
  • Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
  • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
  • Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
  • Support log ingestion, index management, lifecycle/retention, and search performance tuning.
  • Participate in an on-call rotation; support occasional weekend/after-hours needs.

Requirements

  • US citizenship required; dual citizenship not permitted.
  • 8 years of experience in SRE/DevOps/cloud operations with distributed systems.
  • Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
  • Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
  • Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
  • Strong Linux expertise (SUSE and Ubuntu).
  • Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
  • Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
  • Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
  • Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask. Preferred Qualifications

  • AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
  • Experience with Cloud Foundry environments.
  • Experience with Jenkins, Chef, and/or Terraform.
  • Experience with Prometheus and Grafana.
  • Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
  • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms. Work Environment

  • Collaborative, globally distributed team with cross-training opportunities.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:02 min

Provisioning a cluster with managed OpenSearch

Olena Kutsenko · WWC 2022

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:19 min

Core OpenSearch cluster architecture and terminology

Olena Kutsenko · WWC 2022

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all