> Markdown version of [/jobs/ext/2008268-site-reliability-engineer-opensearch-remote](https://www.wearedevelopers.com/jobs/ext/2008268-site-reliability-engineer-opensearch-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (OpenSearch) - Remote - **Company:** Information Consulting Services - **Location:** Herndon, VA, United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Application Layers, Automation of Tests, Ubuntu (Operating System), Software as a Service, Cloud Computing, Cloud Foundry, Databases, Linux, DevOps, Disaster Recovery, Distributed Systems, Amazon DynamoDB, Identity and Access Management, Internet Protocol, Performance Tuning, Reliability Engineering, Cloud Services, Prometheus, Runbook, Virtualization Technology, Web Services, Apache Zookeeper, SUSE Linux, Data Ingestion, Cloud Monitoring, System Availability, Grafana, Indexer, Amazon Virtual Private Cloud (VPC), Git, Concourse, Amazon Relational Database Service, Kubernetes, Apache Kafka, Route53, Cloudwatch, Terraform, Jenkins - **Published:** August 9, 2026 - **Apply:** https://www.wayup.com/i-j-Site-Reliability-Engineer-OpenSearch-Remote-Information-Consulting-Services-743779963842532/ ## About the Role + US citizenship required; dual citizenship not permitted. + 8 years of experience in SRE/DevOps/cloud operations with distributed systems. + Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production. + Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services). + Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR. + Strong Linux expertise (SUSE and Ubuntu). + Expertise with Git and Concourse (pipeline setup, management, troubleshooting). + Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management. + Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols. + Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask. Preferred Qualifications + AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS. + Experience with Cloud Foundry environments. + Experience with Jenkins, Chef, and/or Terraform. + Experience with Prometheus and Grafana. + Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls. + Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms. Work Environment + Collaborative, globally distributed team with cross-training opportunities. ## Description Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team. Key Responsibilities + Provision, build, deploy, monitor, operate, and support cloud services in a global team environment. + Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up. + Optimize OpenSearch for high availability, resiliency, scalability, security, and performance. + Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization. + Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation. + Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure. + Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery. + Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures. + Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services. + Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness. + Support log ingestion, index management, lifecycle/retention, and search performance tuning. + Participate in an on-call rotation; support occasional weekend/after-hours needs. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Search and aggregations made easy with OpenSearch and NodeJS](https://www.wearedevelopers.com/videos/490-search-and-aggregations-made-easy-with-opensearch-and-nodejs) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers)