Site Reliability Engineer OpenSearch

Jtsi (johnson Technology Systems, Inc.
United States
4 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Shift work
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Application Layers Application Services Automation of Tests Big Data Cloud Computing Databases Shard (Database Architecture) DevOps Disaster Recovery
+22 more
Distributed Systems Amazon DynamoDB Identity and Access Management Internet Protocol Node.Js Performance Tuning Query Optimization Reliability Engineering Cloud Services Strategies of Testing Apache Zookeeper Cloud Platform System Data Ingestion Cloud Monitoring Grafana Indexer Amazon Virtual Private Cloud (VPC) Amazon Relational Database Service Information Technology Performance Monitor Azure AKS Cloudwatch

Job description

Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environment Architect, build, deploy, and maintainhigh-performance OpenSearch clusters and platforms from the ground up Administer and optimize OpenSearch environments forhigh availability, resiliency, scalability, security, and performance Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization Analyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layers Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement Maintain the integrity and security of servers, systems, and OpenSearch platform infrastructure Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery Develop and maintain monitoring policies, alerting standards, operational runbooks, and support procedures Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services Ensure proper resource allocation and capacity planning across compute, memory, storage, and network resources Partner with product development and engineering teams to design and enhance service reliability and operational readiness Develop and implement testing strategies and document results for platform changes and operational improvements Support log ingestion, index management, retention policies, lifecycle management, and search performance tuning Work in a diverse environment and cross-train with other global team members Participate in an on-call rotation and support weekend or after-hours operational needs as required

Requirements

We are seeking a Site Reliability Engineer OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services. This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment.

The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments., Expert withKubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services. Proven hands-on expertisedesigning, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments Strong experience withOpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting Experience withindex design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery Strong understanding ofdistributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures Expertise withGit Expertise withConcourse, including setup, management, and troubleshooting of new pipelines Expertise withLinux, specificallySUSEandUbuntu Expertise withKafka, Zookeeper, and Big Data technologies Expert in development of automation for testing, deployment, scalability, and management of cloud services Expertise with building, implementing, and/or supporting cloud monitoring tools Expert knowledge ofcloud computing, infrastructure operations, and databases Expert understanding ofweb services, networking, virtualization, and internet protocols Ability to multitask and handle various projects, deadlines, and changing priorities Excellent communication and prioritization skills Expertise with security fundamentals as they pertain toSaaS multi-tenant application systems Strong interpersonal, presentation, and customer service skills

Desired Qualifications Experience with AWS services includingRoute 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC Experience deploying and operatingOpenSearch in AWS-based environments Experience withCloud Foundry-based environments Experience withJenkins,Chef, and/orTerraform Exposure to and understanding of troubleshootingIP networks and application stacks Experience with observability tools such asPrometheusandGrafana Experience withlog ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls Familiarity withcapacity forecasting, performance benchmarking, and resilience testingfor distributed search platforms

Education BS/BAdegree in Computer Science, Management Information Systems, or related IT discipline preferred Allowable substitution:An additional four (4) years of experience may be substituted for a BS/BA degree 8+ years of experience

Additional Requirements Participation in anon-call rotationfor handlingP1 incidentsis required Flexible schedule which may includeweekend or after-hours work Ability to work effectively in adiverse, collaborative, and globally distributed team environment

About the company

Established in 2003, JTSi is a Professional IT & Engineering Services provider with years of documented experience in the Information Technology and Engineering services field. JTSi has a proven track record for successfully delivering mission critical Professional services to the Government and the industry. JTSi SAP team delivers solutions to its clients by clearly understanding their core business problems. We deliver quality services at equitable rates and focus on constant improvement in all areas of our operation, austerely complying to the customer s desire. We view our-selves more as a business partner than a mere provider of consulting services. At JTSi customer is always first and partnering is our means to customer satisfaction. We do what we say!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:02 min

Provisioning a cluster with managed OpenSearch

Olena Kutsenko · WWC 2022

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:55 min

Technical challenges driving the OpenSearch database migration

Dharin Shah Dharin Shah · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all