Reliability Engineer

Randstad
Roanoke, TX, United States
25 days ago

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$139,360.0 - $141,440.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon Web Services Microsoft Azure Big Data Bioinformatics Mobile Application Development Cloud Computing Data Visualization Query Languages Identity and Access Management Python (Programming Language) Node.Js
+19 more
Systems Development Life Cycle Power BI Prometheus Shell Script Software Engineering Tableau (Software) Datadog Data Logging Scripting Grafana Git Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Splunk Software Version Control Docker Jenkins

Job description

job summary: Advance enterprise resiliency through improved recovery capabilities.

Reduce recovery time via automation.

Enable rehoused recovery into new datacenters.

Strengthen platform reliability through data protection design.

location: Westlake, Texas job type: Contract salary: $67 - 68 per hour work hours: 8am to 5pm education: Bachelors

responsibilities:

  • Bachelor’s Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
  • Production experience running Cloud and on-prem Storage workloads at scale
  • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
  • Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
  • Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
  • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
  • 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
  • Experience building and deploying Docker images including Docker Compose
  • Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
  • Experience with distributed version control systems, Git preferred
  • Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
  • Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
  • Experience in incident/crisis management and supporting critically important applications

qualifications: Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)

Ability to automate with various scripting languages (Python, Shell scripting, etc.)

Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)

Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.

At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com.

Pay offered to a successful candidate will be based on several factors including the candidate’s education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).

This posting is open for thirty (30) days.

,

  • Bachelor’s Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
  • Production experience running Cloud and on-prem Storage workloads at scale
  • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
  • Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
  • Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
  • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
  • 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
  • Experience building and deploying Docker images including Docker Compose
  • Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
  • Experience with distributed version control systems, Git preferred
  • Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
  • Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
  • Experience in incident/crisis management and supporting critically important applications

Requirements

Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.) Ability to automate with various scripting languages (Python, Shell scripting, etc.) Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.randstadusa.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all