SRE DevOps Engineer

Covetus, LLC
Westlake, TX, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
9 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Amazon Web Services Microsoft Azure Big Data Continuous Integration Query Languages Linux DevOps Distributed Systems Python (Programming Language) Node.Js Reliability Engineering
+20 more
Power BI Software Deployment Software Engineering Software Organization Datadog Data Logging Scripting Delivery Pipeline Grafana Reliability of Systems Cloudformation Kubernetes Infrastructure Automation Frameworks Deployment Automation Hardware Infrastructure Api Gateway Terraform Splunk Software Version Control Jenkins

Job description

We are seeking a highly motivated Site Reliability Engineer to help build and operate reliable, scalable and secure services across our platform. This role is designed for someone who combines strong DevOps practices, modern SRE principles and software engineering experience to improve system reliability, automate operations and support high-availability production environments.

The ideal candidate will be passionate about building resilient systems, improving developer productivity and driving operational excellence through automation, observability and engineering best practices. This role partners closely with engineering, platform and product teams to ensure services are built for reliability from the start and remain performant, stable and supportable at scale.

Key Skills - Node.js, Python, DevOps, Jenkins, AWS., * Design, build and operate resilient, scalable systems using DevOps, SRE and software development best practices.

  • Deliver high-availability services through automation, infrastructure as code and proactive reliability engineering.
  • Improve monitoring, logging, alerting and observability for distributed systems.
  • Support CI/CD automation, deployment workflows and production tooling to reduce operational toil.
  • Drive incident response, root cause analysis and recovery improvements to minimize downtime.
  • Partner with engineering teams to embed reliability into the software development lifecycle.
  • Automate provisioning, configuration and self-healing across cloud and on-prem environments.
  • Validate resiliency and performance through testing, chaos engineering and capacity planning.

Requirements

  • 9+ years of demonstrated experience developing and designing products around Site Reliability Engineering (SRE) principles to improve stability and platform availability for containerized workloads and on-premises services using Kubernetes.
  • Experience managing and interpreting large datasets using query languages and creating dashboards and reports with Power BI and Grafana.
  • Strong background in managing cloud and on-premises infrastructure using Infrastructure as Code tools, including Terraform and CloudFormation.
  • Hands-on experience building, operating, monitoring, logging and alerting distributed systems at scale using Datadog and Splunk.
  • Experience supporting DevOps practices for service delivery and operations using Jenkins, Azure DevOps, Team Foundation Version Control and CI/CD automation.
  • Experience developing software and automation solutions to support application delivery, operations and repeatable business processes using Python.
  • Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway
  • Strong development experience in scripting, automation and integration across Linux and Windows-based environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all