DevOps/Site Reliability Engineer (SRE) - Hybrid

The Smart
Westlake, TX, United States
10 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$83,200.0 - $166,400.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Microsoft Windows Agile Methodology Automation of Tests Bash Shell C Sharp (Programming Language) Databases Dynamic Host Configuration Protocol DevOps Distributed Systems Domain Name System (DNS)
+24 more
Fault Tolerance Monitoring of Systems Python (Programming Language) PostgreSQL Linux System Administration Linux Servers Windows Servers MongoDB Routing Oracle (Applications) Windows PowerShell Systems Development Life Cycle Release Management Reliability Engineering YAML Datadog Scripting System Availability Grafana Reliability of Systems Firewalls (Computer Science) Information Technology Influxdb Performance Monitor

Job description

As a DevOps/Site Reliability Engineer (SRE), you will be responsible for maintaining, automating, monitoring, and optimizing enterprise-scale infrastructure and applications to ensure high availability, reliability, and operational excellence. You will leverage automation, observability, and system administration expertise to proactively identify issues, improve system resiliency, and support mission-critical environments., * Design, implement, and support enterprise monitoring, alerting, and observability solutions.

  • Develop automation scripts and tools to streamline operational processes and deployments.
  • Build dashboards and proactive monitoring capabilities using enterprise monitoring platforms.
  • Manage and support Windows and Linux server environments.
  • Troubleshoot infrastructure, application, network, and performance issues.
  • Support deployment, release management, and operational activities across environments.
  • Maintain highly available and scalable distributed systems.
  • Collaborate with development teams to improve system reliability and operational efficiency.
  • Implement process improvements aligned with DevOps and SRE best practices.
  • Support incident response, root cause analysis, and service restoration activities.

Requirements

  • 6-8 years of experience in enterprise infrastructure administration, support, and Site Reliability Engineering.
  • Strong experience with automation scripting using PowerShell, Python, Bash, C#, .NET, Java, YAML, or similar technologies.
  • Experience building monitoring dashboards and alerting solutions using Grafana, Datadog, InfluxDB, Moogsoft, ThousandEyes, or similar tools.
  • Strong understanding of SDLC practices, operational processes, and continuous improvement methodologies.
  • Hands-on experience with Windows Server 2016/2019/2022 and Linux administration.
  • Strong knowledge of networking concepts including DNS, DHCP, firewalls, routing, and connectivity troubleshooting.
  • Experience supporting large-scale distributed systems and high-availability architectures.
  • Knowledge of SQL Server, Oracle, MongoDB, PostgreSQL, or similar database platforms.
  • Bachelor’s degree in Computer Science, Information Technology, or a related field.
  • Strong ownership mindset, troubleshooting abilities, communication skills, and experience in Agile and financial services environments preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 Ā· WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum Ā· WWC Europe 2026

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos Ā· LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley Ā· WWC 2021

1:31 min

Orchestrating generative configurations using standardized YAML files

Han Xiao Ā· WWC 2022

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann Ā· WWC 2023

Videos

See all

Related articles

See all