Site Reliability Engineer

Ethos Group Inc
Irving, TX, United States
26 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Job source

Tech stack

Application Performance Management Systems Engineering Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Linux DevOps Monitoring of Systems Python (Programming Language) Windows PowerShell Reliability Engineering
+10 more
Runbook Software Engineering Scripting Google Cloud Cloudformation Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Terraform

Job description

Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team.

This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development and infrastructure teams to improve platform performance, automation, and operational excellence.

As a Site Reliability Engineer, you will play a critical role in maintaining system uptime, optimizing application performance, automating operational processes, and supporting a modern cloud-based technology environment.

Key Responsibilities

  • Design, implement, and maintain reliable, scalable, and secure cloud infrastructure
  • Monitor application and system performance to ensure maximum availability
  • Automate operational tasks and deployment processes
  • Respond to incidents, troubleshoot production issues, and drive root cause analysis
  • Develop and maintain monitoring, alerting, and observability solutions
  • Partner with software engineering teams to improve application reliability and performance
  • Support CI/CD pipelines and deployment automation initiatives
  • Create operational documentation, runbooks, and best practices
  • Participate in on-call rotations and incident response activities
  • Continuously identify opportunities for system improvements and increased efficiency

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, or related field, or equivalent experience
  • Experience in Site Reliability Engineering, DevOps, Systems Engineering, or Cloud Infrastructure roles
  • Strong understanding of Linux and cloud technologies
  • Direct hands-on experience with AWS, Azure, or Google Cloud Platform
  • Experience writing and maintaining infrastructure-as-code tools such as Azure resource templates, Terraform, or CloudFormation
  • Direct hands-on experience supporting CI/CD pipelines and automation tools
  • Direct hands-on experience with production grade Kubernetes deployments
  • Experience with monitoring and observability platforms
  • Strong scripting skills in Python, PowerShell, Bash, or similar languages
  • Excellent troubleshooting and problem-solving abilities
  • Strong communication and collaboration skills

Preferred Qualifications

  • Direct hands-on experience supporting large-scale production environments
  • Deep knowledge of networking, security, and cloud architecture principles
  • Experience with incident management and root cause analysis
  • Relevant cloud, Kubernetes, or DevOps certifications

Benefits & conditions

Pulled from the full job description 401(k) Health insurance 401(k) matching Paid time off Vision insurance Dental insurance Life insurance, * 401(k)

  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Life insurance
  • Paid time off
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all