Site Reliability Engineer

Hptech Inc.
Schaumburg, IL, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Data Analysis Build Automation Microsoft Azure Cloud Computing Continuous Integration Software Debugging DevOps Disaster Recovery Distributed Systems Github Monitoring of Systems
+33 more
Python (Programming Language) NumPy Performance Tuning Windows PowerShell Reliability Engineering Power BI Ansible Tensorflow Prometheus Software Engineering SQL Databases Tableau (Software) Datadog Data Logging Scripting Google Cloud Pytorch Grafana Software Troubleshooting Reliability of Systems Infrastructure as Code (IaC) Cloudformation Pandas Containerization Gitlab-ci Scikit Learn Infrastructure Automation Frameworks Deployment Automation Terraform Splunk Dynatrace Docker Jenkins

Job description

  • Monitor, maintain, and improve the reliability, availability, and performance of production systems.
  • Design and implement monitoring, alerting, logging, and observability solutions.
  • Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
  • Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
  • Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
  • Perform capacity planning, performance tuning, and scalability assessments.
  • Support CI/CD pipelines and deployment automation initiatives.
  • Implement high-availability, disaster recovery, and failover strategies.

Requirements

Note - We are seeking a highly motivated Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and availability of critical production systems. The ideal candidate will combine software engineering and operations expertise to build automation, improve system resilience, reduce operational toil, and enhance service reliability.

Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management, * Strong experience with Linux/Unix administration.

  • Proficiency in scripting languages such as Python, Shell, or PowerShell.
  • Hands-on experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
  • Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
  • Strong troubleshooting, debugging, and problem-solving skills.
  • Understanding of networking, security, and distributed systems concepts. Experience

  • 8-10+ years of overall IT experience.
  • 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all