> Markdown version of [/jobs/ext/2025912-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2025912-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Hptech Inc. - **Location:** Schaumburg, IL, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Data Analysis, Build Automation, Microsoft Azure, Cloud Computing, Continuous Integration, Software Debugging, DevOps, Disaster Recovery, Distributed Systems, Github, Monitoring of Systems, Python (Programming Language), NumPy, Performance Tuning, Windows PowerShell, Reliability Engineering, Power BI, Ansible, Tensorflow, Prometheus, Software Engineering, SQL Databases, Tableau (Software), Datadog, Data Logging, Scripting, Google Cloud, Pytorch, Grafana, Software Troubleshooting, Reliability of Systems, Infrastructure as Code (IaC), Cloudformation, Pandas, Containerization, Gitlab-ci, Scikit Learn, Infrastructure Automation Frameworks, Deployment Automation, Terraform, Splunk, Dynatrace, Docker, Jenkins - **Published:** August 11, 2026 - **Apply:** https://www.dice.com/job-detail/92b4d455-5224-4b1f-8a9a-c1ac99476e34 ## About the Role Note - We are seeking a highly motivated Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and availability of critical production systems. The ideal candidate will combine software engineering and operations expertise to build automation, improve system resilience, reduce operational toil, and enhance service reliability. Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management, * Strong experience with Linux/Unix administration. * Proficiency in scripting languages such as Python, Shell, or PowerShell. * Hands-on experience with cloud platforms (AWS, Azure, or Google Cloud Platform). * Experience with containerization technologies such as Docker and Kubernetes. * Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog. * Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps. * Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation. * Strong troubleshooting, debugging, and problem-solving skills. * Understanding of networking, security, and distributed systems concepts. Experience * 8-10+ years of overall IT experience. * 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles. ## Description * Monitor, maintain, and improve the reliability, availability, and performance of production systems. * Design and implement monitoring, alerting, logging, and observability solutions. * Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. * Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC). * Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews. * Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience. * Perform capacity planning, performance tuning, and scalability assessments. * Support CI/CD pipelines and deployment automation initiatives. * Implement high-availability, disaster recovery, and failover strategies. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)