Site Reliability Engineer
Shrive Technologies Llc
Schaumburg, United States
1 day ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Data Analysis
Microsoft Azure
Cloud Computing
Continuous Integration
Software Debugging
DevOps
Disaster Recovery
Distributed Systems
Github
Monitoring of Systems
Python (Programming Language)
+31 more
NumPy
Performance Tuning
Windows PowerShell
Reliability Engineering
Power BI
Ansible
Tensorflow
Prometheus
SQL Databases
Tableau (Software)
Datadog
Data Logging
Scripting
Google Cloud
Pytorch
Grafana
Software Troubleshooting
Reliability of Systems
Infrastructure as Code (IaC)
Cloudformation
Pandas
Containerization
Gitlab-ci
Scikit Learn
Infrastructure Automation Frameworks
Deployment Automation
Terraform
Splunk
Dynatrace
Docker
Jenkins
Job description
Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management, * Monitor, maintain, and improve the reliability, availability, and performance of production systems.
- Design and implement monitoring, alerting, logging, and observability solutions.
- Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
- Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
- Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
- Perform capacity planning, performance tuning, and scalability assessments.
- Support CI/CD pipelines and deployment automation initiatives.
- Implement high-availability, disaster recovery, and failover strategies.
Requirements
- Strong experience with Linux/Unix administration.
- Proficiency in scripting languages such as Python, Shell, or PowerShell.
- Hands-on experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
- Experience with containerization technologies such as Docker and Kubernetes.
- Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
- Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
- Strong troubleshooting, debugging, and problem-solving skills.
-
Understanding of networking, security, and distributed systems concepts. Experience
- 5-10+ years of overall IT experience.
- 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
LM
Luis Minvielle
Top-Paying Tech Jobs (with Salaries)
over 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago
KM
Kaleb McKelvey
The Best Software Developer Blogs to Read
over 3 years ago