Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
We are seeking a skilled Site Reliability Engineer (SRE) to support highly available, business-critical applications across on-premises and AWS cloud environments. The ideal candidate will have strong expertise in DevOps, cloud infrastructure, automation, CI/CD, monitoring, and troubleshooting complex production systems. This role offers the opportunity to work with modern cloud technologies and contribute to the reliability, scalability, and performance of enterprise applications., * Manage and optimize data streaming and API components in OpenShift (On-Premises) and AWS.
- Review application APIs and processes to identify performance optimization opportunities.
- Automate testing, including data quality validation, production deployments, and release processes.
- Develop integrations between on-premises, AWS, and third-party tools such as ServiceNow, VersionOne, and Sumo Logic.
- Collaborate with teams to define and implement SLIs and SLOs.
- Monitor production environments, troubleshoot performance issues, conduct root cause analysis, and document findings.
- Design, build, and maintain CI/CD pipelines for application artifacts, APIs, and data processing jobs.
- Configure monitoring, alerting, and observability solutions to enable proactive issue detection.
- Implement AWS security best practices, including IAM, HSM, encryption, and access controls.
- Monitor cloud costs, generate usage reports, and recommend cost optimization strategies.
- Design and implement solutions to address security vulnerabilities and compliance requirements.
- Analyze infrastructure capacity and performance to support scalable and resilient systems.
- Develop backup and disaster recovery strategies for critical applications and data.
- Collaborate with architecture, infrastructure, and application teams to continuously improve system performance, reliability, and security.
Requirements
- Strong experience with AWS cloud services and cloud operations.
- Hands-on experience with OpenShift, CloudFormation, Terraform, Ansible, Shell scripting, and Python.
- Experience with Linux administration and enterprise infrastructure.
- Knowledge of virtualization, networking, load balancers, firewalls, storage, backup, and monitoring tools.
- Experience with CI/CD tools such as GitLab, GitHub, Jenkins, Maven, Gradle, and Nexus.
- Experience with Software Release Management.
- Strong troubleshooting and incident management skills for mission-critical systems.
- Experience with automation, infrastructure orchestration, and configuration management., * Bachelor’s degree in Computer Science or a related technical field (or equivalent experience).
- 3+ years of DevOps/SysOps engineering experience with a focus on AWS.
- 2+ years of application development experience involving data streaming and high-availability applications.
- 1+ year of experience in a Site Reliability Engineering (SRE) environment preferred.
- Overall 4 6 years of IT experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
React Developer Salary [2023]
Find a Developer Job: 12 Best Job Sites For Developers