Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
Site Reliability Engineering is a pivotal role in the success of this project. Our SREs ensure that the platform software and platform automation is robust, reliable, and scalable.
As a Site Reliability Engineer, you will:
-
triage and remediate production incidents
-
provide engineering-level support for issues reported by users
-
work closely with development teams to improve the observability of the system
-
aggressively automate remediations for common problems
-
build tools to facilitate rapid triage and troubleshooting
-
build tools to enable continuous monitoring of production systems
-
measure service level objectives
-
define and improve service level objectives
Requirements
- Linux Administration
- AWS
- Demonstrated ability to write programs using Java based technologies/Scala.
- Experience in shell scripting using Python or Shell
- Must have experience managing cloud production distributed application stack in AWS/Azure/Google cloud
- Docker and Container Orchestration experience
- Hands on Experience using configuration management tools like Ansible, Chef or Puppet is a big plus
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on ibainfotech.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top-Paying Tech Jobs (with Salaries)
Dev Digest 137 - AI'm not sure about this
The Best Software Developer Blogs to Read
Is Software Engineering Over-Saturated?