Site Reliability Engineer SRE (contract)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
o Lead or participate in triaging and resolving alerts related to high availability applications and platforms o Serve was a key escalation point for complex production issues and coordinate recovery across teams o Develop and enhance monitoring dashboards mentions reporting and alert runbooks/playbooks o Analyze trends from historical incidents and help drive root cause analysis and long-term fixes o Build automation scripts and tools to improve response time and reduce manual intervention o Coordinate with application owners, Engineering teams and service providers to ensure timely and accurate resolution of issues o Participate in on-call support rotation as needed and contribute to building resilient operations model o Continuously improve onboarding documentation alert definitions and service catalogs to ensure support readiness
Requirements
In this contingent resource assignment, you may: Consult on or participate in moderately complex initiatives and deliverables within Systems Operations Engineering and contribute to large-scale planning related to Systems Operations Engineering deliverables. Review and analyze moderately complex Systems Operations Engineering challenges that require an in-depth evaluation of variable factors. Contribute to the resolution of moderately complex issues and consult with others to meet Systems Operations Engineering deliverables while leveraging solid understanding of the function, policies, procedures, and compliance requirements. Collaborate with client personnel in Systems Operations Engineering. Required Qualifications: Systems Engineering or Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work or consulting experience, training, military experience, education., * Applicants must be authorized to work for ANY employer in the U.S. This position is not eligible for visa sponsorship.
o Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education o Experience in handling incident response, troubleshooting live issues and restoring services quickly o Hands on experience with monitoring tools such as AppDynamics, CRIBL, Dynatrace, Splunk, Thousandeyes, Prometheus, Grafana etc o Strong experience working with ServiceNow or other ITSM platforms for Incident, Change and Problem management o Proven experience in writing runbooks playbooks and working across teams to streamline support processes o Familiarity with scripting languages such as PowerShell Python or bash for automation, o Prior experience in DevOps/SRE Environments or Platform teams supporting enterprise scale applications. Familiarity with container platforms and cloud services (Azure, AWS, GCP) o Experience with CI/CD pipelines deployment processes and observability best practices o Excellent communication and collaboration skills, with proactive approach to issue resolution Strong organizational and documentation skills to support knowledge-based creation and escalation paths
Benefits & conditions
Pulled from the full job description
- Health insurance
- Life insurance
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Find a Developer Job: 12 Best Job Sites For Developers