SRE Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will provide hands-on Level 1/2 production support for AWS-hosted applications, prioritizing rapid incident containment and service restoration. You will work closely with on-call rotations and cross-functional teams to diagnose root causes, escalate defects, and improve overall reliability. You’ll own monitoring and dashboards to maintain performance, and you will help shape proactive health checks and incident post-mortems. This role suits a reliability-minded engineer ready to thrive in a fast-paced, 24/7 environment. Compensation / Benefits * Provide Level 1 and Level 2 production incident support for AWS-hosted applications and infrastructure * Triage incidents to identify root causes and restore service within SLAs * Escalate defects with diagnostics and impact assessments to development teams * Participate in on-call rotations, major incident bridges, and post-incident reviews * Investigate application defects, configuration issues, and infrastructure anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements *
Requirements
restoration. anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
Software Engineer Salary London
Dev Digest 120 - Apple and peers