SRE Engineer

LTM Inc
Atlanta, GA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Application Layers Databases Monitoring of Systems Delivery Pipeline Grafana Software Troubleshooting Amazon Virtual Private Cloud (VPC) Cloudwatch Dynatrace

Job description

Experteer Overview In this role you will provide hands-on Level 1/2 production support for AWS-hosted applications, prioritizing rapid incident containment and service restoration. You will work closely with on-call rotations and cross-functional teams to diagnose root causes, escalate defects, and improve overall reliability. You’ll own monitoring and dashboards to maintain performance, and you will help shape proactive health checks and incident post-mortems. This role suits a reliability-minded engineer ready to thrive in a fast-paced, 24/7 environment. Compensation / Benefits * Provide Level 1 and Level 2 production incident support for AWS-hosted applications and infrastructure * Triage incidents to identify root causes and restore service within SLAs * Escalate defects with diagnostics and impact assessments to development teams * Participate in on-call rotations, major incident bridges, and post-incident reviews * Investigate application defects, configuration issues, and infrastructure anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements *

Requirements

restoration. anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all