> Markdown version of [/jobs/ext/1969124-sre-engineer](https://www.wearedevelopers.com/jobs/ext/1969124-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Engineer - **Company:** LTM Inc - **Location:** Atlanta, GA, United States - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Application Layers, Databases, Monitoring of Systems, Delivery Pipeline, Grafana, Software Troubleshooting, Amazon Virtual Private Cloud (VPC), Cloudwatch, Dynatrace - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/sre-engineer-atlanta-ga-usa-58831752 ## About the Role restoration. anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements * ## Description Experteer Overview In this role you will provide hands-on Level 1/2 production support for AWS-hosted applications, prioritizing rapid incident containment and service restoration. You will work closely with on-call rotations and cross-functional teams to diagnose root causes, escalate defects, and improve overall reliability. You'll own monitoring and dashboards to maintain performance, and you will help shape proactive health checks and incident post-mortems. This role suits a reliability-minded engineer ready to thrive in a fast-paced, 24/7 environment. Compensation / Benefits * Provide Level 1 and Level 2 production incident support for AWS-hosted applications and infrastructure * Triage incidents to identify root causes and restore service within SLAs * Escalate defects with diagnostics and impact assessments to development teams * Participate in on-call rotations, major incident bridges, and post-incident reviews * Investigate application defects, configuration issues, and infrastructure anomalies reported by monitoring tools or users * Perform regular health checks across applications, infrastructure, and AWS services * Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes * Maintain and improve monitoring and observability dashboards Tasks * Strong experience supporting production systems on AWS (EC2, VPC, ALB/NLB, RDS, Lambda, EKS) * Hands-on incident management and 24/7 production support experience * Proficiency with monitoring/observability tools (CloudWatch, Dynatrace, Quantum Metric) * Experience building and maintaining monitoring dashboards * Strong troubleshooting across infrastructure, networking, and application layers * Working knowledge of CI/CD pipelines and AWS deployment processes * Experience with databases and Unix/Linux environments Key requirements * ## Related Videos - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)