> Markdown version of [/jobs/ext/2617428-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2617428-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** Atlanticus Services Corporation - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $100,000.0 - $115,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Application Performance Management, Bash Shell, Batch Processing, Cloud Computing, Databases, DevOps, Distributed Systems, Domain Name System (DNS), Github, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Identity and Access Management, Java Virtual Machine (JVM), Java Web Services, Python (Programming Language), Linux System Administration, Log Analysis, MySQL, Networking Basics, Octopus Deploy, Oracle Databases, Performance Tuning, Reliability Engineering, Software Tools, Prometheus, Software Deployment, Systems Integration, TCP/IP, Datadog, Scripting, Transport Layer Security, Application Enhancement Tool, Load Balancing, Autoscaling, Grafana, Reliability of Systems, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), Amazon Relational Database Service, Kubernetes, Route53, BIG-IP Access Policy Manager (APM), Cloudwatch, Terraform, Splunk, Docker, Jenkins, Microservices - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5a8a5fc6717625ae ## About the Role * 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role. * Strong experience supporting Java-based applications in production environments. * Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC. * Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments. * Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis. * Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar). * Experience supporting MySQL and Oracle databases from an application support perspective. * Proficiency in Python, Bash, or other scripting languages for automation. * Strong Linux system administration and troubleshooting skills. * Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure. * Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls. * Experience with incident management, problem management, and change management processes. * Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency. Preferred * Experience with Helm and GitOps deployment models. * Experience with Terraform or Infrastructure as Code. * Familiarity with Prometheus, Grafana, or OpenTelemetry. * Experience supporting microservices architectures. * Knowledge of JVM tuning and Java performance optimization. * Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler. ## Description We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability. The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services. The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency., * Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response. * Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools. * Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact. * Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents. * Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines. * Execute production deployments, application releases, hotfixes, and rollbacks following change management processes. * Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution. * Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics. * Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues. * Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks. * Build self-healing and automated operational processes to improve system reliability and reduce manual intervention. * Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues. * Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks. * Partner with Development teams to improve application reliability, resiliency, scalability, and performance. * Continuously improve operational processes, monitoring coverage, automation, and deployment practices. * Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)