> Markdown version of [/jobs/ext/654298-java-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/654298-java-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Java Site Reliability Engineer - **Company:** Arthur Grand Technologies - **Location:** McLean, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Application Performance Management, Confluence, JIRA, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Information Systems, Databases, Continuous Integration, Disaster Recovery, Distributed Systems, Github, Monitoring of Systems, Python (Programming Language), Octopus Deploy, Performance Tuning, Release Management, Reliability Engineering, Ansible, Prometheus, Shell Script, SQL Databases, Datadog, Data Logging, Google Cloud, Java Application Server, System Availability, Grafana, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Cloud Migration, Terraform, Splunk, Docker, Jenkins, Servicenow, Microservices - **Published:** June 26, 2026 - **Apply:** https://www.dice.com/job-detail/1079c3fb-50fe-4879-ae97-960eb0787bae ## About the Role · 16-20 years of experience in Site Reliability Engineering (SRE), Production Engineering, Platform Engineering, or Application Support. · Strong experience supporting large-scale enterprise production environments. Proven background in incident management, problem management, and operational support. · Experience working within banking, financial services, fintech, or other highly regulated industries. Hands-on experience supporting mission-critical applications with stringent availability and performance requirements. Required Skills · Java · Linux/Unix Administration · Kubernetes and Container Platforms · Docker · Cloud Platforms (AWS, Azure, or Google Cloud Platform) · CI/CD Tools (Jenkins, GitHub Actions, GitLab CI/CD, ArgoCD) · Infrastructure as Code (Terraform, Ansible) · Monitoring & Observability Tools (Splunk, Datadog, Grafana, Prometheus, Moog soft) · ServiceNow, JIRA, Confluence · Python, Bash, or Shell Scripting · SQL and Database Troubleshooting · Application Performance Monitoring (APM) · Production Release Management · Disaster Recovery and High Availability Architectures, · Bachelor''''s degree in Computer Science, Information Systems, Engineering, or a related technical discipline ## Description · Support and maintain highly available production platforms across cloud and distributed environments. Drive incident management, root cause analysis, problem management, and platform stability initiatives. · Monitor and maintain uptime of Java applications and microservices. · Proactively identify and resolve application performance bottlenecks. · Conduct root cause analysis (RCA) for application outages and incidents. · Implement resiliency patterns including circuit breakers, retries, and failover mechanisms. · Lead reliability engineering efforts focused on system availability, performance optimization, and operational excellence. Implement and enhance observability solutions including monitoring, logging, alerting, and incident response automation. · Collaborate with development, infrastructure, and cloud engineering teams to improve deployment reliability and operational efficiency. Support infrastructure modernization, cloud transformation, and platform automation initiatives. · Coordinate disaster recovery testing, resiliency validation, capacity planning, and production readiness reviews. Provide technical leadership and mentor offshore/onshore engineering teams. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [GitOps for the people](https://www.wearedevelopers.com/videos/815-gitops-for-the-people) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)