> Markdown version of [/jobs/ext/2306018-sre-ii-east-coast](https://www.wearedevelopers.com/jobs/ext/2306018-sre-ii-east-coast). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE II (East Coast) - **Company:** Insight Global - **Location:** Fairfax County, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Layers, Systems Engineering, Bash Shell, Cloud Computing, Cloud Engineering, Continuous Delivery, Linux, DevOps, Distributed Systems, Elasticsearch, Python (Programming Language), Network Troubleshooting, Transport Layer, Linux System Administration, Red Hat Enterprise Linux, Reliability Engineering, Cloud Services, Prometheus, Software Engineering, Load Balancing, System Availability, Grafana, Software Troubleshooting, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Bare Metal, Kibana, Terraform, Golang - **Published:** August 30, 2026 - **Apply:** https://dejobs.org/x/x/8B49B22FBA844C408D9C52A6C5C16491/job/ ## About the Role 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Systems Engineering, Infrastructure Engineering, or a related technical field. -Strong Linux administration experience, including RHEL and Amazon Linux. -Hands-on AWS experience supporting production workloads. -Experience supporting Kubernetes environments, including Amazon EKS. -Experience building, maintaining, and troubleshooting GitLab CI/CD pipelines. -Strong automation and scripting skills utilizing Bash and Python. -Working knowledge of Go. Hands-on experience with: -Prometheus -Grafana -Alertmanager -Strong understanding of TCP/IP networking fundamentals. -Experience troubleshooting Layer 4 and Layer 7 networking issues. -Experience supporting load balancing technologies and traffic routing. -Experience supporting 24x7 production environments and responding to critical incidents. -Strong troubleshooting, incident management, and operational support experience. -Ability to work overnight, weekend, holiday, and rotational shift schedules. -Experience supporting FedRAMP, , or other regulated cloud environments. -Experience supporting secure government or compliance-driven infrastructures. -Experience operating and troubleshooting bare metal environments. -Experience supporting air-gapped environments. -Experience with Infrastructure as Code tools such as Terraform. -Familiarity with distributed systems and cloud-native architectures. -Experience supporting mission-critical customer-facing applications and platforms. ## Description We are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Cloud Operations team. In this role, you will provide round-the-clock, eyes-on-glass monitoring, proactive incident response, operational maintenance, and continuous compliance support for our FedRAMP-authorized cloud platforms across AWS (Commercial and GovCloud) and Kubernetes (EKS) environments. As an SRE II, your primary responsibility is maintaining strict system availability and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 platform uptime. Responsibilities: -Monitor and support production cloud services in a 24x7x365 operations environment. -Perform incident response, service restoration, and production troubleshooting activities. -Execute documented operational procedures and runbooks during service-impacting events. -Escalate incidents to senior engineering and infrastructure teams when necessary. -Investigate infrastructure, application, networking, and platform alerts. -Support Kubernetes-based workloads running within AWS environments. -Utilize Prometheus, Grafana, and Alertmanager to monitor system health and performance. -Analyze and troubleshoot Layer 4 and Layer 7 networking issues. -Support load balancing, traffic routing, and service availability initiatives. -Develop and maintain operational tooling and automation using Bash, Python, and Go. -Participate in post-incident reviews and continuous improvement efforts. -Create and maintain technical documentation, monitoring standards, and operational runbooks. -Collaborate with platform, infrastructure, software engineering, and security teams to maintain service reliability and performance. We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)