SRE II (East Coast)

Insight Global
Fairfax County, VA, United States
1 day ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

Amazon Web Services Application Layers Systems Engineering Bash Shell Cloud Computing Cloud Engineering Continuous Delivery Linux DevOps Distributed Systems Elasticsearch Python (Programming Language)
+20 more
Network Troubleshooting Transport Layer Linux System Administration Red Hat Enterprise Linux Reliability Engineering Cloud Services Prometheus Software Engineering Load Balancing System Availability Grafana Software Troubleshooting Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Bare Metal Kibana Terraform Golang

Job description

We are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our

24/7 Cloud Operations team. In this role, you will provide round-the-clock, eyes-on-glass

monitoring, proactive incident response, operational maintenance, and continuous compliance support for our FedRAMP-authorized cloud platforms across AWS (Commercial and GovCloud) and Kubernetes (EKS) environments.

As an SRE II, your primary responsibility is maintaining strict system availability and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 platform uptime.

Responsibilities:

-Monitor and support production cloud services in a 24x7x365 operations environment.

-Perform incident response, service restoration, and production troubleshooting activities.

-Execute documented operational procedures and runbooks during service-impacting events.

-Escalate incidents to senior engineering and infrastructure teams when necessary.

-Investigate infrastructure, application, networking, and platform alerts.

-Support Kubernetes-based workloads running within AWS environments.

-Utilize Prometheus, Grafana, and Alertmanager to monitor system health and performance.

-Analyze and troubleshoot Layer 4 and Layer 7 networking issues.

-Support load balancing, traffic routing, and service availability initiatives.

-Develop and maintain operational tooling and automation using Bash, Python, and Go.

-Participate in post-incident reviews and continuous improvement efforts.

-Create and maintain technical documentation, monitoring standards, and operational runbooks.

-Collaborate with platform, infrastructure, software engineering, and security teams to maintain service reliability and performance.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

5+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Systems Engineering, Infrastructure Engineering, or a related technical field.

-Strong Linux administration experience, including RHEL and Amazon Linux.

-Hands-on AWS experience supporting production workloads.

-Experience supporting Kubernetes environments, including Amazon EKS.

-Experience building, maintaining, and troubleshooting GitLab CI/CD pipelines.

-Strong automation and scripting skills utilizing Bash and Python.

-Working knowledge of Go.

Hands-on experience with:

-Prometheus

-Grafana

-Alertmanager

-Strong understanding of TCP/IP networking fundamentals.

-Experience troubleshooting Layer 4 and Layer 7 networking issues.

-Experience supporting load balancing technologies and traffic routing.

-Experience supporting 24x7 production environments and responding to critical incidents.

-Strong troubleshooting, incident management, and operational support experience.

-Ability to work overnight, weekend, holiday, and rotational shift schedules. -Experience supporting FedRAMP,

, or other regulated cloud environments.

-Experience supporting secure government or compliance-driven infrastructures.

-Experience operating and troubleshooting bare metal environments.

-Experience supporting air-gapped environments.

-Experience with Infrastructure as Code tools such as Terraform.

-Familiarity with distributed systems and cloud-native architectures.

-Experience supporting mission-critical customer-facing applications and platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · World Congress 2022

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all