Sr. SRE (Reston, VA)

Insight Global
Reston, VA, United States
1 day ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Ubuntu (Operating System) Cloud Computing Linux DevOps Federal Information Processing Standards (FIPS) Identity and Access Management Python (Programming Language) Linux System Administration Red Hat Enterprise Linux
+16 more
Reliability Engineering Cloud Services Prometheus Runbook Shell Script Software Vulnerability Management Scripting System Availability Delivery Pipeline Grafana Amazon Virtual Private Cloud (VPC) Gitlab-ci Kubernetes Route53 Terraform Data Pipelines

Job description

Insight Global is seeking a Senior Site Reliability Engineer (Senior SRE) for a top cloud technology and federal compliance-focused client. This individual will serve as a senior technical leader responsible for designing, scaling, and maintaining highly available cloud infrastructure across AWS Commercial and GovCloud environments. The ideal candidate will have deep expertise in Kubernetes, Amazon EKS, Terraform, and FedRAMP compliance, while playing a critical role in audit readiness, platform reliability, incident response, and automation initiatives. This is an opportunity to work on mission-critical cloud platforms, influence infrastructure architecture, drive security and compliance strategy, and mentor engineering teams in a highly visible and impactful role.

Day to day:

Design and implement highly available AWS Commercial and AWS GovCloud infrastructure solutions

Architect and manage multi-region Amazon EKS/Kubernetes environments using Terraform

Lead technical preparation and participation for FedRAMP 3PAO audits and compliance reviews

Oversee vulnerability management, CVE remediation, and compliance initiatives

Participate in a rotational 24/7 on-call schedule and lead response to critical production incidents

Perform root cause analysis and drive post-mortem improvements

Build and maintain observability platforms using Prometheus, Alertmanager, and Grafana

Define SLIs, SLOs, and error budgets across cloud services

Design and optimize GitLab CI/CD deployment pipelines and GitOps workflows

Manage Kubernetes application deployments using Helm

Create architectural documentation, SOPs, runbooks, and audit artifacts

Automate operational processes using Python, Go, or shell scripting

Mentor junior and mid-level SRE engineers and provide technical leadership

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

5-8+ years of experience in Site Reliability Engineering, Cloud DevOps, or Cloud Infrastructure Architecture

-Strong AWS expertise including VPC, IAM, EC2, S3, KMS, and Route53

-Experience with AWS GovCloud environments

-Hands-on FedRAMP High/Moderate authorization experience

-Experience leading 3PAO audits and Continuous Monitoring (ConMon) activities

-Strong knowledge of NIST SP 800-53 controls and FedRAMP compliance requirements

-Deep expertise with Amazon EKS and Kubernetes administration

-Terraform Infrastructure as Code experience

-Advanced Linux administration experience (RHEL, Amazon Linux, Ubuntu)

-Prometheus and Grafana monitoring/observability expertise

-GitLab CI/CD pipeline development experience

-Experience supporting production environments in a 24/7 on-call capacity

-Strong technical documentation and incident management skills -Python development or automation scripting experience

-Go programming experience

CISSP certification

-Certified Kubernetes Administrator (CKA) certification

-AWS Certified Solutions Architect - Professional certification

-Experience with DISA STIG hardening standards

-Experience with FIPS 140-2/3 requirements and vulnerability remediation programs

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all