> Markdown version of [/jobs/ext/1293590-site-reliability-release-pipeline-engineer](https://www.wearedevelopers.com/jobs/ext/1293590-site-reliability-release-pipeline-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability/ Release Pipeline Engineer - **Company:** Gridiron IT Solutions LLC - **Location:** Arlington, VA, United States (Remote available) - **Experience:** Experienced - **Salary:** $190,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Build Automation, Continuous Delivery, DevOps, Identity and Access Management, Python (Programming Language), Octopus Deploy, Reliability Engineering, Software Deployment, Data Streaming, TypeScript, Scripting, Delivery Pipeline, Large Language Models, Grafana, Prompt Engineering, Amazon Virtual Private Cloud (VPC), Amazon Relational Database Service, Gitlab-ci, Kubernetes, Information Technology, Low Latency, Document Classification, Dynatrace, Artifactory - **Published:** July 16, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9033947/site-reliability-release-pipeline-engineer ## About the Role 3+ years of professional experience in SRE, DevOps, platform, or infrastructure engineering Experience operating Kubernetes (EKS) in production, including in DoD/IC classified environments Experience packaging and deploying applications with Helm (authoring and maintaining charts, not only consuming them) Experience with Flux (or an equivalent GitOps controller - e.g., Argo CD) driving continuous delivery of Helm releases Experience with AWS compute and managed services (e.g., EKS, RDS, S3, IAM/IRSA, EC2 Image Builder) Experience with infrastructure-as-code; TypeScript/CDK experience specifically, or demonstrated ability to work in a TypeScript IaC codebase Experience with production observability and distributed tracing (e.g., OpenTelemetry, Grafana/Tempo, or equivalent), used to diagnose failures from telemetry rather than guesswork Experience leading incident diagnosis and resolution, including identifying and confirming root cause before remediating Experience with GitLab CI/CD pipelines for build automation and deployment Proficiency in at least one scripting language (e.g., Python) for tooling and automation, Experience building STIG-compliant AMI pipelines using EC2 Image Builder with DoD security baseline validation Experience with Artifactory integration for AMI/container image distribution across AWS Organizations Experience with cross-domain solutions (AWS Diode, CDS) and SIPRNet/JWICS environments Experience with AWS Organizations, SCPs, OU design, and multi-account governance Experience with certificate lifecycle management (ACM, Private CA) and automated renewal workflows Experience deploying or operating ML/LLM-serving infrastructure (e.g., Amazon Bedrock endpoints, SageMaker, model-inference endpoints) and reasoning about token/latency/cost from traces Experience with an ECR- or registry-based GitOps model (charts/images mirrored to a registry, controller reconciling from it) Experience diagnosing failed or stuck Helm/Flux reconciliations (HelmRelease not converging, drift between desired and live state) Experience with Envoy/gateway and certificate/TLS management in classified environments Proficiency using AI coding assistants as a daily driver, with the judgment to validate their output before it reaches a cluster Bachelor's degree in Computer Science or a related field, or equivalent practical experience Program Qualification Run an engineering practice driven by evaluation, testing, and verification - changes are proven with tests, traces, and metrics before they are called done Operate agentic systems fluently (MCP tools, Bedrock/LLM calls, streaming, distributed traces) in support of GenAI IDP document classification and field extraction workflows Work within DCSA classified environments (IL5/IL6) following DoD security baselines and STIG compliance requirements Support the GenAI IDP solution: prompt engineering, model fine-tuning, evaluation logic for document section classification and extracted field validation Use AI coding assistants heavily as daily drivers, while remaining skeptical of their output and holding it to the same evidentiary bar as any other code Current, active Top Secret security clearance Current, active Security+ or equivalent certificate for privileged user access ## Description Owns the path from local development to deployed AWS cluster across DCSA's GovCloud (IL2/IL5) and classified (IL6/Secret) partitions, and the health of those clusters. These environments are currently in IATT and working towards scale and ATO. Deliberately a single role: at current team size, build/release and run/operate are the same person's problem. Release side: CDK stacks, Flux GitOps, Helm chart authoring, Envoy Gateway routes, EKS/IRSA, container builds, GitLab CI pipeline automation, and the local-to-AWS promotion path. Run side: the observability/tracing stack, cross-service tracing, runtime health, credential rotation, incident response, and the preflight/diagnose/QA loops. This role directly supports OY2 Workstream 1 (Image Builder Pipeline - STIG-compliant AMI automation, Artifactory integration, GitLab CI deployment), Workstream 6 (Import Account Standardization - Baseline Account Pipeline, tagging compliance, centralized VPC endpoints), and the GenAI IDP deployment pipeline (deploying the IDP Solution to Customer non-production and production environments via IaC). ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)