Platform Engineer (Site Reliability Engineering - SRE)

U.S. Navy
Winchester, VA, United States
3 months ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$78,400.0 - $123,200.0
Working hours
Regular working hours

Tech stack

Active Directory Application Programming Interfaces (APIs) Agile Methodology Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Continuous Delivery Continuous Integration DevOps Github Monitoring of Systems
+26 more
Python (Programming Language) Key Management OpenShift Operational Databases Role-Based Access Control Red Hat Enterprise Linux Reliability Engineering Azure Active Directory Ansible Prometheus Software Engineering Systems Integration VMware Infrastructure Scripting Google Cloud System Availability Grafana Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Bicep Terraform Splunk Azure Resource Manager Docker

Job description

You will leverage deep technical expertise of container and virtual infrastructure, GitOps practices, Continuous Integration, Continuous Deployment (CI/CD) automation, and infrastructure standardization and optimization. You have strong experience with on-premise and cloud infrastructure to include installation, configuration, administration, support and maintenance of IT software and hardware systems and internal databases for our enterprise clients. You will play a critical role in ensuring performance and resource availability of critical infrastructure supporting customers across Navy Federal Credit Union., * Manage and automate OCP (OpenShift Container Platform) and ARO (Azure Red Hat OpenShift) cluster upgrades, ensuring alignment with upstream releases and zero-downtime rolling updates

  • Configure tools like OADP (OpenShift API for Data Protection) to handle backup, restore, and failover across on-prem OCP and ARO regions
  • Monitor compute, storage, and networking capacity to prevent bottlenecks in hybrid Kubernetes environments
  • Define Service Level Indicators (SLIs) and Service Level Objectives (SLOs) specifically for cluster control planes and underlying cloud resources.
  • Use and tune frameworks (e.g., Prometheus, Grafana) to surface actionable alerts and track core signals like latency, traffic, errors, and saturation
  • Participate in on-call rotations to triage cluster-level issues (e.g., node evictions, Distributed ETC Directory (ETCD) latency, certificate expirations)
  • Manage continuous deployment pipelines using ArgoCD or Flux to ensure declarative, reproducible configurations for all cluster workloads
  • Abstract underlying Kubernetes complexities through self-service portals, standardized namespaces, and automated CI/CD guardrails
  • Consult with development teams to resolve deployment failures, routing, and pod-crash issues
  • Enforce Role-Based Access Control (RBAC) across environments, integrating with Active Directory or Azure Entra ID
  • Implement platform-level security guardrails using OpenShift policies and Open Policy Agent (OPA) to automatically detect configuration drifts or vulnerabilities
  • Integrate enterprise vaults or Azure Key Vault seamlessly with native OpenShift secrets
  • Treat the platform as software by managing ARO clusters and underlying Azure resources using Terraform or Bicep
  • Write automation scripts (Python, Go, or Ansible) to streamline common Day 2 configuration tasks like deploying custom operators, storage classes (e.g., ODF), and monitoring add-ons
  • Collaborate with development, security, and operations teams to improve platform reliability and developer experience using iterative Agile processes and practices
  • Create and maintain technical documentation, standards, and operational procedures using Docs as Code
  • Support continuous improvement initiatives focused on scalability, resiliency, and automation of both on-prem and cloud environments

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience
  • 5+ years of experience in platform engineering, DevOps, Site Reliability Engineering (SRE), or related roles
  • Hands-on experience administering and supporting OpenShift or Kubernetes environments
  • Experience implementing GitOps workflows using Argo, Flux
  • Strong experience building CI/CD pipelines with Tekton and GitHub Actions
  • Experience with container technologies such as Docker and Kubernetes
  • Proficiency with Infrastructure as Code tools such as Terraform, Ansible, or similar
  • Strong scripting skills in Bash, Python, or similar languages
  • Experience with cloud platforms such as Amazon Web Servicers (AWS), Azure, or Google Cloud Platform
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, ELK, or Splunk
  • Strong problem-solving, communication, and collaboration skills

Desired Qualifications

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Microsoft Certified: Azure Solutions Architect Expert
  • Red Hat Certified System Administrator (RHCSA)

About the company

Navy Federal provides much more than a job. We provide a meaningful career experience, including a culture that is energized, engaged and committed; and fierce appreciation for our teams, who are rewarded with highly competitive pay and generous benefits and perks.

Our approach to careers is simple yet powerful: Make our mission your passion.

  • FORTUNE 100 Best Companies to Work For 2026

  • Yello and WayUp Top 100 Internship Programs 2025

  • Computerworld Best Places to Work in IT 2026

  • Most Loved Workplace - America’s Top Most Loved Workplaces 2025

  • 2025 PEOPLE Companies That Care

  • Newsweek Most Trustworthy Companies in America 2026

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on localjobnetwork.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all