> Markdown version of [/jobs/ext/2484600-sme-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2484600-sme-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SME Platform Engineer - **Company:** General Dynamics Information Technology - **Location:** Arlington, VA, United States - **Salary:** $191,250.0 - $258,750.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Business Analytics Applications, Bash Shell, Big Data, Command-Line Interface, Cloud Computing, Computer Clusters, Computer Networks, Continuous Delivery, Continuous Integration, Distributed Computing Environment, Federal Information Processing Standards (FIPS), Identity and Access Management, Python (Programming Language), Linux System Administration, Machine Learning, Octopus Deploy, OpenID, Performance Tuning, Role-Based Access Control, Ansible, Prometheus, Software Engineering, Management of Software Versions, Data Logging, Network Routers, High Performance Computing, Delivery Pipeline, Grafana, Kubernetes Helm Charts, Infrastructure as Code (IaC), Git, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Data Analytics, Machine Learning Operations, Terraform, Block Storage, Docker, Programming Languages - **Published:** August 21, 2026 - **Apply:** https://dejobs.org/x/x/DCC0FBE305614B6198869478511A71CC/job/ ## About the Role CI/CD,Cloud Infrastructure,Cluster Administration,Kubernetes,Linux Server Administration, 15 + years of related experience, * Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar. * Strong Linux systems administration skills, including the ability to manage, diagnose, and troubleshoot infrastructure and platform services through the command line (CLI). * Experience implementing authentication and authorization solutions using technologies such asKeycloak, Open Policy Agent (OPA), OIDC, RBAC, or comparable identity and access management frameworks. * Hands-on experience with Git-based development and deployment workflows, including Helm charts, CI/CD pipelines, andGitOpspractices. * Experience with Argo CD or similar tools for declarative,GitOps-based continuous delivery. * Experience managing Kubernetes storage solutions, including distributed block storage and object storage; experience with Longhorn or comparable technologies preferred. * Experience hosting and administering data science or analytics platforms; experience with Posit Workbench, Posit Connect, HiveMetastore, or similar technologies. * Experience configuring and operating systemsin accordance withFIPS or comparable security and compliance requirements. * Demonstrated experienceidentifying, prioritizing, and remediating critical and high-severity system and application vulnerabilities. * Experience administering GPU-enabled compute environments supporting AI/ML, high-performance computing, or other compute-intensive workloads. * Experience supporting distributed AI/ML workloads using Ray or comparable distributed computing frameworks, including model training and fine-tuning use cases. * Experience deploying or supporting large language model inference and serving technologies; experience withvLLMand related routing capabilities preferred. * Experience pulling, managing, securing, and troubleshooting container images using private or public registries such as Harbor, Docker Hub, Container Yard, or equivalent container registry platforms. DESIRED SKILLS * Hands-on experience with Infrastructure as Code (IaC) and configuration management tools such as Terraform, Ansible, or comparable technologies. * Proficiencyin scripting or programming languages such as Python, Go, or Bash to support infrastructure automation, platform operations, and troubleshooting. * Advanced knowledge of Linux system administration, networking concepts, and protocols within complex orhighly availableinfrastructure environments. * Experience implementing andmaintainingmonitoring, logging, and observability solutions using tools such as Prometheus, Grafana, or comparable platforms. * Familiarity withMLOpspractices and the machine learning lifecycle, including model development, deployment, monitoring, versioning, and operational support. * Experience automating infrastructure provisioning, configuration, deployment, and operational workflows in secure or regulated environments. * Familiarity with performance tuning, capacity planning, and resource optimization for Kubernetes, GPU, or other compute-intensive environments., * Clearance:Current TS/SCI Clearance with current or willingness to obtain CI polygraph * Experience:15+years of related experience * Education:Bachelor's degree in Computer Science, Software Engineering, or a related field (or equivalent experience) * Role requirements:Work is onsite in Crystal City, VAwith optional CONUS travel * **Due to US Government Contract Requirements, only US Citizens are eligible for this role ** ## Description Iron EagleX is seeking a SME Platform Engineer to support our Engineering team in Crystal City, VA. This role will lead the design, implementation, and management of our secure, on-premises cloud infrastructure. In this role, you will be the driving force behind our advanced computing environments, ensuring the seamless orchestration of containerized applications and large-scale data science platforms. You will work at the intersection of infrastructure, security, and machine learning, managing robust compute clusters and providing foundational support for AI model training and deployment. The ideal candidate has deep expertise in Kubernetes ecosystem tools, GitOps methodologies, and strict compliance standards., As a SME Platform Engineer, your work will directly empower our data science and engineering teams to push the boundaries of machine learning and data analytics. By building and maintaining resilient GPU and Ray clusters, you will accelerate the fine-tuning and deployment of advanced models. Your commitment to security and compliance will ensure our critical systems remain protected against vulnerabilities, providing a safe, compliant, and highly performant foundation for the organization's most impactful technical initiatives. You will not just be managing infrastructure; you will be enabling innovation., * Infrastructure & Orchestration: Architect, deploy, and manage on-premises cloud infrastructure using RKE2 andmaintainstorage solutions like Longhorn and Objectstorage. * Platform Enablement: Host andmaintainrobust data science environments, including software such as POSIT Workbench/Connect and HiveMetastore. * AI/ML Infrastructure: Manage and scale robust GPU clusters, Ray Clusters for fine-tuning machine learning models, and VLLM Routers for efficient model inference. * CI/CD & Automation: Build,maintain, andoptimizeCI/CD pipelines using Git, Helm charts, andArgoCDfor reliable software delivery. * Security & Compliance: Ensure continuous FIPS compliance across the environment. Actively manage and mitigate critical and high-level vulnerabilities. * Identity & Access: Implement andmaintainrobust authentication and authorization mechanisms usingKeycloakand Open Policy Agent (OPA). * System Administration: Pull and manage container images from secure registries such as Harbor,DockerHub, orContaineryard. Manage all core capabilities and troubleshoot issues effectively via the command-line console. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)