> Markdown version of [/jobs/ext/1323639-lead-platform-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1323639-lead-platform-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Platform Infrastructure Engineer - **Company:** UFS LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon S3, Application Release Automation, Configuration Management, Information Systems, Computer Networks, Continuous Integration, Linux, Github, Machine Learning, Networking Basics, Nginx, Public Key Infrastructure, Ansible, Prometheus, Software Deployment, Load Balancing, Grafana, CADdy, Gitlab-ci, Kubernetes, Information Technology, Bare Metal, Hashicorp, Terraform, Docker - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1d420510dc5e4565 ## About the Role To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required. * 8-12+ years building and operating production infrastructure, with 2+ years leading or setting direction for a platform or infrastructure function * Demonstrated experience shipping and operating software in customer-managed or on-premises environments - deploying and operating inside environments you do not control * Deep, hands-on Docker and Kubernetes in production (deployments, networking, storage, upgrades), plus strong Linux and networking fundamentals * Fluency with IaC and configuration management as a discipline - modules, state, idempotency, and reproducible images * A track record of making infrastructure reliable, secure, and automated rather than hand-tuned Core Technologies * Containers & orchestration: Docker, Kubernetes, K3s, Helm * IaC & config: Terraform, Packer, Ansible * CI/CD: Gitea Actions (or GitHub Actions / GitLab CI) * Secrets & certs: HashiCorp Vault, SOPS; PKI / mTLS (or equivalent) * Networking & ingress: Caddy or NGINX, MetalLB * Storage: S3 and MinIO (object storage) * Observability: Prometheus, Grafana, OpenTelemetry, Loki (or equivalent) * On-prem / bare-metal: strong familiarity required; cloud (AWS) experience useful for tooling and CI/CD pipelines Nice to Have * Prior experience deploying software into regulated or air-gapped customer environments (banking, healthcare, or government) * Experience with Kubernetes network policy, namespace isolation, and multi-environment rollout strategies * GPU scheduling and cost optimization for model inference Education and/or Experience * Bachelor's degree in computer science, information systems, or a related technical field, or equivalent hands-on experience * Experience in the financial services industry or a regulated technology environment strongly preferred Work Structure & Expectations * Full-time role combining ongoing platform operations with initiative-based build-out of deployment tooling and automation * Close collaboration with engineering, security, and client success teams; on-call rotation as deployments reach production Physical Demands The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. While performing the duties of this job, the employee is regularly required to sit and use hands to finger, handle, or touch objects, tools, or controls. The employee frequently is required to talk or hear. The employee is occasionally required to stand; walk; and stoop, kneel, crouch, or crawl. The employee must occasionally lift and/or move up to 10 pounds, usually waist high, up to 50 feet away. Specific vision abilities required by this job include close vision and the ability to adjust focus. ## Description * Design and own the Infrastructure-as-Code baseline that enables dedicated single-tenant or on-premises deployment with minimal per-customer effort * Stand up and operate Docker and Kubernetes (and K3s for small or in-bank installs), ingress, load balancing, and the CI/CD pipeline * Build and maintain the secrets and certificate story and the observability stack * Define environment promotion, release automation, and rollback procedures - making deployments boring and repeatable * Partner with Security on encryption, network isolation, and exam-ready hardening aligned to regulatory guidance * Right-size and operate GPU and inference infrastructure as the AI workload grows * Establish on-call, SLOs, and capacity planning practices ahead of beta scale-up * Document deployment architecture and operating procedures to support customer due diligence and audit readiness Core Competencies * Infrastructure-as-Code discipline - modules, state, idempotency, and reproducible images * Security-first design: encryption, network isolation, and least-privilege access aligned to NIST CSF 2.0 and NIST SP 800-53 principles * Operational ownership - on-call, SLOs, observability, and incident response * Cross-functional collaboration with security, data, and AI/ML engineering teams Key Performance Indicators (KPIs) * Deployment repeatability: new bank environments stood up within defined time and effort targets * Platform availability and SLO attainment across all production deployments * Security and hardening posture - zero critical findings in customer security reviews * CI/CD pipeline reliability and mean time to recovery for infrastructure incidents * GPU and compute cost per inference request as AI workloads scale ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Post-Quantum Cryptography: Preparing for Q-Day](https://www.wearedevelopers.com/videos/100179-post-quantum-cryptography-preparing-for-q-day) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)