Lead Platform Infrastructure Engineer

UFS LLC
United States
29 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Application Release Automation Configuration Management Information Systems Computer Networks Continuous Integration Linux Github Machine Learning Networking Basics
+15 more
Nginx Public Key Infrastructure Ansible Prometheus Software Deployment Load Balancing Grafana CADdy Gitlab-ci Kubernetes Information Technology Bare Metal Hashicorp Terraform Docker

Job description

  • Design and own the Infrastructure-as-Code baseline that enables dedicated single-tenant or on-premises deployment with minimal per-customer effort
  • Stand up and operate Docker and Kubernetes (and K3s for small or in-bank installs), ingress, load balancing, and the CI/CD pipeline
  • Build and maintain the secrets and certificate story and the observability stack
  • Define environment promotion, release automation, and rollback procedures - making deployments boring and repeatable
  • Partner with Security on encryption, network isolation, and exam-ready hardening aligned to regulatory guidance
  • Right-size and operate GPU and inference infrastructure as the AI workload grows
  • Establish on-call, SLOs, and capacity planning practices ahead of beta scale-up
  • Document deployment architecture and operating procedures to support customer due diligence and audit readiness

Core Competencies

  • Infrastructure-as-Code discipline - modules, state, idempotency, and reproducible images
  • Security-first design: encryption, network isolation, and least-privilege access aligned to NIST CSF 2.0 and NIST SP 800-53 principles
  • Operational ownership - on-call, SLOs, observability, and incident response
  • Cross-functional collaboration with security, data, and AI/ML engineering teams

Key Performance Indicators (KPIs)

  • Deployment repeatability: new bank environments stood up within defined time and effort targets
  • Platform availability and SLO attainment across all production deployments
  • Security and hardening posture - zero critical findings in customer security reviews
  • CI/CD pipeline reliability and mean time to recovery for infrastructure incidents
  • GPU and compute cost per inference request as AI workloads scale

Requirements

To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required.

  • 8-12+ years building and operating production infrastructure, with 2+ years leading or setting direction for a platform or infrastructure function
  • Demonstrated experience shipping and operating software in customer-managed or on-premises environments - deploying and operating inside environments you do not control
  • Deep, hands-on Docker and Kubernetes in production (deployments, networking, storage, upgrades), plus strong Linux and networking fundamentals
  • Fluency with IaC and configuration management as a discipline - modules, state, idempotency, and reproducible images
  • A track record of making infrastructure reliable, secure, and automated rather than hand-tuned

Core Technologies

  • Containers & orchestration: Docker, Kubernetes, K3s, Helm
  • IaC & config: Terraform, Packer, Ansible
  • CI/CD: Gitea Actions (or GitHub Actions / GitLab CI)
  • Secrets & certs: HashiCorp Vault, SOPS; PKI / mTLS (or equivalent)
  • Networking & ingress: Caddy or NGINX, MetalLB
  • Storage: S3 and MinIO (object storage)
  • Observability: Prometheus, Grafana, OpenTelemetry, Loki (or equivalent)
  • On-prem / bare-metal: strong familiarity required; cloud (AWS) experience useful for tooling and CI/CD pipelines

Nice to Have

  • Prior experience deploying software into regulated or air-gapped customer environments (banking, healthcare, or government)
  • Experience with Kubernetes network policy, namespace isolation, and multi-environment rollout strategies
  • GPU scheduling and cost optimization for model inference

Education and/or Experience

  • Bachelor’s degree in computer science, information systems, or a related technical field, or equivalent hands-on experience
  • Experience in the financial services industry or a regulated technology environment strongly preferred

Work Structure & Expectations

  • Full-time role combining ongoing platform operations with initiative-based build-out of deployment tooling and automation
  • Close collaboration with engineering, security, and client success teams; on-call rotation as deployments reach production

Physical Demands

The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

While performing the duties of this job, the employee is regularly required to sit and use hands to finger, handle, or touch objects, tools, or controls. The employee frequently is required to talk or hear. The employee is occasionally required to stand; walk; and stoop, kneel, crouch, or crawl. The employee must occasionally lift and/or move up to 10 pounds, usually waist high, up to 50 feet away. Specific vision abilities required by this job include close vision and the ability to adjust focus.

About the company

Navanta is the trusted technology and services partner for community financial institutions, unifying critical systems, security, cloud infrastructure, and support into one seamless, purpose built experience. With more than 35 years of banking expertise - from Managed IT to Core Banking, CRM, and Advisory Services - Navanta helps institutions simplify complexity, reduce risk, and strengthen daily operations. Navanta empowers community bankers and their people to thrive together. Go Bankers, Go.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:33 min

Debugging application telemetry with unified network flow logs

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all