AI/ML Infrastructure Engineer

Predii Inc.
San Lorenzo, CA, United States
3 days ago
Apply on www.disabledperson.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Apple Mac Systems Microsoft Azure Backup Devices Bash Shell Cloud Computing Continuous Integration Linux DevOps Disaster Recovery Domain Name System (DNS)
+28 more
Github Identity and Access Management Virtual Private Networks (VPN) Python (Programming Language) Network Security Windows Servers Role-Based Access Control Ansible Prometheus Runbook TCP/IP Data Logging Google Cloud Load Balancing Okta Grafana Multi-Cloud Firewalls (Computer Science) Git Gitlab-ci Kubernetes Azure AKS Machine Learning Operations Cloud Optimization Terraform Devsecops Docker Vulnerability Analysis

Job description

We need a AI/ML Infrastructure Engineer who wants more than tickets - someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You’ll touch:

  • Multi-cloud infra (Azure, GCP, AWS)
  • Kubernetes, CI/CD, automation-everything
  • Security, compliance, access - keeping the house locked
  • Monitoring & reliability - catching problems before customers do
  • Incident response & disaster recovery
  • Cloud cost optimization (yes, we care about the bill too)

Senior folks: expect to shape architecture and mentor the team, not just execute someone else’s roadmap., * Cloud & Platform - build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.

  • DevOps & CI/CD - ship pipelines that are reliable and repeatable, not held together with duct tape.
  • Reliability & Observability - build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.
  • Security & Compliance - RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.
  • Resilience & Ops - own backups, DR, capacity planning, cloud costs, and production support.
  • Keep Leveling Up - evaluate new tools across DevOps, infra, and DevSecOps; you’re not stuck with 2019’s stack.
  • IT Support - jump in on Windows Server / macOS support when needed.

Requirements

  • Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKS
  • Automation: Terraform, Ansible, Helm, Bash, Python
  • CI/CD: GitHub Actions, GitLab CI, Azure DevOps (or similar)
  • Observability: Grafana, Prometheus, ELK (or equivalent)
  • Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentation
  • Identity: RBAC, cloud IAM, Okta, multi-tenant auth
  • Systems: Linux, Windows Server, macOS
  • Security & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR, * 3-5+ years hands-on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production - not just labs.
  • Strong cloud + Kubernetes chops - Azure/AKS preferred; GCP/AWS/Docker is a big plus.
  • Solid networking and security fundamentals.
  • Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).
  • Real Git-based CI/CD experience - automated deploys, security scanning included.
  • Battle-tested on observability & prod ops - monitoring, incident response, RCA, runbooks, backup, DR.
  • Working knowledge of security/compliance across Linux, Windows Server, macOS.
  • A self-starter mindset - comfortable working independently across a distributed US-India team., * Auth/IAM experience, especially multi-tenant setups.
  • Background in automotive or data-heavy platforms.
  • Been through a SOC 2 Type 2 audit before.
  • Startup or small-team energy - you’ve worn more than one hat.
  • Cloud, Kubernetes, or Terraform certs.
  • Senior folks: mentoring or technical leadership experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.disabledperson.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:33 min

Introduction to security advocacy and automation testing

Chris Heilmann +2 · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

4:37 min

Architecting single sign-on flows across multiple application domains

Gift Egwuenu · World Congress 2023

Videos

See all

Related articles

See all