Principal Platform Engineer

JOB POINT
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Airflow BigQuery Cloud Computing Continuous Integration Data Distribution Service DevOps Elasticsearch Github Monitoring of Systems Identity and Access Management PostgreSQL Machine Learning
+17 more
Ansible Azure Machine Learning Google Cloud Istio Delivery Pipeline Grafana Model Validation Backend Containerization Kubernetes Machine Learning Operations Front End Software Development Vertica Api Gateway Terraform Virtual Private Clouds Jenkins

Job description

We’re looking for a Principal Platform Engineer to architect and lead the infrastructure strategy for our next-generation Production ML platform on Google Cloud. In this role, you will be the backbone of our high-performance machine learning workloads, ensuring our systems are elastic, secure, and resilient. You won’t just maintain the status quo; you’ll build the “paved road” for our engineers, automating everything from model deployment to complex networking perimeters. We are a high-trust, outcome-focused team that moves quickly to solve some of the most challenging problems in the ML space., * Infrastructure Management: Design, deploy, and maintain elastic scaling cloud infrastructure (GCP) and containerization tools like Kubernetes for high-performance ML workloads.

  • CI/CD Pipeline Development and maintenance: Build automated pipelines for training, testing, and deploying machine learning models using tools like Jenkins, GitHub Actions, or Airflow.
  • Model Monitoring & Maintenance: Implement observability tools to track model drift, accuracy, latency, and performance degradation in production.
  • Collaboration: Bridge the gap between data engineers, ML engineers, Backend and Frontend engineers to ensure smooth production operation.
  • ML Observability: Implement comprehensive monitoring for system health (latency/uptime) alongside ML-specific metrics, such as feature drift, prediction accuracy, and data distribution shifts, to ensure long-term model reliability. Non ML workload and production metrics monitoring.
  • Deploy tools that empower individual teams to monitor their workloads.
  • Participate in on-call rotation, help manage posture to ensure compliance with standards such as SOC.

Requirements

Do you have experience in Virtual Private Clouds?, * Senior Expertise: 8 - 10+ years in DevOps/Platform Engineering, with at least 2 years of experience specifically operating and maintaining production ML workloads.

  • GCP & K8s Mastery: Deep, hands-on experience with GCP (VPC-SC, IAM, Organization Policies) and GKE (Cluster topology, Helm, Kustomize, and in-cluster operators like ArgoCD).
  • Service Mesh Excellence: High proficiency with Istio (VirtualServices, mTLS, sidecar injection) and API Gateways (specifically Kong).
  • Infrastructure as Code: Expert-level Terraform skills, specifically using an Atlantis/GitOps workflow across a massive, multi-hundred-file estate.
  • Secrets & Identity: Experience managing enterprise-grade identity and secrets (Auth0, Dex, ESO, or SOPS).
  • Data/ML Tooling: Experience operating Airflow in production and an ML-serving stack (e.g., Triton, vLLM, MLflow).
  • Database Management: Comfortable managing Cloud SQL (PostgreSQL), BigQuery, and in-cluster datastores like Elasticsearch or ClickHouse.
  • At least an upper-intermediate level of spoken and written English.

It would be great if you also had:

  • ML Observability: Past experience with continuous monitoring of model accuracy and detecting data/concept drift.
  • Automation Savvy: Experience with Ansible for cluster bootstrap and recovery.
  • Advanced Certifications: Kubernetes (CKA/CKS) or GCP Professional Cloud Architect/Security Engineer certifications.
  • Modern Stack Exposure: Familiarity with Loki, Grafana, or managing ClickHouse at scale.

About the company

Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions. Our vision is to become THE go-to resource for every cyber protection need individuals may face - today and in the future., As part of Point Wild, you will:

Solve real customer problems. Point Wild’s point solutions allow consumers to address their immediate cyber protection needs. Our mandate is to continuously anticipate our customers’ evolving digital security needs to create best-in-class solutions aimed at keeping them safe.

See your impact. We are a scrappy, nimble organization where individual contributions are needed and valued. You will see your impact every day.

Accelerate your career. As we expand, you will have the opportunity to learn new technologies, products, and markets in a fast-paced, growth-oriented environment.

Most importantly, you’ll get to work with other talented people at a company where people matter. If you want to put your fingerprint on an organization and leapfrog your growth, this is the place for you.

In keeping with our beliefs and goals, no employee or applicant will face discrimination or harassment based on race, color, ancestry, national origin, religion, age, gender, marital domestic partner status, sexual orientation, gender identity, disability status, or veteran status. Above and beyond discrimination or harassment based on “protected categories,” Point Wild is committed to being an inclusive community where all feel welcome. Whether blatant or hidden, barriers to success have no place at Point Wild.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo ¡ LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch ¡ WWC Europe 2026

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum ¡ WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters ¡ WWC 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou ¡ Coffee With Developers

7:15 min

Installing Istio programmatically with bash scripts

Thomas SßdbrÜcker ¡ LIVE

Videos

See all

Related articles

See all