Principal Platform Engineer

Bliss Verna
Us, France
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Monitoring of Systems Prometheus Scientific Computating AI Infrastructure Datadog Pulumi Grafana Backend
+5 more
Kubernetes Azure AKS Machine Learning Operations Virtual Agents Terraform

Job description

  • Architect, build, and operate Kubernetes infrastructure supporting thousands of long-running, stateful AI agent workloads.
  • Design and implement custom Kubernetes operators and CRDs using tools such as Kubebuilder, Operator SDK, and controller-runtime.
  • Own the broader infrastructure platform across cloud infrastructure, observability, networking, storage, and scaling strategy.
  • Lead and mentor a small platform engineering team while remaining heavily hands-on technically.
  • Partner closely with backend, ML, and research teams to design reliable infrastructure patterns for AI and scientific computing workloads.
  • Establish production-grade infrastructure practices around monitoring, alerting, observability, and incident response.
  • Drive infrastructure reliability, scalability, and operational excellence across the platform.

Requirements

This is a highly hands-on leadership role for someone who has deep Kubernetes expertise, enjoys operating with high ownership in startup environments, and wants to build infrastructure that directly enables breakthrough scientific research., * 10+ years of experience in infrastructure or platform engineering with deep hands-on Kubernetes expertise.

  • Proven experience architecting and operating Kubernetes platforms for long-running, stateful workloads rather than only stateless services or short-lived inference jobs.
  • Strong experience building custom Kubernetes operators and CRDs using Kubebuilder, Operator SDK, controller-runtime, or similar tooling.
  • Recent startup or small-team experience with end-to-end ownership of infrastructure platforms.
  • Experience leading or mentoring small engineering teams.
  • Strong cloud infrastructure experience across AWS, GCP, or Azure, including managed Kubernetes services (EKS/GKE/AKS).
  • Familiarity with infrastructure-as-code tooling such as Terraform or Pulumi.
  • Experience building observability and monitoring systems using tools like Prometheus, Grafana, Datadog, or similar.
  • Background supporting ML infrastructure, scientific computing, or data-intensive platforms is highly valued.
  • Strong communication skills and high EQ, with the ability to work cross-functionally with technical and non-technical stakeholders.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eu-recruit.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all