Principal Platform Engineer
Bliss Verna
Us, France
3 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Amazon Web Services
Microsoft Azure
Cloud Computing
Monitoring of Systems
Prometheus
Scientific Computating
AI Infrastructure
Datadog
Pulumi
Grafana
Backend
+5 more
Kubernetes
Azure AKS
Machine Learning Operations
Virtual Agents
Terraform
Job description
- Architect, build, and operate Kubernetes infrastructure supporting thousands of long-running, stateful AI agent workloads.
- Design and implement custom Kubernetes operators and CRDs using tools such as Kubebuilder, Operator SDK, and controller-runtime.
- Own the broader infrastructure platform across cloud infrastructure, observability, networking, storage, and scaling strategy.
- Lead and mentor a small platform engineering team while remaining heavily hands-on technically.
- Partner closely with backend, ML, and research teams to design reliable infrastructure patterns for AI and scientific computing workloads.
- Establish production-grade infrastructure practices around monitoring, alerting, observability, and incident response.
- Drive infrastructure reliability, scalability, and operational excellence across the platform.
Requirements
This is a highly hands-on leadership role for someone who has deep Kubernetes expertise, enjoys operating with high ownership in startup environments, and wants to build infrastructure that directly enables breakthrough scientific research., * 10+ years of experience in infrastructure or platform engineering with deep hands-on Kubernetes expertise.
- Proven experience architecting and operating Kubernetes platforms for long-running, stateful workloads rather than only stateless services or short-lived inference jobs.
- Strong experience building custom Kubernetes operators and CRDs using Kubebuilder, Operator SDK, controller-runtime, or similar tooling.
- Recent startup or small-team experience with end-to-end ownership of infrastructure platforms.
- Experience leading or mentoring small engineering teams.
- Strong cloud infrastructure experience across AWS, GCP, or Azure, including managed Kubernetes services (EKS/GKE/AKS).
- Familiarity with infrastructure-as-code tooling such as Terraform or Pulumi.
- Experience building observability and monitoring systems using tools like Prometheus, Grafana, Datadog, or similar.
- Background supporting ML infrastructure, scientific computing, or data-intensive platforms is highly valued.
- Strong communication skills and high EQ, with the ability to work cross-functionally with technical and non-technical stakeholders.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on eu-recruit.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EF
Elizabeth Fuentes Leone, AWS Developer Advocate, GenAI
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
9 months ago
BR
Benjamin Ruschin
Navigating the AI Shift
11 months ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago