> Markdown version of [/jobs/ext/1344501-senior-ai-platform-engineer](https://www.wearedevelopers.com/jobs/ext/1344501-senior-ai-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Platform Engineer - **Company:** eMFusion Global - **Location:** Berlin, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Audit Trail, Software as a Service, Encodings, Computer Networks, Extract Transform Load (ETL), Identity and Access Management, Key Management, Routing, Nginx, OAuth, OpenID, Role-Based Access Control, Prometheus, Runbook, Data Streaming, Tripwire, Management of Software Versions, WebSocket, AI Infrastructure, Datadog, Policy as Code, Data Ingestion, Istio, Large Language Models, Grafana, Database Optimization, Caching, Amazon Virtual Private Cloud (VPC), AI Platforms, Kubernetes, Linkerd (Service Mesh), Machine Learning Operations, Terraform, Stream Processing, New Relic (SaaS), Software Version Control, Data Pipelines, Static Application Security Testing, Vulnerability Analysis, Dynamic Application Security Testing - **Published:** July 19, 2026 - **Apply:** https://www.adzuna.de/details/5665735872 ## About the Role * Proven patterns for tenant isolation (DB-per-tenant, schema-per-tenant, row-level security), tenant-aware caching, noisy-neighbor protection * OIDC/OAuth2, tenant-aware RBAC/ABAC, SCIM provisioning, and audit logging for B2B SaaS, * You have shipped and operated customer-facing SaaS products at scale with real users * You have owned end-to-end ML/AI infrastructure - from data ingestion through to production monitoring * You enable engineers and data scientists to move faster through self-service platforms and automated workflows * You have a track record of designing systems that scale globally across regions and traffic patterns * You are comfortable with incident response, on-call rotations, and stabilising critical production systems * You think with a product mindset - customer value, reliability, and speed-to-market over technology for its own sake * You have a strong bias for automation and eliminating manual operational toil * Excellent communication skills - async collaboration, documentation, and explaining technical decisions to non-technical audiences ## Description * Design and evolve a multi-tenant SaaS architecture with tenant isolation, per-tenant controls, and enterprise security * Build automated tenant provisioning, safe rollouts (canary/feature flags), and noisy-neighbor protection * Operationalise LLMs end-to-end - fine-tuning, evaluation, high-performance serving, monitoring, and embeddings workflows * Drive MLOps foundations: automated training pipelines, experiment tracking, and scalable model deployment * Manage Kubernetes clusters, GPU-heavy workloads, and autoscaling on AWS * Build unified CI/CD pipelines shipping ML and application code seamlessly * Implement comprehensive observability: logs, metrics, traces, model/data drift detection * Embed enterprise security and compliance - IAM, RBAC, VPC design, secrets management, encryption - at every layer * Design well-architected ETL/ELT pipelines, streaming systems, feature store integration, and workflow orchestration, * Deep Kubernetes: cluster ops, HPA/VPA, node pools, GPU scheduling, Karpenter, PDBs, network policies, multi-AZ design * Service mesh (Istio/Linkerd), ingress patterns (ALB/Nginx), secure egress, mTLS * Infrastructure as Code beyond basics: Terraform modules, Terragrunt, policy-as-code (OPA/Conftest), secrets automation * GitOps (ArgoCD/Flux), progressive delivery (Argo Rollouts/Flagger), feature flags, canary and blue/green deployments MLOps & Model Lifecycle * Model lifecycle tooling: MLflow/W&B, model registry, experiment tracking, reproducible training, dataset versioning (DVC/lakeFS) * Pipeline orchestration: Airflow, Prefect, or Dagster + artifact stores * Model serving: KServe, Seldon, BentoML, or Ray Serve - online, async/batch inference, autoscaling, rollback LLMOps * Prompt and version management, offline + online evaluation harnesses, RAG evaluation (retrieval metrics, groundedness), guardrails, red-teaming basics * Streaming inference (SSE/WebSockets), caching, routing, fallback models * Vector DB experience: pgvector, Pinecone, Weaviate, or Milvus - embedding lifecycle, backfills, re-embedding, indexing strategies Observability & Security * OpenTelemetry, tracing, SLOs - Prometheus/Grafana, Loki/ELK, Datadog/New Relic * Incident management: postmortems, runbooks, error budgets * GDPR, encryption at rest/in transit, secrets management (AWS Secrets Manager/Vault), KMS, key rotation * SOC 2 / ISO 27001 familiarity, vulnerability scanning (Trivy/Grype), SBOMs, SAST/DAST ## Related Videos - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Post-Quantum Cryptography: Preparing for Q-Day](https://www.wearedevelopers.com/videos/100179-post-quantum-cryptography-preparing-for-q-day) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) - [AI-Augmented DevOps with Platform Engineering](https://www.wearedevelopers.com/videos/1614-ai-augmented-devops-with-platform-engineering) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)