Platform Engineer

Allianz Group
Spain
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

Kubernetes Security Artificial Intelligence Airflow Application Performance Management Automation of Tests Microsoft Azure Computer Networks Continuous Integration Data Infrastructure DevOps Programming Tools Github
+31 more
Python (Programming Language) PostgreSQL Linux System Administration Networking Basics Performance Tuning Queueing Systems Role-Based Access Control Redis Cloud Services Prometheus Workflow Management Systems Data Logging Scripting Cloud Platform System Cloud Monitoring GitHub Copilot Autoscaling Istio Grafana Backend Kubernetes Bicep Linkerd (Service Mesh) Machine Learning Operations Celery Virtual Agents Terraform Dynatrace Azure Resource Manager Key Vault Databricks

Job description

Implement and operate Kubernetes infrastructure (AKS): cluster lifecycle, networking, resource management, auto-scaling, and multi-tenancy patterns. - Build and maintain CI/CD pipelines using GitHub Actions and ArgoCD for automated testing, container builds, and GitOps deployments. - Develop Infrastructure as Code (Terraform, Bicep) to provision and manage Azure resources with consistency and auditability. - Operate container registries (ACR), artifact management, and image security scanning workflows. - Implement and maintain observability infrastructure: Azure Monitor, Application Insights, Prometheus, Grafana-including dashboards, alerting, and distributed tracing. - Manage async processing infrastructure: Celery workers, Redis queues, and workflow orchestration patterns supporting AI agent execution. - Implement platform security controls: network policies, pod security standards, Key Vault integration, RBAC, and private endpoint configurations. - Support database

Requirements

infrastructure: PostgreSQL management, backup/recovery, connection pooling, and performance tuning. - Create self-service tooling and templates that enable development teams to deploy and operate services with minimal friction. - Diagnose and resolve infrastructure issues across clusters, pipelines, and cloud services; perform root-cause analysis and implement preventative improvements. - Collaborate with Platform Architects, Backend Engineers, and ML Engineers to translate architecture designs into reliable infrastructure. What You Bring - 5+ years professional experience in platform engineering, SRE, or DevOps roles; experience supporting AI/ML workloads is a strong plus. - Strong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and troubleshooting. - Solid Infrastructure as Code experience with Terraform, Bicep, or equivalent tools. - Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Virtual Networks, Private Endpoints, and Azure Policy. - Strong CI/CD experience: GitHub Actions (self-hosted runners, reusable workflows), ArgoCD, or similar GitOps tooling. - Proficiency in Python for automation, scripting, and tooling. - Experience with container security: image scanning, runtime security, network policies, and least-privilege patterns. - Experience with observability stack: Prometheus, Grafana, centralized logging, and alerting configuration. - Familiarity with async task processing: Celery, Redis, or equivalent message queue patterns. - Strong Linux systems administration and networking fundamentals. - Operational mindset with strong troubleshooting skills across infrastructure layers. Ways of Working - Comfortable in agile, iterative delivery environments with ownership and accountability. - Clear communicator and collaborator across global, cross-functional stakeholders. - Strong focus on reliability and automation: you measure success by system uptime and reduced manual toil. - Proactive learner with pragmatic adoption of AI-assisted developer tools (GitHub Copilot, Claude Code) to improve automation and delivery. Nice to Have - Experience supporting AI/ML infrastructure: GPU scheduling, model serving platforms, or ML pipeline orchestration. - Service mesh experience (Istio, Linkerd) for traffic management and security. - Experience with Databricks or similar data platform infrastructure. - Familiarity with workflow orchestration (Temporal, Airflow) for complex AI pipelines. - Experience with cost optimization: FinOps practices, reso

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on es.trabajo.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all