Senior Platform Engineer - Aws / Terraform / Linux

Capitole
Madrid, Spain
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Airflow Amazon Web Services Cloud Computing Cloud Engineering Software Debugging Linux DevOps Domain Name System (DNS) Identity and Access Management Subnetting Linux System Administration
+19 more
Networking Basics Routing Scrum Methodology Reliability Engineering Cloud Services Azure Machine Learning Workflow Management Systems Data Logging Load Balancing Cloud Platform System Large Language Models Grafana Generative AI Firewalls (Computer Science) AI Platforms Machine Learning Operations Terraform Data Pipelines Docker

Job description

About the role We are looking for aSeniorCloud Platform Engineerto join a technology team within a large international environment, working on the platform foundation that supports modern AI, data, and cloud-based systems.This role is highly focused on cloud infrastructure, platform engineering, AWS, Terraform, Linux troubleshooting, networking, automation, and production operations.The position is not primarily focused on building AI models or developing GenAI applications.Instead, you will work on the infrastructure and platform layer that enables technical teams to deploy, operate, monitor, and scale reliable systems in production.You will be expected to bring a strongPlatform / DevOps / Cloud Engineering mindset , with the ability to analyse real production issues, debug infrastructure problems, design AWS-based solutions, and make sound technical decisions in complex environments.Experience with AI/ML platforms, MLOps, SageMaker, MLflow, or LLM tooling is valuable, but the core of the role isAWS Cloud Platform Engineering .ResponsibilitiesDesign, build, and maintain cloud infrastructure solutions on AWSWork with Terraform / Infrastructure as Code to provision, manage, and standardize infrastructureAnalyse and troubleshoot production issues across cloud infrastructure, Linux systems, networking, storage, permissions, deployments, and platform servicesInvestigate infrastructure drift, unexpected production changes, Terraform state inconsistencies, and configuration mismatchesDesign AWS-based solutions, selecting the right services and explaining technical trade-offs around scalability, reliability, security, cost, and maintainabilitySupport and improve CI/CD pipelines for infrastructure, platform services, and cloud workloadsWork with Linux environments, including debugging issues related to disk usage, permissions, processes, logs, networking, and system performanceContribute to monitoring, observability, alerting, logging, and operational readiness of production systemsCollaborate with engineering, data, AI, and platform teams to ensure systems are reliable, automated, secure, and scalableApply DevOps, SRE, and platform engineering best practices to improve reliability, automation, and operational excellenceSupport cloud environments that may include AI/ML workloads, MLOps tooling, training/inference environments, or AI platform componentsQualificationsSolid experience in Platform Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or SREStrong hands?on experience with AWS in production environmentsStrong experience with Terraform and Infrastructure as CodeGood understanding of cloud infrastructure design, including networking, compute, storage, IAM/security, monitoring, and scalabilityStrong troubleshooting skills in Linux environmentsAbility to debug real infrastructure issues using command?line tools, logs, metrics, system resources, and cloud?native servicesExperience with CI/CD pipelines and automationUnderstanding of networking fundamentals, including VPCs, subnets, routing, DNS, load balancers, security groups, firewalls, and connectivity troubleshootingExperience with production operations, incident analysis, root cause investigation, and reliability improvementAbility to design technical solutions in AWS and explain the reasoning behind the selected services and architectureStrong ownership mindset and ability to work independently in complex technical environmentsGood communication skills and ability to explain technical decisions clearlyExperience with MLOps / AI Platform environments, SageMaker, MLflow, feature stores, model deployment, model serving, or training/inference platformsExperience with Docker and KubernetesFamiliarity with LLM tooling such as LangChain, Langfuse, LangSmith, or similarExperience with observability tools, monitoring platforms, logging, tracing, and alerting systemsExperience with cost optimisation in AWS environmentsExperience with data pipelines or workflow orchestration tools such as Airflow or PrefectKnowledge of security, governance, compliance, and best practices for cloud platformsExperience working in Agile / Scrum environmentsHybrid modelHybrid model: 2 days onsite per weekInformation Security NoticeThe employee will have access to confidential information related to Capitole and the assigned project.Compliance with internal security and information protection policies is mandatory.#J-*****-Ljbffr

Requirements

Solid experience in Platform Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or SRE Strong hands?on experience with AWS in production environments Strong experience with Terraform and Infrastructure as Code Good understanding of cloud infrastructure design, including networking, compute, storage, IAM/security, monitoring, and scalability Strong troubleshooting skills in Linux environments Ability to debug real infrastructure issues using command?line tools, logs, metrics, system resources, and cloud?native services Experience with CI/CD pipelines and automation Understanding of networking fundamentals, including VPCs, subnets, routing, DNS, load balancers, security groups, firewalls, and connectivity troubleshooting Experience with production operations, incident analysis, root cause investigation, and reliability improvement Ability to design technical solutions in AWS and explain the reasoning behind the selected services and architecture Strong ownership mindset and ability to work independently in complex technical environments Good communication skills and ability to explain technical decisions clearly Experience with MLOps / AI Platform environments, SageMaker, MLflow, feature stores, model deployment, model serving, or training/inference platforms Experience with Docker and Kubernetes Familiarity with LLM tooling such as LangChain, Langfuse, LangSmith, or similar Experience with observability tools, monitoring platforms, logging, tracing, and alerting systems Experience with cost optimisation in AWS environments Experience with data pipelines or workflow orchestration tools such as Airflow or Prefect Knowledge of security, governance, compliance, and best practices for cloud platforms Experience working in Agile / Scrum environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all