Senior Platform Engineer

Axiomatic AI
Madrid, Spain
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
7 years minimum
Working hours
Regular working hours
Languages
English, Spanish

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Microsoft Azure Bash Shell Cloud Computing Computer Networks Continuous Integration Data Governance Software Debugging Linux
+48 more
DevOps Disaster Recovery Github Hardware Design Monitoring of Systems Identity and Access Management Python (Programming Language) Key Management PostgreSQL Linux System Administration Open Source Technology Redis Reliability Engineering Prometheus Azure Machine Learning Runbook Software Deployment Software Vulnerability Management Datadog Data Logging Scripting Google Cloud Istio Delivery Pipeline Grafana Deep Learning Multi-Cloud Reliability of Systems Backend Fastapi Build Management AI Platforms Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Bare Metal Data Analytics Linkerd (Service Mesh) Machine Learning Operations Cloud Optimization Api Gateway Terraform Software Version Control Dynatrace Engineering Base Jenkins Microservices

Job description

About Us Axiomatic AI is building a new class of AI systems designed to reason with the rigor of the scientific method.By combining deep learning with formal logic and physics-based modeling, we create verifiable, interpretable AI systems that collaborate with and support human researchers in high-stakes scientific and engineering workflows.Our mission, 30x30, is to deliver a 30x improvement in the speed, accessibility, and cost of semiconductor and photonic hardware development by **.We aim to revolutionize hardware design and simulation in these industries and are building a team of highly motivated professionals to bring these innovations from research into commercial products.Position OverviewAs a Senior Platform Engineer at Axiomatic, you will own the reliability, deployment, and operational excellence of our AI platform.This role focuses primarily on infrastructure, CI/CD, and operations, with additional responsibilities for automation and tooling development.You WillLead deployment strategies and CI/CD pipelines across multiple environmentsArchitect and maintain multi-cloud infrastructure (Azure, AWS, GCP) and on-premise deploymentsOwn infrastructure as code using Terraform to automate provisioning and configurationBuild comprehensive observability systems: monitoring, metrics, logging, and alertingImplement security controls, compliance frameworks, and data governance policiesDevelop automation tools, APIs, and scripts (Python) to improve operational efficiencyEnsure system reliability, performance, and scalabilityDrive incident response, postmortems, and continuous improvementTroubleshoot infrastructure and application issues across multiple environmentsDeployment & CI/CDDesign and implement deployment pipelines for multi-environment releases (dev, staging, production)Own the full deployment lifecycle: build, test, release, and rollback strategiesImplement blue-green deployments, canary releases, and progressive rolloutsBuild automated deployment tooling and workflowsEnsure zero-downtime deployments and rollback capabilitiesOptimize build and deployment performanceManage artifact repositories and container registriesInfrastructure & Cloud OperationsDesign and operate multi-cloud infrastructure across Azure, AWS, and GCPArchitect and deploy on-premise solutions for enterprise customers (Linux-based)Manage Kubernetes clusters, container orchestration, and networkingImplement disaster recovery, backup strategies, and business continuityOptimize cloud costs and resource utilizationDefine and track SLIs, SLOs, and error budgets for critical servicesInfrastructure as CodeWrite and maintain Terraform modules for infrastructure provisioningImplement GitOps workflows for infrastructure changesAutomate infrastructure scaling, updates, and operationsEnsure reproducible and version-controlled infrastructureObservability & MonitoringDesign comprehensive monitoring, logging, and alerting (Prometheus, Grafana, Datadog, or similar)Build dashboards for system health, performance, and business metricsImplement distributed tracing for microservicesConduct capacity planning and performance analysisDrive reliability improvements through data-driven insightsSecurity & ComplianceImplement security best practices: identity management, secrets management, network policiesWork towards or maintain security certifications (SOC 2, ISO **, or similar)Conduct security audits and vulnerability remediationImplement data governance policies for AI pipelines and user dataEnsure compliance with data privacy regulations (GDPR, CCPA)Automation & Tooling DevelopmentWrite automation scripts and tools in Python for operational tasksBuild internal tooling for deployments, monitoring, and incident responseDevelop runbooks, automation, and self-healing systemsCreate APIs for infrastructure operations when neededMaintain high code quality and testing standards for toolingReliability & Incident ManagementParticipate in on-call rotation and lead incident responseConduct blameless postmortems and drive action itemsBuild and maintain incident response playbooksImprove system resilience and failure modesCollaborationPartner with engineering teams on deployment strategies and architectureWork with security team on compliance and governanceMentor engineers on operational best practicesDocument systems, procedures, and runbooksKey Requirements7+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Infrastructure Engineering rolesDeep experience with CI/CD pipelines, release strategies, and production deployments at scaleHands-on experience with Azure and AWS (GCP is a plus)Linux system administration, bare-metal provisioning, networking for on-premise deploymentsExpert proficiency writing and maintaining Terraform as codeProven track record building monitoring, alerting, and metrics platformsExperience implementing security controls and best practices; security certification preferred (CISSP, CEH, AWS/Azure Security Specialty, or similar)Understanding of data privacy, residency requirements, and governance frameworksBackend/scripting skills: Python (preferred) or Go for automation, tooling, and operational scriptsExperience with Kubernetes and container orchestration in productionStrong Linux/Unix administration and scripting (Bash, Python)Familiarity with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, or similar)Version control and GitOps practicesStrong problem-solving and debugging skillsFluent in English (Spanish is a plus)Nice-to-HavePython proficiency for automation and internal toolingExperience with cloud AI platforms (Vertex AI, Azure ML, AWS SageMaker)Service mesh experience (Istio, Linkerd) or API gatewaysExperience with GPU workloads and ML infrastructureFinOps and cloud cost optimizationCompliance frameworks experience (SOC 2, ISO ***, HIPAA, FedRAMP)Database operations: PostgreSQL, Redis administrationExperience with FastAPI or similar frameworks for internal toolsContributions to open-source infrastructure projectsBackground in hardware or semiconductor industriesWork model & location expectationsHybrid work model, open to remotePrimary location: Preferential timezone EUOn-site expectations: hybrid or remote, ~2 days per week in the office (with flexibility).Occasional travel to our Barcelona or Boston office may be required if remote.Why join us?At Axiomatic AI, you will be working on technology that drives innovation in AI for scientific and engineering applications in line with our 30x30 mission.This is your opportunity to contribute to the development of new AI architectures that can reason coherently and produce interpretable and verifiable solutions.Consequently, see those ideas commercialized into products that will shape the future of hardware and computing, while collaborating with a global team of engineers and AI specialists.We believe in pushing the boundaries of what is possible and continuously seek to redefine the intersection of AI, with focus on formal consistency.If you’re ready to take your expertise in artificial intelligence and physics to the next level, we want to hear from you!Worried about not meeting every qualification?Studies show that women and people of color are less likely to apply for jobs unless they meet every listed requirement.At Axiomatic-AI, we are dedicated to creating a diverse, inclusive, and authentic workplace.If this role excites you but your background doesn’t perfectly match every qualification, we still encourage you to apply.You could be the perfect fit for this position or another opportunity with us.#J-**-Ljbffr

Requirements

7+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles Deep experience with CI/CD pipelines, release strategies, and production deployments at scale Hands-on experience with Azure and AWS (GCP is a plus) Linux system administration, bare-metal provisioning, networking for on-premise deployments Expert proficiency writing and maintaining Terraform as code Proven track record building monitoring, alerting, and metrics platforms Experience implementing security controls and best practices; security certification preferred (CISSP, CEH, AWS/Azure Security Specialty, or similar) Understanding of data privacy, residency requirements, and governance frameworks Backend/scripting skills: Python (preferred) or Go for automation, tooling, and operational scripts Experience with Kubernetes and container orchestration in production Strong Linux/Unix administration and scripting (Bash, Python) Familiarity with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, or similar) Version control and GitOps practices Strong problem-solving and debugging skills Fluent in English (Spanish is a plus) Nice-to-Have Python proficiency for automation and internal tooling Experience with cloud AI platforms (Vertex AI, Azure ML, AWS SageMaker) Service mesh experience (Istio, Linkerd) or API gateways Experience with GPU workloads and ML infrastructure FinOps and cloud cost optimization Compliance frameworks experience (SOC 2, ISO *****, HIPAA, FedRAMP) Database operations: PostgreSQL, Redis administration Experience with FastAPI or similar frameworks for internal tools Contributions to open-source infrastructure projects Background in hardware or semiconductor industries Work model & location expectations Hybrid work model, open to remote

About the company

At Axiomatic AI, you will be working on technology that drives innovation in AI for scientific and engineering applications in line with our 30x30 mission. This is your opportunity to contribute to the development of new AI architectures that can reason coherently and produce interpretable and verifiable solutions. Consequently, see those ideas commercialized into products that will shape the future of hardware and computing, while collaborating with a global team of engineers and AI specialists. We believe in pushing the boundaries of what is possible and continuously seek to redefine the intersection of AI, with focus on formal consistency. If you’re ready to take your expertise in artificial intelligence and physics to the next level, we want to hear from you! Worried about not meeting every qualification? Studies show that women and people of color are less likely to apply for jobs unless they meet every listed requirement. At Axiomatic-AI, we are dedicated to creating a diverse, inclusive, and authentic workplace. If this role excites you but your background doesn’t perfectly match every qualification, we still encourage you to apply. You could be the perfect fit for this position or another opportunity with us.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all