> Markdown version of [/jobs/ext/1316670-senior-platform-engineer](https://www.wearedevelopers.com/jobs/ext/1316670-senior-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Engineer - **Company:** Axiomatic_AI Inc. - **Location:** Cambridge, MA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Analysis, Microsoft Azure, Bash Shell, Cloud Computing, Computer Networks, Continuous Integration, Data Governance, Software Debugging, Linux, DevOps, Disaster Recovery, Github, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, PostgreSQL, Linux System Administration, Open Source Technology, Redis, Reliability Engineering, Prometheus, Azure Machine Learning, Runbook, Software Deployment, Software Vulnerability Management, Datadog, Data Logging, Scripting, Google Cloud, Istio, Delivery Pipeline, Grafana, Multi-Cloud, Reliability of Systems, Backend, Fastapi, AI Platforms, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Bare Metal, Linkerd (Service Mesh), Machine Learning Operations, Cloud Optimization, Api Gateway, Terraform, Software Version Control, Dynatrace, Jenkins, Microservices - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=675661d633a7ce19 ## About the Role * 7+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles * Deployment expert: Deep experience with CI/CD pipelines, release strategies, and production deployments at scale * Multi-cloud expertise: Hands-on experience with Azure and AWS required (GCP is a plus) * On-premise deployment experience: Linux system administration, bare-metal provisioning, networking * Terraform expert: Deep experience writing and maintaining infrastructure as code * Observability systems: Proven track record building monitoring, alerting, and metrics platforms * Security mindset: Experience implementing security controls and best practices. Security certification preferred (CISSP, CEH, AWS/Azure Security Specialty, or similar) * Data governance: Understanding of data privacy, residency requirements, and governance frameworks * Backend/scripting skills: Python (preferred) or Go for automation, tooling, and operational scripts * Kubernetes and container orchestration in production * Strong Linux/Unix administration and scripting (Bash, Python) * CI/CD platforms: GitHub Actions, GitLab CI, Jenkins, or similar * Version control and GitOps practices * Strong problem-solving and debugging skills * Fluent in English (Spanish is a plus) Nice-to-Have * Python proficiency for automation and internal tooling * Experience with cloud AI platforms (Vertex AI, Azure ML, AWS SageMaker) * Service mesh experience (Istio, Linkerd) or API gateways * Experience with GPU workloads and ML infrastructure * FinOps and cloud cost optimization * Compliance frameworks experience (SOC 2, ISO 27001, HIPAA, FedRAMP) * Database operations: PostgreSQL, Redis administration * Experience with FastAPI or similar frameworks for internal tools * Contributions to open-source infrastructure projects * Background in hardware or semiconductor industries ## Description As a Senior Platform Engineer at Axiomatic, you will own the reliability, deployment, and operational excellence of our AI platform. This role focuses primarily on infrastructure, CI/CD, and operations, with additional responsibilities for automation and tooling development. You will: * Lead deployment strategies and CI/CD pipelines across multiple environments * Architect and maintain multi-cloud infrastructure (Azure, AWS, GCP) and on-premise deployments * Own infrastructure as code using Terraform to automate provisioning and configuration * Build comprehensive observability systems: monitoring, metrics, logging, and alerting * Implement security controls, compliance frameworks, and data governance policies * Develop automation tools, APIs, and scripts (Python) to improve operational efficiency * Ensure system reliability, performance, and scalability * Drive incident response, postmortems, and continuous improvement * Troubleshoot infrastructure and application issues across multiple environments. Your mission Deployment & CI/CD * Design and implement deployment pipelines for multi-environment releases (dev, staging, production) * Own the full deployment lifecycle: build, test, release, and rollback strategies * Implement blue-green deployments, canary releases, and progressive rollouts * Build automated deployment tooling and workflows * Ensure zero-downtime deployments and rollback capabilities * Optimize build and deployment performance * Manage artifact repositories and container registries Infrastructure & Cloud Operations * Design and operate multi-cloud infrastructure across Azure, AWS, and GCP * Architect and deploy on-premise solutions for enterprise customers (Linux-based) * Manage Kubernetes clusters, container orchestration, and networking * Implement disaster recovery, backup strategies, and business continuity * Optimize cloud costs and resource utilization * Define and track SLIs, SLOs, and error budgets for critical services Infrastructure as Code * Write and maintain Terraform modules for infrastructure provisioning * Implement GitOps workflows for infrastructure changes * Automate infrastructure scaling, updates, and operations * Ensure reproducible and version-controlled infrastructure Observability & Monitoring * Design comprehensive monitoring, logging, and alerting (Prometheus, Grafana, Datadog, or similar) * Build dashboards for system health, performance, and business metrics * Implement distributed tracing for microservices * Conduct capacity planning and performance analysis * Drive reliability improvements through data-driven insights Security & Compliance * Implement security best practices: identity management, secrets management, network policies * Work towards or maintain security certifications (SOC 2, ISO 27001, or similar) * Conduct security audits and vulnerability remediation * Implement data governance policies for AI pipelines and user data * Ensure compliance with data privacy regulations (GDPR, CCPA) Automation & Tooling Development * Write automation scripts and tools in Python for operational tasks * Build internal tooling for deployments, monitoring, and incident response * Develop runbooks, automation, and self-healing systems * Create APIs for infrastructure operations when needed * Maintain high code quality and testing standards for tooling Reliability & Incident Management * Participate in on-call rotation and lead incident response * Conduct blameless postmortems and drive action items * Build and maintain incident response playbooks * Improve system resilience and failure modes Collaboration * Partner with engineering teams on deployment strategies and architecture * Work with security team on compliance and governance * Mentor engineers on operational best practices * Document systems, procedures, and runbooks ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [AI-Augmented DevOps with Platform Engineering](https://www.wearedevelopers.com/videos/1614-ai-augmented-devops-with-platform-engineering) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)