> Markdown version of [/jobs/ext/2416425-principal-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2416425-principal-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Platform Engineer - **Company:** ELLKAY, LLC. - **Location:** United States (Remote available) - **Salary:** $160,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Configuration Management, Computer Programming, Software Debugging, DevOps, Identity and Access Management, Python (Programming Language), Network Segmentation, PCI Data Security Standards, Ansible, Prometheus, Management of Software Versions, Datadog, Policy as Code, Data Logging, Delivery Pipeline, Grafana, Mttr, Multi-Cloud, Kubernetes, Deployment Automation, Hardware Infrastructure, Puppet, Terraform, Splunk, Docker - **Published:** August 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d902ff56ea6bf913 ## About the Role * 15+ years in platform engineering, infrastructure engineering, DevOps, or SRE roles, with demonstrated staff-level scope and impact * Deep, production-grade experience with both AWS and Azure, plus experience managing on-premises infrastructure in a hybrid model * Expert-level Terraform experience - module design, state management, workspace/environment strategy at scale * Strong experience building CI/CD pipelines for Kubernetes and Docker-based workloads * Hands-on experience with configuration management tooling (Ansible, Chef, Puppet, or similar) * Proven track record implementing observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry, ELK/Splunk) and defining meaningful SLIs/SLOs * Demonstrated experience leading production incident response and driving reliability improvements * Experience designing cloud cost governance and security/compliance frameworks in a multi-cloud or hybrid environment * Strong scripting/programming ability (Python, Go, or Bash) for automation and tooling * Excellent cross-functional communication - able to work directly with SRE, security, and engineering leadership Preferred * Experience with policy-as-code frameworks (OPA/Gatekeeper, Sentinel) * Relevant certifications (AWS, Azure, CKA/CKAD) * Experience operating infrastructure under regulatory or compliance requirements (SOC 2, HIPAA, PCI-DSS, ISO 27001) * Prior experience in an on-call rotation for critical production systems * History of mentoring engineers or leading infrastructure initiatives across multiple teams ## Description We're looking for a Principal Platform Engineer to own the design, build, and operational health of our infrastructure across AWS, Azure, and on-premises environments. This is a hands-on, high-ownership role for someone who thinks in systems, codifies everything, and is equally comfortable writing Terraform modules, debugging a production incident at 2 AM, and advising leadership on cost and security posture., Infrastructure as Code * Design, build, and maintain reusable Terraform modules to provision and manage infrastructure across AWS, Azure, and on-prem environments * Establish IaC standards, module versioning strategy, state management practices, and review processes across engineering teams * Drive migration of manually managed infrastructure to fully codified, version-controlled definitions Deployment Pipelines & Containerized Services * Architect and implement CI/CD pipelines for containerized and Kubernetes-based services, from build through progressive production rollout * Define deployment strategies (blue/green, canary, rolling) and the automation that supports them * Partner with development teams to streamline the path from commit to production while maintaining safety and auditability Configuration Management * Own configuration management tooling and practices across hybrid environments (e.g., Ansible, Chef, Puppet, or equivalent) to ensure consistency, repeatability, and drift detection * Standardize secrets management, environment configuration, and golden image/baseline practices across cloud and on-prem fleets Observability * Design and implement observability patterns (metrics, logging, tracing) that provide actionable signal across distributed, hybrid-cloud services * Define SLIs/SLOs in partnership with SRE and product teams, and build the dashboards and alerting that make them actionable * Reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR) through better instrumentation, not just more of it Production Reliability * Act as a senior escalation point for complex production incidents, driving root cause analysis and durable remediation * Lead or contribute to postmortems and translate findings into infrastructure, process, or tooling improvements * Proactively identify and remediate reliability risk before it becomes an incident Cost & Security Governance * Design and implement cost governance practices - tagging standards, budget alerting, rightsizing, and reserved capacity strategy across AWS, Azure, and on-prem infrastructure * Partner with Security to define and enforce infrastructure security guardrails (IAM least-privilege, network segmentation, secrets handling, compliance controls) * Build automated policy enforcement (e.g., policy-as-code) so governance scales with infrastructure rather than depending on manual review Cross-Team Leadership * Serve as the primary point of contact between Platform Engineering and SRE teams, aligning on standards, priorities, and shared tooling * Mentor senior and mid-level engineers on infrastructure design, operational excellence, and IaC best practices * Influence infrastructure architecture and technical roadmap at the organizational level ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [One Platform Could Not Fit Them All](https://www.wearedevelopers.com/videos/1919-one-platform-could-not-fit-them-all) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)