> Markdown version of [/jobs/ext/3073246-cloud-systems-engineer](https://www.wearedevelopers.com/jobs/ext/3073246-cloud-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Systems Engineer - **Company:** Lunar Outpost - **Location:** Arvada, CO, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, DevOps, Domain Name System (DNS), Github, Identity and Access Management, Key Management, Uptime, Site Reliability Engineering Practices, Cloud Services, Load Balancing, Cloud Platform System, Autoscaling, Kubernetes Helm Charts, Amazon Virtual Private Cloud (VPC), Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Performance Monitor, Terraform, Jenkins - **Published:** September 25, 2026 - **Apply:** https://www.juju.com/job/16_9210b5d41 ## About the Role * 5+ years of production DevOps/SRE experience with demonstrable track record of maintaining high-availability systems * Kubernetes administration experience with elevated cluster access in production environments * Strong proficiency writing and maintaining Helm charts for complex, multi-component applications * Hands-on experience implementing canary deployments, blue-green deployments, and other progressive delivery patterns * Deep knowledge of Kubernetes infrastructure management: persistent volumes, DNS/networking, load balancers, and secrets management * Production experience with GitOps workflows and Flux CD * Proven track record maintaining 99.99%+ uptime in production environments * Excellent judgment and decision-making skills when working with production systems * Familiarity with Aetherflux for declarative canary orchestration and progressive rollout gating Preferred Qualifications: * Experience with AWS cloud services, particularly EKS (Elastic Kubernetes Service), Secrets Manager, VPC networking, IAM, and AWS Load Balancers * Experience with Karpenter for Kubernetes node autoscaling and cluster optimization * Experience with OpenTelemetry instrumentation and observability platforms * Kubernetes certifications (CKA, CKAD, or CKS) * Experience building and maintaining CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.) * Knowledge of infrastructure-as-code tools (Terraform, CDK) * Experience implementing SRE practices including SLIs, SLOs, and error budgets ## Description * Own and manage Stargate production releases and deployment pipelines using GitOps practices * Drive operational excellence initiatives including metrics collection, log aggregation, uptime monitoring, KPI tracking, and SIM Integration * Maintain and achieve 99.99% (four nines) to 99.999% (five nines) uptime SLAs * Design, develop, and maintain Helm charts for Stargate and related infrastructure components * Implement and manage progressive deployment strategies including canary deployments and blue-green deployments * Oversee critical Kubernetes infrastructure including volume management, DNS configuration, load balancer provisioning, and secret monitoring/management * Manage and optimize Kubernetes deployments and related AWS services * Implement and maintain observability stack using OpenTelemetry for comprehensive monitoring and alerting * Collaborate with engineering teams to establish and enforce operational best practices and reliability standards ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)