Cloud Systems Engineer

Lunar Outpost
United States
3 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services DevOps Domain Name System (DNS) Github Identity and Access Management Key Management Uptime Site Reliability Engineering Practices Cloud Services Load Balancing Autoscaling Kubernetes Helm Charts
+8 more
Amazon Virtual Private Cloud (VPC) Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Performance Monitor Terraform Jenkins

Job description

  • Own and manage Stargate production releases and deployment pipelines using GitOps practices

  • Drive operational excellence initiatives including metrics collection, log aggregation, uptime monitoring, KPI tracking, and SIM Integration

  • Maintain and achieve 99.99% (four nines) to 99.999% (five nines) uptime SLAs

  • Design, develop, and maintain Helm charts for Stargate and related infrastructure components

  • Implement and manage progressive deployment strategies including canary deployments and blue-green deployments

  • Oversee critical Kubernetes infrastructure including volume management, DNS configuration, load balancer provisioning, and secret monitoring/management

  • Manage and optimize Kubernetes deployments and related AWS services

  • Implement and maintain observability stack using OpenTelemetry for comprehensive monitoring and alerting

  • Collaborate with engineering teams to establish and enforce operational best practices and reliability standards

Requirements

  • 5+ years of production DevOps/SRE experience with demonstrable track record of maintaining high-availability systems

  • Kubernetes administration experience with elevated cluster access in production environments

  • Strong proficiency writing and maintaining Helm charts for complex, multi-component applications

  • Hands-on experience implementing canary deployments, blue-green deployments, and other progressive delivery patterns

  • Deep knowledge of Kubernetes infrastructure management: persistent volumes, DNS/networking, load balancers, and secrets management

  • Production experience with GitOps workflows and Flux CD

  • Proven track record maintaining 99.99%+ uptime in production environments

  • Excellent judgment and decision-making skills when working with production systems

Preferred Qualifications:

  • Experience with AWS cloud services, particularly EKS (Elastic Kubernetes Service), Secrets Manager, VPC networking, IAM, and AWS Load Balancers

  • Experience with Karpenter for Kubernetes node autoscaling and cluster optimization

  • Experience with OpenTelemetry instrumentation and observability platforms

  • Kubernetes certifications (CKA, CKAD, or CKS)

  • Experience building and maintaining CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.)

  • Knowledge of infrastructure-as-code tools (Terraform, CDK)

  • Experience implementing SRE practices including SLIs, SLOs, and error budgets

Benefits & conditions

Compensation & Benefits: Compensation level and base salary are competitively structured and thoughtfully determined based on factors such as relevant skills, experience, education, and the scope of the role.

  • Comprehensive health coverage: Medical, dental, and vision benefits, with 70% of premiums covered by the employer
  • Paid time off: Three (3) weeks per year of vacation
  • Retirement plan: Up to 4% employer match on 401(k) contributions
  • Paid holidays: 11 company-recognized holidays
  • Parental leave
  • Educational reimbursement opportunities to support company objectives, continued learning, and career development

About the company

Are you passionate about shaping the future of humanity’s presence in space? Lunar Outpost, an industry leader in space robotics and planetary vehicles, invites you to join our team! Lunar Outpost is dedicated to creating a permanent presence in space, while also driving positive impacts here on Earth. Currently, we are seeking a Senior Cloud Systems Engineer to contribute to our mission in a dynamic startup environment. The main responsibilities of this role include managing Stargate deployments in production, ensuring high availability and uptime, executing reliable releases, and driving operational excellence through comprehensive monitoring, metrics, and infrastructure management. Stargate is a next-generation Command and Control (C2) platform-the ground software that enables and empowers all Lunar Outpost missions, including the Lunar Terrain Vehicle (LTV) program. As mission-agnostic software used by all operators in mission control, Stargate’s reliability and uptime are critical to mission success.

Take the #NextLeap with Lunar Outpost and work on the Pegasus LTV, which will carry NASA astronauts farther than they’ve ever been before on the lunar surface!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:22 min

Leveraging unique cultural backgrounds in engineering design

Ixchel Ruiz · LIVE

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

9:06 min

Questions on career paths and continuous delivery orchestration platforms

Zan Markan Zan Markan · LIVE

Videos

See all

Related articles

See all