Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
- Help design and build Aalyria’s centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry).
- Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
- Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries).
- Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD).
- Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
- Develop and lead the company’s monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
Requirements
This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you., * Active Top Secret (TS/SCI) security clearance.
- 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.
- Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
- Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
- Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
- Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
- Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems., * Experience operating a multi-cloud environment, specifically GCP and AWS.
- Hands-on experience with GitLab CI for CI/CD pipelines.
- Working knowledge of service mesh technologies such as Istio or Linkerd.
- Familiarity with instrumenting applications written in Go and C++.
- An active Secret clearance, or higher, is preferred for this position.
- Experience with JVM observability (tuning, monitoring) for Java-based applications.
Benefits & conditions
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance. The summary above was generated by AI About Aalyria, * Innovative Environment: Work at a cutting-edge company shaping the future of aerospace communications.
- Impactful Work: Directly contribute to critical national security programs and initiatives.
- Growth Opportunities: Expand your career with opportunities for professional development and advancement.
- Inclusive Culture: Be part of a collaborative, supportive, and inclusive workplace where your contributions matter.
- Flexibility: Flexible working arrangements including hybrid remote/in-office schedules.
- Compensation and Equity: Competitive salary, comprehensive benefits (401(k), dental, vision, health, life insurance), paid time off, and equity options.
ITAR/EAR Requirements:
This position involves access to export-controlled information. To comply with U.S. government export regulations, applicants must meet one of the following criteria:
(A) Qualify as a U.S. person, which includes:
- U.S. citizen or national
- U.S. lawful permanent resident (green card holder)
- Refugee under 8 U.S.C. 1157
- Asylee under 8 U.S.C. 1158
About the company
Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. With technology acquired from Google, Aalyria is at the forefront of innovation in satellite and airborne mesh networks, as well as cislunar and deep-space communications. We are revolutionizing the orchestration and management of planetary mesh networks using any radio or optical spectrum, any orbit, and any hardware across land, sea, air, and space.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
Find a Developer Job: 12 Best Job Sites For Developers