> Markdown version of [/jobs/ext/2590659-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2590659-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** AALYRIA TECHNOLOGIES, INC. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $125,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, C++ (Programming Language), Software Debugging, Distributed Systems, Java Virtual Machine (JVM), Java Web Services, Performance Tuning, Prometheus, System Programming, Data Logging, Computer Networking Systems, Google Cloud, Istio, System Availability, Grafana, Multi-Cloud, Infrastructure as Code (IaC), SC Clearance, Gitlab-ci, Kubernetes, Linkerd (Service Mesh), Terraform, Dynatrace - **Published:** August 28, 2026 - **Apply:** https://ats.rippling.com/aalyria-careers/jobs/27e9f512-e14b-4ba4-aefb-0f05a63fa21a ## About the Role This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you., * Active Top Secret (TS/SCI) security clearance. * 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems. * Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems. * Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes. * Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD). * Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling. * Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems., * Experience operating a multi-cloud environment, specifically GCP and AWS. * Hands-on experience with GitLab CI for CI/CD pipelines. * Working knowledge of service mesh technologies such as Istio or Linkerd. * Familiarity with instrumenting applications written in Go and C++. * An active Secret clearance, or higher, is preferred for this position. * Experience with JVM observability (tuning, monitoring) for Java-based applications. ## Description * Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). * Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready. * Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries). * Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD). * Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments. * Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)