> Markdown version of [/jobs/ext/2709128-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2709128-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** SKY I.T. SOLUTIONS, Inc. - **Location:** San Mateo, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $240,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Information Systems, Databases, Continuous Integration, Linux, DevOps, Distributed Systems, Github, Identity and Access Management, Subnetting, Python (Programming Language), PostgreSQL, Octopus Deploy, Reliability Engineering, Software Deployment, Spinnaker, Data Streaming, Technical Data Management Systems, Datadog, Load Balancing, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Terraform, Jenkins - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-site-reliability-engineer-skydio-com-8994483 ## About the Role * 8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer or equivalent infrastructure role. * Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters. * Experience managing Kubernetes/EKS upgrades and production clusters. * Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases. * Production experience with Terraform or similar infrastructure-as-code tooling. * Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins. * Experience diagnosing production infrastructure and networking problems. * Experience solving meaningful scaling or reliability challenges. * This position requires access to export-controlled technical data, restricted government information, and/or information systems subject to U.S. government security and access-control requirements. Employment in this role is contingent upon verification of U.S. person status and the ability to access controlled or restricted information as required for the position. Bonus points: * Helm and GitOps experience. * Datadog or similar observability tooling. * PostgreSQL/database operations experience. * Multi-region infrastructure experience. * On-premises or disconnected deployment experience. * Streaming or high-throughput distributed systems experience. ## Description We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability. You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain. How you'll make an impact: * Build, operate, and troubleshoot production Kubernetes/EKS clusters. * Perform Kubernetes upgrades, node rollouts, and cluster maintenance. * Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage. * Define and maintain infrastructure using Terraform. * Build and operate CI/CD and deployment infrastructure. * Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases. * Build monitoring, alerting, and observability for critical infrastructure. * Participate in on-call rotations and respond to production incidents. * Identify and solve infrastructure scaling and reliability problems. * Automate operational work using Python, Go, or similar languages. * Help expand infrastructure across new regions and deployment environments. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)