> Markdown version of [/jobs/ext/1312845-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1312845-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Latitude Inc - **Location:** Pittsburg, CA, United States - **Experience:** Expert - **Salary:** $179,200.0 - $268,800.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computer Engineering, Software Debugging, Linux, Elasticsearch, Python (Programming Language), Machine Learning, Motion Planning, Reliability Engineering, Prometheus, Subsystems, Transmission Control Protocol (TCP), Data Logging, Data Processing, Google Cloud, Cloud Platform System, Cloudformation, Kubernetes, Information Technology, Terraform - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/a80b39f4-dddb-48a5-906b-94055aa577f2 ## About the Role * Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, Robotics or a related field and 4+ years of relevant experience (or Master's degree and 2+ years of relevant experience, or PhD) * Fundamental understanding of Linux operating system internals, TCP/IP networking, and storage subsystems * Hands on development in Go or Python to create robust software that can run reliably in production * Strong experience scaling and securing services in the cloud (AWS, Google Cloud Platform) or cloud native environments * Experience using infrastructure-as-code principles to automate the creation of infrastructure resources (e.g. Terraform, CloudFormation) * Experience authoring and maintaining Kubernetes Controllers in Go * Experience running Kubernetes and related core components in a large-scale, production environment * Experience with metrics (e.g. Prometheus), logging (e.g. Elasticsearch, Loki) and tracing (e.g. Jaeger, Tempo) systems * Understanding of engineering design limitations and ability to provide guidance to teams to scale their services to achieve desired performance within budget * A focus on increasing service reliability through defining and adhering to SLOs * Strong communication skills and the ability to work effectively in a diverse and distributed team ## Description As a Site Reliability Engineer on the team, you will be responsible for helping to build and run these mission critical systems. Through the implementation of monitoring and automation, you will constantly ensure the health, reliability, scalability, and performance of the platforms. The Site Reliability team interacts with engineering teams including ingest/data processing, mapping, labeling, triage, machine learning (detection, prediction, tracking), motion planning/control, offline simulation, and release/deployment teams to provide uniform service observability and incident response. What you'll do: * Build monitoring to ensure our platform is healthy and its reliability measurable * Build alerting and a set of runbooks to enable faster detection and remediation of platform issues * Debug complex issues that may combine multiple components of the stack and ensure proper fixes are implemented to prevent these issues from happening again * Participate in an on-call rotation and culture of continuous improvement through blameless postmortems * Design and implement components of the platform to enable features that make the work of our customers possible, simpler and more efficient * Build Kubernetes controllers to automate operations ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london)