> Markdown version of [/jobs/ext/2071398-platform-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2071398-platform-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Reliability Engineer - **Company:** Bright Vision Technologies - **Location:** Bellevue, WA, United States (Remote available) - **Salary:** $100,000.0 - $150,000.0 - **Contract:** Temporary to permanent - **Skills:** Java (Programming Language), Amazon Web Services, Microsoft Azure, Computer Programming, Linux, DevOps, Distributed Systems, Python (Programming Language), Load Testing, Performance Tuning, Reliability Engineering, Prometheus, Software Engineering, Datadog, Cloud Platform System, Istio, System Availability, Grafana, Kubernetes, Information Technology, Linkerd (Service Mesh) - **Published:** August 15, 2026 - **Apply:** https://www.careerjet.com/jobad/us23dbc11b60a75b7bb6f4a66d93d14dd1 ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related technical discipline. * Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems. * Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling. * Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting. * Production experience operating Kubernetes and container-based workloads. * Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents. * Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications. * Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics. * Demonstrated experience leading incident response and conducting effective post-incident reviews. * Excellent communication and documentation skills., * Experience defining and operationalizing SLOs and error budgets in real production environments. * Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. * Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP). * Background in capacity planning, performance engineering, or large-scale load testing. * Familiarity with service mesh technologies such as Istio, Linkerd, or Consul. ## Description We are seeking an experienced Platform Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern., Title: Reliability and Maintenance Engineer Location: Redmond, WA Duration: 6 months, possible extensions Compensation: $55-60hr Work Requirements: Export Control Requirement: … + 1 day ago + ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)