> Markdown version of [/jobs/ext/2642097-telecommute-infrastructure-cloud-devops-sre](https://www.wearedevelopers.com/jobs/ext/2642097-telecommute-infrastructure-cloud-devops-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Infrastructure/Cloud DevOps-SRE - **Company:** Bayside Solutions - **Location:** Cupertino, CA, United States (Remote available) - **Salary:** $114,400.0 - $135,200.0 - **Contract:** Permanent contract - **Skills:** Cloud Computing, Continuous Integration, DevOps, Distributed Systems, Monitoring of Systems, Python (Programming Language), Network Troubleshooting, Log Analysis, Reliability Engineering, Prometheus, Datadog, Data Logging, Computer Networking Systems, Grafana, Kubernetes, Deployment Automation, Splunk, Golang - **Published:** August 31, 2026 - **Apply:** https://www.dice.com/job-detail/d2e623ec-fc77-4abc-b66b-0da5e3ed8091 ## About the Role * Strong hands-on experience with Kubernetes platforms such as: + EKS + GKE + AKS or similar * Experience running and supporting applications on Kubernetes at scale * Strong understanding of containerized infrastructure and distributed systems * Experience with monitoring and observability tools, preferably: + Grafana + Prometheus * Experience with CI/CD pipelines and deployment automation * Experience with Splunk logging, log analysis, and troubleshooting * Strong scripting and automation experience using Python and/or Golang * Experience troubleshooting production systems under pressure * Strong communication and collaboration skills * Self-starter mentality with strong ownership and accountability, * Experience operating Ray clusters/services * Strong networking and troubleshooting experience * Experience with cloud infrastructure and platform services * Experience with Infrastructure as Code and automation frameworks * Experience supporting high-scale production systems * Familiarity with SRE principles and operational best practices ## Description We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services., * Design, build, automate, and support scalable Kubernetes-based platforms and services * Operate and troubleshoot production environments running at scale * Develop automation and tooling to improve operational efficiency and reliability * Monitor platform health, performance, and availability using observability tooling * Troubleshoot infrastructure, application, and networking issues across distributed systems * Work closely with engineering teams to improve deployment, reliability, and scalability practices * Participate in operational support, incident response, and root cause analysis * Improve CI/CD workflows and deployment automation * Drive operational excellence through documentation, automation, and process improvements * Take ownership of projects and independently drive deliverables to completion ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023)