DevOps / Site Reliability Engineer

Bayside Solutions
Cupertino, CA, United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source

Tech stack

Cloud Computing Continuous Integration DevOps Distributed Systems Monitoring of Systems Python (Programming Language) Network Troubleshooting Log Analysis Reliability Engineering Prometheus Datadog Data Logging
+6 more
Computer Networking Systems Grafana Kubernetes Deployment Automation Splunk Golang

Job description

We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services., * Design, build, automate, and support scalable Kubernetes-based platforms and services

  • Operate and troubleshoot production environments running at scale
  • Develop automation and tooling to improve operational efficiency and reliability
  • Monitor platform health, performance, and availability using observability tooling
  • Troubleshoot infrastructure, application, and networking issues across distributed systems
  • Work closely with engineering teams to improve deployment, reliability, and scalability practices
  • Participate in operational support, incident response, and root cause analysis
  • Improve CI/CD workflows and deployment automation
  • Drive operational excellence through documentation, automation, and process improvements
  • Take ownership of projects and independently drive deliverables to completion

Requirements

  • Strong hands-on experience with Kubernetes platforms such as EKS, GKE, and AKS or similar
  • Experience running and supporting applications on Kubernetes at scale
  • Strong understanding of containerized infrastructure and distributed systems
  • Experience with monitoring and observability tools, preferably Grafana or Prometheus
  • Experience with CI/CD pipelines and deployment automation
  • Experience with Splunk logging, log analysis, and troubleshooting
  • Strong scripting and automation experience using Python and/or Golang
  • Experience troubleshooting production systems under pressure
  • Strong communication and collaboration skills
  • Self-starter mentality with strong ownership and accountability, * Experience operating Ray clusters/services
  • Strong networking and troubleshooting experience
  • Experience with cloud infrastructure and platform services
  • Experience with Infrastructure as Code and automation frameworks
  • Experience supporting high-scale production systems
  • Familiarity with SRE principles and operational best practices

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all