DevOps / Site Reliability Engineer
Bayside Solutions
Cupertino, CA, United States
13 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source
Tech stack
Cloud Computing
Continuous Integration
DevOps
Distributed Systems
Monitoring of Systems
Python (Programming Language)
Network Troubleshooting
Log Analysis
Reliability Engineering
Prometheus
Datadog
Data Logging
+6 more
Computer Networking Systems
Grafana
Kubernetes
Deployment Automation
Splunk
Golang
Job description
We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services., * Design, build, automate, and support scalable Kubernetes-based platforms and services
- Operate and troubleshoot production environments running at scale
- Develop automation and tooling to improve operational efficiency and reliability
- Monitor platform health, performance, and availability using observability tooling
- Troubleshoot infrastructure, application, and networking issues across distributed systems
- Work closely with engineering teams to improve deployment, reliability, and scalability practices
- Participate in operational support, incident response, and root cause analysis
- Improve CI/CD workflows and deployment automation
- Drive operational excellence through documentation, automation, and process improvements
- Take ownership of projects and independently drive deliverables to completion
Requirements
- Strong hands-on experience with Kubernetes platforms such as EKS, GKE, and AKS or similar
- Experience running and supporting applications on Kubernetes at scale
- Strong understanding of containerized infrastructure and distributed systems
- Experience with monitoring and observability tools, preferably Grafana or Prometheus
- Experience with CI/CD pipelines and deployment automation
- Experience with Splunk logging, log analysis, and troubleshooting
- Strong scripting and automation experience using Python and/or Golang
- Experience troubleshooting production systems under pressure
- Strong communication and collaboration skills
- Self-starter mentality with strong ownership and accountability, * Experience operating Ray clusters/services
- Strong networking and troubleshooting experience
- Experience with cloud infrastructure and platform services
- Experience with Infrastructure as Code and automation frameworks
- Experience supporting high-scale production systems
- Familiarity with SRE principles and operational best practices
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago