Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
This is a global materials science leader with a 170 plus year history, operating dozens of manufacturing and R&D sites worldwide. This role sits within a research and development group supporting advanced computing infrastructure behind ongoing materials science innovation.
Top 3 Skills
- Kubernetes cluster operations and management, including provisioning, upgrades, and troubleshooting across on premises and cloud environments
- Rancher for Kubernetes platform management
- Linux systems administration, including performance tuning, scripting, and networking
What You’ll Do
- Maintain and enhance Kubernetes platforms across on premises and cloud environments
- Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher
- Provide deep Linux systems administration support, including performance tuning, troubleshooting, and automation
- Develop and maintain infrastructure as code solutions to standardize and automate platform deployment
- Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration
- Collaborate with developers, scientists, and infrastructure teams to deliver reliable platform services
- Identify opportunities to improve platform resilience, observability, security, and maintainability
Requirements
- 5 plus years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering
- Hands on experience operating Kubernetes platforms in production environments, both on premises and cloud based
- Experience with Rancher for Kubernetes cluster management
- Strong Linux systems administration skills, including troubleshooting, scripting, and system performance analysis
- Experience implementing infrastructure as code solutions for platform provisioning and lifecycle management
Preferred/Bonus
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or related field
- Experience with Cluster API (CAPI)
- Experience with hybrid infrastructure spanning on premises and public cloud (AWS, Azure, Google Cloud Platform)
- Familiarity with Kubernetes observability, logging, monitoring, and alerting tooling
- Experience supporting scientific research, high performance computing, or computational science environments
- Experience with Agile teams (Scrum, Kanban)
About the company
Elevait Solutions was founded by veterans who believe that how you treat people is the only thing that actually matters in this industry. We’re a team that stays in your corner before, during, and after placement. We show up for the communities we work in, we tell you the truth, and we work hard to make sure every placement is a good fit for both sides. If that sounds like the kind of team you want behind you, we’d like to talk.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Why Upskilling And Reskilling is Important For Developers
Find a Developer Job: 12 Best Job Sites For Developers