Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Kelly Services Inc.
United States
24 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Bash Shell Cloud Computing Computer Programming Continuous Integration Linux Distributed Systems Python (Programming Language) Linux System Administration Reliability Engineering Prometheus
+10 more
Azure Machine Learning AI Infrastructure Scripting Cloud Platform System Grafana Containerization Kubernetes Hardware Infrastructure Terraform Golang

Job description

Join a rapidly growing B2B AI infrastructure company powering large-scale machine learning and AI workloads for more than one million developers worldwide. As a Site Reliability Engineer, you’ll help improve the reliability, scalability, and performance of a cloud platform built on Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation technologies. This full-time remote opportunity offers the chance to work on critical infrastructure supporting AI applications on a global scale.

As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization. This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.

Requirements

  • 5+ years of experience within major public cloud environment like AWS, GCP

  • Strong Linux systems administration experience

  • Strong networking fundamentals and troubleshooting skills

  • Experience supporting containerized environments (Kubernetes preferred)

  • Experience with monitoring, alerting, and observability tools

  • Experience defining and managing SLIs, SLOs, and reliability metrics

  • Incident response and postmortem experience

  • Scripting or programming experience. Python, Go, Bash, or similar technologies

  • Distributed systems and failure scenarios

Desired Skills & Experience

  • Kubernetes

  • Prometheus, Grafana, or similar monitoring platforms

  • Experience supporting GPU infrastructure or AI/ML platforms

  • Infrastructure as Code experience (Terraform preferred)

  • CI/CD pipeline experience, Applicants must be currently authorized to work in the US on a full-time basis now and in the future. Sponsorship is not available for this position

About the company

Motion Recruitment Partners (MRP) is an Equal Opportunity Employer. All applicants must be currently authorized to work on a full-time basis in the country for which they are applying, and no sponsorship is currently available. Employment is subject to the successful completion of a pre-employment screening. Accommodation will be provided in all parts of the hiring process as required under MRP’s Employment Accommodation policy. Applicants need to make their needs known in advance.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Loading talks and stories from around this role…