Site Reliability Engineer (On Prem)
Role details
Job location
Tech stack
Job description
Insight Global is looking for a Site Reliability Engineer (on prem) to join one of our largest clients in the Bay Area. This person will:
-Provide day-to-day operational support for production and pre-production environments
-Administer and support Jenkins (user management, pipeline reliability, upgrades, troubleshooting)
-Manage and operate Kubernetes clusters at scale (deployments, scaling, upgrades, monitoring)
-Maintain system reliability, uptime, and performance through proactive monitoring and incident response
-Participate in on-call rotations to support critical systems and resolve production issues
-Work closely with engineering teams to support CI/CD and platform reliability
-Support on-site operations 3-4 days per week (collaboration with infra and hardware teams)
Requirements
5+ years of experience in SRE, DevOps, or Systems Engineering roles
-Strong hands-on experience with Kubernetes in production
-Strong hands-on experience with Jenkins administration and CI/CD operations
-Experience supporting Linux-based systems in high-availability environments
-Comfort operating and troubleshooting complex infrastructure under SLA pressure -Experience supporting GPU-based infrastructure (NVIDIA GPUs or similar)
-Hands-on exposure to GPU scheduling, health monitoring, and workload reliability
-Supporting ML/AI, compute-intensive, or accelerator-based workloads is a strong plus
-Familiarity with GPU drivers, firmware, and integration within Kubernetes environments preferred
Benefits & conditions
This role can pay between $60-$75/hour depending on years of experience + skillset.