Site Reliability Engineer (On Prem)

Insight Global
Santa Clara, CA, United States
about 2 months ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$124,800.0 - $156,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Continuous Integration DevOps Firmware Linux System Administration Reliability Engineering Graphics Processing Unit (GPU) Reliability of Systems Kubernetes Performance Monitor User Administration
+1 more
Jenkins

Job description

Insight Global is looking for a Site Reliability Engineer (on prem) to join one of our largest clients in the Bay Area. This person will:

-Provide day-to-day operational support for production and pre-production environments

-Administer and support Jenkins (user management, pipeline reliability, upgrades, troubleshooting)

-Manage and operate Kubernetes clusters at scale (deployments, scaling, upgrades, monitoring)

-Maintain system reliability, uptime, and performance through proactive monitoring and incident response

-Participate in on-call rotations to support critical systems and resolve production issues

-Work closely with engineering teams to support CI/CD and platform reliability

-Support on-site operations 3-4 days per week (collaboration with infra and hardware teams)

Requirements

5+ years of experience in SRE, DevOps, or Systems Engineering roles

-Strong hands-on experience with Kubernetes in production

-Strong hands-on experience with Jenkins administration and CI/CD operations

-Experience supporting Linux-based systems in high-availability environments

-Comfort operating and troubleshooting complex infrastructure under SLA pressure -Experience supporting GPU-based infrastructure (NVIDIA GPUs or similar)

-Hands-on exposure to GPU scheduling, health monitoring, and workload reliability

-Supporting ML/AI, compute-intensive, or accelerator-based workloads is a strong plus

-Familiarity with GPU drivers, firmware, and integration within Kubernetes environments preferred

Benefits & conditions

This role can pay between $60-$75/hour depending on years of experience + skillset.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all